The Schema Nobody Owns

The working assumption in most organizations right now is that the data is ready.

The tools are live. The dashboards are populated. The AI copilot is answering questions. So the assumption follows naturally: the data underneath all of it must be in reasonable shape, or none of this would be working.

That assumption is the problem.

What most organizations have built is a data environment that functions well enough to produce output, but not well enough to be trusted. The difference between those two things is definition: what a field actually means, who decided it, and whether anyone in the organization can answer those questions without a two-day archaeology project.

They usually can't.

I've been in sessions where a room full of executives is looking at two reports that contradict each other. Same time period. Same system. Different numbers. And when someone finally digs into it, the answer is almost always the same: two teams defined the same metric differently, at different points in time, for different purposes, and neither definition was ever formally documented. Both queries were technically correct. Neither answer was trustworthy.

That's not a data quality problem. It's an ownership problem.

Every field in an enterprise data environment represents a decision: what this thing is called, what it contains, what it excludes, who defined it, and when. In a governed environment, those decisions are documented, deliberate, and owned. In most enterprise environments, they accumulated. Someone made a naming decision seven years ago for a reason that made sense at the time. That person left. The field stayed. The reason disappeared.

Now an AI agent is reading that field and drawing conclusions from it.

This is the gap that exists underneath many AI deployments. The technical infrastructure exists. The accountability layer does not. Nobody has been explicitly assigned to own what the data means, which means nobody can be held accountable when the meaning drifts.

And it always drifts.

Three things happen in organizations that close this gap, and none of them are purely technical decisions:

First, they treat data definitions as a business asset, not an IT byproduct. The meaning of a field belongs to the business unit that uses it, not the team that built the table. Ownership follows accountability. If Finance owns the revenue metric, Finance owns the definition of revenue. That means a Finance leader, not a data engineer, signs off on what counts as recognized revenue, what gets excluded, and when the definition changes. The data team documents and enforces it. The business unit defines it.

Second, they establish a change protocol. Definition changes don't happen silently. When a field is renamed, deprecated, or redefined, there's a documented process, a notification chain, and an impact assessment. Most organizations have change management processes for software releases. Almost none apply the same discipline to the data that software reads. A field renamed with no downstream communication is an incident waiting to surface three months later in a board report.

Third, they build for discoverability. A definition that lives in one person's head, or in a spreadsheet someone emailed three years ago, is not a governed definition. It's institutional memory waiting to retire. Governed organizations make their data definitions findable, versioned, and auditable. Not because auditors will ask, but because agents will query. If the system can't find the definition, it will invent one.

The reason this matters now, more than at any prior point in the data governance conversation, is agentic AI. A human analyst encountering an ambiguous field will pause. They'll ask a colleague. They'll flag the discrepancy in their output. An AI agent will not. It will make an inference, proceed with full confidence, and scale that inference across every downstream decision it influences. By the time the error surfaces, the agent has already acted on it dozens of times, across dozens of outputs, some of which triggered real business actions.

The meaning layer nobody owns is the foundation the agent is standing on.

Most governance conversations get stuck on access control: who can see what. That's necessary, but it's not sufficient. The harder question, and the one most organizations haven't answered, is who decides what the data means, and how does that decision get enforced when a machine is the one doing the reading.

Ownership is not a technical problem. It's a governance decision that most organizations have deferred because nobody made it urgent.

AI just made it urgent.

If you found this briefing valuable, share it with a colleague who is navigating the shift from AI hype to operational reality.

Previous
Previous

The Pipeline That Lied

Next
Next

The Lineage Nobody Is Tracing