Skip to main content

The Operating Model for AI-Generated Context

Teams at ASOS, Mastercard and Sophos found that writing context is no longer the bottleneck. What that context depends on still is.

Nandini Tyagi

Nandini Tyagi

Customer Experience Manager, Strategic Initiatives

October 1, 2026·13 min read

Earlier this month, I talked with three people who each used AI to document their company’s data at a scale their teams could never have staffed by hand. Zhenni Hu is a manager in AI & data governance at Mastercard. Rupal Sumaria heads data management at ASOS. Keith Guyett is a principal solutions architect at Sophos.

Before Zhenni said anything about what her team had built, she asked the audience to answer two questions in the chat.

Have you used AI at work this week? Thumbs-ups came up quickly, and a lot of them.

Would you be comfortable letting AI make a decision you were held accountable for?

The yes responses quickly stopped, which was what she had expected. “Not surprisingly,” she said, “I didn’t see any thumbs up come in yet.”

Everyone uses AI, but few data practitioners will sign their name to what it produces. The models were not the question, and nobody spent a minute arguing about them. The question was whether the material underneath could be vouched for.

Metadata was built for humans, but now AI uses it just as much. This means that the data catalog is changing, as Prukalpa recently argued. Today agents write context, they read it back at runtime, and quality compounds along the way, because better evidence at one layer produces better context at the next. But agents producing context at scale is not the same thing as that context holding up when other agents have to make a decision with it.

That’s the question these three data experts have been living in: not whether agents can write context, which was a unanimous yes, but what it takes to get context you would actually put your name on.


The problem: Metadata built for speed, not agents

Sophos had more than 600,000 assets in its catalog by March. When the team started enriching, the first nine thousand went quickly. This sounds like a good thing, but speed turned out to be the wrong thing to optimize for.

“Doing it fast was not the correct answer,” Keith told us. “Doing it fast just gave us static restatement of the underlying data model. It did not provide any business purpose.”

A restatement is what comes back when the model has nothing to read except the schema. It says that the customer_id column contains customer identifiers, and when a reviewer skims it, it says that nothing in it is wrong. And yet that doesn’t help anyone decide anything.

Keith put the underlying reason more precisely than I had heard anyone put it.

“A good enough description for a human is not necessarily a good enough description for an AI agent, because the human has that business knowledge ingrained. The AI agent has to reason through it.”

Think about what a person does with a thin description. They read that a table holds order data, and they already know which of the three near-identical order tables is the one finance trusts. They know the field changed meaning after the migration two years ago. They know that when risk says exposure they mean something narrower than what the commercial team means. The description was never carrying the meaning by itself. It was a pointer, and the reader supplied the rest from memory.

An agent has no memory to supply. It reasons over what is written down, and only that.

A description that was perfectly adequate for a decade is not adequate now, and neither is anything built on top of it. The standard moved because the reader changed, and most estates are full of documentation that doesn’t make sense to AI.

Side-by-side comparison of a human reader and an agent reader looking at the same order data record. The human reader resolves ambiguity by drawing on the trusted table, the migration history and the business meaning. The agent reader has the same record but three empty question-mark boxes in place of that knowledge, and arrives at an ambiguous result.

The record is identical. Only one of the two readers brings anything to it. | Source: Context and Chaos


Widen the control scope to add tribal knowledge

The first phase of an agentic data catalog is creating context for agents, but as these three practitioners found, it’s not easy. The first idea that they used to solve agentic metadata creation was control scope, or what the agent sees and uses to create data descriptions.

Sophos fixed its first run by changing the inputs rather than the order of the work. The model was told to give each description a business purpose, to state the grain and scope, to look at metric relevance, to use examples where they existed, and to read the lineage and source context, which Keith called the “metadata around the metadata”. With all of that information, the same model on the same assets produced something a reviewer could work with.

Zhenni from Mastercard had reached the same conclusion from the other direction. If the agent sees only column names, what comes back is a polished restatement of the schema, which is not what a steward needs. What made the drafts usable at Mastercard was reasoning over real signals: how the data actually gets queried, what it connects to upstream and downstream, and the glossary terms, classifications and data products that already existed.

The difference between a generic description and a useful one is not the model. It is what the model had to work from. And the economics follow from that: correcting a draft that is already close takes minutes, while rewriting one that is too generic takes as long as writing it yourself, which is the job nobody had time for in the first place.

Sophos went back and re-enriched some of those first nine thousand, added more on top, and reached almost 43,000 enriched assets between April and early May. They are near 100,000 now.


Think left to right when building a context pipeline

The second idea these practitioners talked about is the importance of the order in which data assets are enriched before they are handed over to agents for consumption, which is the second phase of an agentic data catalog.

Rupal from ASOS put it: “I prefer following the chain of my pipeline because I think you build knowledge and transformation on top of each other… For me, it’s left to right.”

Descriptions improve at the source layer. Those descriptions are what the semantic model gets built from. The semantic model is what grounds the metric layer in the formulas the business actually uses, rather than the ones someone assumed. Each layer is the evidence the next layer reasons over, which means the evidence problem and the ordering problem are the same problem seen twice, once inside a single run and once across the stack.

Four stages running left to right. Source descriptions carrying business purpose, grain and scope, usage and lineage. Then the semantic model with entities, relationships and definitions. Then the metric layer with approved formulas, ownership and business scope. Then the agent producing an answer, a decision and an action.

Each stage is built from the one before it, which is why the order cannot be rearranged. | Source: Context and Chaos

Mastercard ran the same sequence as program order: build the foundation, enrich it, make it discoverable, and only then hand it to agents.

Zhenni was blunt about what happens when teams start at the far end, where the demos are and where leadership is pointing. “Handing organizational context to an autonomous agent where you have not verified that context yourself is a fast way to lose trust in the whole program.”

An agent enriching your semantic layer is reading your source descriptions. If those are restatements, it is reasoning over restatements, and what it produces one level up will be a more confident restatement. Errors do not sit still. They become the input to the next layer, and the next one after that. The compounding has a direction, and nobody gets to pick it.


Consider how much AI-created context can be reviewed

The third idea for better data enrichment is the one teams discover last: acknowledging how many drafted descriptions a human can really get through. This is a critical part of the third phase of agentic data catalogs, learning from failures and compounding feedback.

Once enrichment agents work, running them across the entire data estate costs just one click, but Zhenni held back on this obvious step. Generating hundreds of thousands of descriptions overnight builds a review queue no team will ever clear, and the first time a steward opens it and finds something wrong, doubt spreads across the whole batch. Instead, Mastercard took a bounded set of critical assets, small enough to review properly and visible enough that finishing it would register with the business. This two-week pilot covered approximately 30,000 assets across roughly 700 tables and views, which they estimated at about 6,200 hours of steward effort saved.

A large grid of AI-generated descriptions narrowing into a much smaller bounded review batch of critical assets, which passes through subject matter expert review and splits into three outcomes: accept, edit, or rewrite.

Accept and edit mean the draft was worth having. Rewrite means the time was never saved. | Source: Context and Chaos

Each team then had its own way of deciding whether a layer was solid enough to build on. Mastercard asked its subject matter experts two questions: what percentage could you accept with no changes, and for the rest, were you editing or rewriting from scratch? Editing means the draft was worth having, rewriting means the time was never saved. On the critical assets they sampled, roughly half came back untouched. The output was strongest on tables people curated often, where there was real lineage and real usage to reason from. It was weakest on assets nobody touches, and weak on terms that mean different things to different teams. The layer performs where the evidence beneath it is dense.

Similarly, Sophos had 43,000 descriptions and no way to put them in front of the business one at a time, so the team built a test harness in Cortex Code that graded every one of them from A to E. Around 80% came back A or B, which is what let Keith hand over a pre-screened set of data descriptions rather than an open queue.

As Rupal from ASOS put it, “Perfection is not how your business runs. Your business runs on, can we get 80% there, so we operate the same way.”


Building these ideas into the context development lifecycle

To ensure these three principles stayed true over time, Sophos wrote them into their deployment lifecycle. The team treated them as durable requirements and codified them, from data’s dev through QA into production. Descriptions and terms now have to be defined as part of a data product’s rollout, or the product does not move forward.

A data product moving through dev, QA and production with context travelling alongside it. A context check sits at each gate. Dev requires a description defined and terms attached, QA requires lineage verified and definitions reviewed, and production is reached only once context is certified.

Context stops being applied after the build and becomes a condition of it. | Source: Context and Chaos

Think of this the same as making tests a merge condition. Nobody argues about whether to write them because the pipeline will not let you skip it.

This means that, unlike before, bad context can no longer fail quietly. In the past, a person would spot a bad definition, work around it, and never file a ticket. But with clear requirements and expectations, new agentic data definitions will often fail in public, fast enough that someone traces it back. At ASOS, this has led to business teams asking for metric ownership to be codified and settled, just like any other development process.


Where everyone is still stuck

ASOS, Mastercard, and Sophos have all come a long way in using AI to document their company’s data, but there are still plenty of issues to tackle. Their three principles are valuable for enriching each data asset, but what about looking beyond that?

As Zhenni from Mastercard put it, “Enriching the individual asset might not be enough. The real context also comes with the relationships, whether you call it lineage or ontology, how the assets are meant to be used together, which ones to join and how, what the agreed metrics are, what guidance or prompts or instructions apply to them.”

As we get better at agentic enrichment, the data asset will stop being the unit. Most of what replaces it is not sitting in a structured field anywhere. It is in documents and threads and people’s heads, which is where all of this started.

This isn’t just relevant to enrichment, but also consumption. Sophos is building routing so that when someone asks an assistant a question about the business, it goes to the governed catalog first and gets the real definition before the model starts reasoning on its own.

Both of those are the same problem one level up, and neither has an answer yet.

It’s undeniable that agents have changed the economics of producing context. However, they did not change what that context depends on, and they did not remove the judgment that makes it worth trusting. Describing an estate got dramatically cheaper. Understanding one did not, at anything like the same rate, and the gap between those two is where the work now sits.


The Cats of Context & Chaos

Two-panel cartoon titled certify a sample, ship, repeat. In the first panel an orange tabby in a hoodie sits at a desk with a red pen beside a towering stack labelled 100,000 AI-generated descriptions and says I will review all 100,000 myself. In the second panel the same cat is slumped at the same desk under cobwebs with torn January, February and March calendar pages on the floor and the stack barely touched, while a grey professor cat in tweed and glasses leans through the doorway and says we shipped the first 500 in March.

We shipped the first 500 in March. | Source: Context and Chaos




About Context & Chaos

Context & Chaos isn’t just a newsletter. It’s shared community space where practitioners, builders, and thinkers come together to share stories, lessons, and ideas about what truly matters in the world of data and AI: context engineering, governance, architecture, discovery, and the human side of doing meaningful work.

Our goal is simple, to create a space that cuts through the noise and celebrates the people behind the amazing things that are happening in the data & AI domain.

Whether you’re solving messy problems, experimenting with AI, or figuring out how to make data more human, Context & Chaos is your place to learn, reflect, and connect.


Got something on your mind? We’d love to hear from you.

Share this article