A context layer is the infrastructure that sits between your data systems and the AI that queries them. It supplies the business meaning, identity, provenance, and permissions a model cannot infer from raw tables: what each metric means, which records refer to the same real-world entity, where a number came from, and what an agent is allowed to do with it.
The edges of that definition are still contested. WTF is the Context Layer? is where the people building these layers argue them out.
Without one, an agent reads your warehouse the way a new hire reads a spreadsheet on their first morning. Every column is there. Which table is authoritative, what “active customer” means in your company, and what happens if someone acts on the number are not.
| What it is | The governed layer of meaning, identity, provenance, and policy that sits between enterprise data and AI |
| What it sits between | Your data systems and every AI agent that queries them |
| What it holds | Metadata, semantics, taxonomy, ontology, knowledge graph, and decision context |
| Who needs it | Teams putting agents in front of business data that more than one system describes |
| What it is not | A data catalog, a semantic layer alone, a vector store, or a one-time project |
| Who builds it | Data and AI platform teams, with ownership federated by domain |
| Open standards in play | Model Context Protocol, Open Semantic Interchange, Apache Iceberg |
The layer supplies what a model cannot infer from raw tables: meaning, identity, provenance, and permissions.
Source: Atlan
What a context layer is
A context layer is the system that turns what your organization knows into something a machine can use correctly. It is a layer in the architectural sense: a durable, governed tier that other things are built on, maintained once and read by every agent, rather than assembled by hand inside each application.
The content of that layer is three kinds of knowledge that live almost entirely outside your databases.
Knowledge is what your organization has established as true. Which revenue figure is the reported one. Which customer table survived the last migration. Which of the four dashboards named “churn” the board actually looks at.
Expertise is how work really gets done, the reasoning an experienced analyst applies without being asked. They know that the number looks wrong in February because of how the fiscal calendar is set, and that you exclude internal test accounts before reporting anything externally.
Norms are the rules. Who may see which rows, what an agent is permitted to do on its own, and where a human has to sign off.
Sanjeev Mohan, who spent years as a Gartner research VP and now runs the advisory firm SanjMo, draws the boundary with a purchase he made recently. Semantics can tell you the transaction happened. It cannot tell you why he chose the MacBook Air M5 over the Pro, which sites he read first, or what tipped the decision.
“All the ‘why’ questions that help complete a transaction have got nothing to do with semantics. That’s context.”
Source: Why is Context so Important? with Sanjeev Mohan
That distinction is the reason the layer exists. Semantics standardizes what happened. Context carries why it happened, and what may be done about it now.
Defining the Context Layer, with Prukalpa Sankar
Prukalpa Sankar defines the layer and explains why a capable model with no context is the more dangerous one.
Source: WTF is the Context Layer? Ep. 01
Why context needs its own layer
For thirty years, software was built for people, and people quietly supplied the missing context themselves. A support rep reading a customer record assembles a dozen scattered signals without noticing they are doing it: the account is in renewal, this contact is unhappy, the discount was approved by someone who has since left. None of that is in the record.
An agent cannot do that. Everything a person would have inferred has to be assembled and handed over explicitly. That is the whole job of the layer, and it is new work, because nobody ever had to write it down before.
The obvious hope was that better models would close the gap on their own. They have not, and the reason is structural. Gartner expects 60% of AI projects to be abandoned through 2026 for want of AI-ready data, ahead of any model choice. Atlan co-founder Prukalpa Sankar frames agent performance as intelligence multiplied by context. Intelligence is the variable the market has already solved: it improves every few months, it gets cheaper, and it arrives at your competitors on exactly the schedule it arrives at you. Context is the variable only you can supply.
Her analogy is people. Cognitive ability explains roughly 10% of the variance in job performance, and nobody names their best colleague as the one who scored highest on a standardized test.
“You ask any human this, and it’s a very obvious answer as to what it takes to create the best teammate. That’s what we believe is the missing layer.”
There is a sharper version of this. A more capable model with no context is more dangerous than a weaker one, because it produces confident, well-argued, wrong answers faster.
“I actually think the worst and the most dangerous is high intelligence and low context systems.”
Source: Defining the Context Layer with Prukalpa Sankar
One support agent, tuned to maximize customer satisfaction scores, worked out on its own that issuing refunds raised them. Nothing in the model was broken. It optimized exactly what it was told to optimize. Nobody had encoded what it was allowed to do.
What a context layer contains
Mohan lays the pieces out as six layers. Each answers a question the one before it cannot, and the order matters more than the count.
| Layer | The question it answers | What breaks without it |
|---|---|---|
| Metadata | What exists, who owns it, where it came from | The agent cannot tell an authoritative table from an abandoned experiment |
| Semantics | What each metric means, and how it is computed | Finance and sales bring two revenue numbers to the same meeting |
| Taxonomy | What the agreed terms are | “Other” becomes a permanent category nobody can define |
| Ontology | What the things are, and how they may relate | Cross-system questions join on the wrong key, silently |
| Knowledge graph | The instances, connected | Relationships live on a whiteboard and nowhere a machine can read |
| Context | Why a decision was made, and what is allowed now | The agent is technically correct and practically wrong |
Metadata and semantics are enough to start. The rest add rigor as complexity grows.
Source: Atlan
The trap is reading that list as a gate you have to clear before starting. Mohan’s advice is the opposite, and it was the most repeated guidance across six recorded interviews with the people building these layers.
“Pick a business use case. Collect your metadata, create the semantic layer, and then do context engineering on top. Ontology, knowledge graph, taxonomy, those are important, but if they slow you down, get started without them.”
Source: Why is Context so Important? with Sanjeev Mohan
His reasoning is commercial. The more elaborate the framework looks, the higher the odds a team decides it lacks the budget or the specialist skills, and nothing ships at all.
Who needs a context layer
Not every team needs one yet. The requirement appears at a specific point, and it is worth being honest about where that point is.
You need one when any of these are true. An agent answers business questions using company data, and someone acts on the answer. The same real-world entity, a customer or a product or an employee, lives in more than one system under more than one identifier. You are running more than one agent platform and need them to agree. A non-analyst is asking questions in plain language against governed data. Or the data carries obligations, so who saw what has to be provable afterwards.
You probably do not need one yet if a single team is asking questions of a single well-documented dataset, a human reads every answer before anything happens, and nothing is regulated. Add the layer when the second system or the second team arrives, which is usually sooner than planned.
Who feels the absence first is fairly predictable.
| Role | How the gap shows up |
|---|---|
| AI and platform leaders | Pilots demo well and stall before production, and the blocker is never the model |
| Data engineers | Every new agent means hand-building the same context again in a new place |
| Analytics leaders | Two teams get two numbers for the same metric, from the same warehouse |
| Governance leads | No answer to what the agent saw, what it was permitted to do, or why it decided |
| Domain experts | They become the bottleneck, answering the same definitional question by hand |
Four signals. One is enough to justify building the layer.
Source: Atlan
What you get out of a context layer
The returns are narrower and more measurable than the category talk suggests.
Start with accuracy on the questions people actually ask. In Atlan’s AI Labs benchmark, adding governed context improved text-to-SQL accuracy by 38%. The model did not change. The inputs did.
Adoption is the return teams underestimate. One enterprise launched roughly a thousand conversational analytics workspaces and abandoned about ninety percent of them within a month. The technology worked. Users could not get past seventy to eighty percent accuracy on ordinary questions, so they stopped trusting it and went back to asking a person. Trust is a threshold rather than a gradient, and below it usage goes to zero.
Then there is cost per answer. Jessica Talisman’s point is that context encoded once is reused, so agents stop re-deriving your meaning on every single run. Most teams discover this backwards, when the invoice shows how many tool calls sat behind one basic question.
Teams running several agent tools find them returning contradictory answers about the same data, each speaking a slightly different dialect of the same business. A shared layer is what makes them agree.
And the audit trail finally exists. Decision context makes it possible to answer what the agent saw and which rule applied, after the fact, without reconstructing it from logs.
The spending pattern follows the same logic. Gartner research finds high-performing organizations invest nearly twice as much in data foundations as in AI tooling.
Context layer vs semantic layer vs RAG vs data catalog
This is where most of the confusion sits, and the honest answer is that these stack rather than compete.
| What it does | What it does not do | |
|---|---|---|
| Data catalog | Inventories what exists, who owns it, and where it came from | Does not define business meaning or govern agent behavior |
| Semantic layer | Standardizes metric definitions and computes them deterministically | Does not carry procedural knowledge, policy, or history |
| RAG | Retrieves relevant passages from a corpus at query time | Does not persist state, resolve conflicts, or decide permissions |
| Context layer | Governs meaning, identity, policy, and decision history across all of it | Does not replace any of the three; it sits above them |
Each answers a different question. They stack rather than compete.
Source: Atlan
David Mariani co-founded AtScale and has spent thirteen years building independent semantic layers. Asked whether a semantic layer is enough on its own, his answer is no, and he is unusually precise about where his own layer stops.
“A semantic layer’s job is to deliver deterministic query results. It’s not its job to reason and plan on which queries to run. That’s what the agent or the LLM does.”
That division of labor is the practical design principle underneath all of this.
“You are giving it a curated set of metrics with which AI can be creative and probabilistic, which is what you want it to be. But not probabilistic when you’re calculating revenue or gross margin.”
Source: Is Semantic Layer = Context Layer? with David Mariani
Is a Semantic Layer the Same as a Context Layer? with David Mariani, AtScale
David Mariani draws the line between what a semantic layer computes and what the agent decides.
Source: WTF is the Context Layer? Ep. 03
The accuracy numbers make the case better than the architecture diagrams do. One major platform put its own ontology accuracy at 84.5% on stage. Being wrong more than fifteen times in a hundred is not a rounding error when someone reports the answer upward. Worse, single-run accuracy hides a second failure: a model can answer correctly on Monday and differently on Thursday, which is why Mariani’s team now measures determinism over time rather than accuracy on one pass.
How to build a context layer
Across those six sessions, with people who disagreed about nearly every architectural question, the starting advice barely varied.
- Begin from a business problem, not a technology. Pick one question the business already argues about, solve it end to end, then work backward to the concepts and vocabulary it required. This is the step teams skip, and skipping it is why context projects become permanent.
- Mine the context you already have. Metric definitions, query history, lineage, and BI semantics already encode most of your meaning. Mohan’s framing: “You’re not embarking on a different project. You’re upskilling.”
- Write the questions before the model. Talisman calls these competency questions, the specific things the layer must be able to answer. They scope the work, and they become the evaluations you test against later. “We do not build ontologies until we form competency questions.” (source)
- Model people, not roles. Define person as the class and employment as a property. Teams who model employee as the class erase people the moment they leave, along with every decision those people made.
- Keep humans on the certification step. Models propose definitions and joins at a scale people cannot match. Deciding which proposal is correct stays a human job, and that is the step that earns trust.
- Insist on portability. Ask whether your context is open, whether it is interoperable, and what happens to it if you walk away from the vendor.
A realistic first ninety days: one use case, the metadata and semantics behind it, roughly a dozen competency questions, and a human review step. Ontology and graph come later, when the second and third use cases show you what actually repeats.
Prove one useful context loop before building ontology and graph.
Source: Atlan
Deeper on the build: how to implement an enterprise context layer, building context for LLMs, how much context is enough, and a context engineering framework to structure the work.
What the people building it still disagree about
The category is young enough that the practitioners building it disagree in public about three things. Knowing where the disagreements are is more useful than pretending there is a consensus.
Does it need a graph database? Emil Eifrem coined the term graph database two decades ago and runs Neo4j, and even he does not make the strong claim. Because the popular database models are isomorphic, a knowledge graph can live in Postgres, Mongo, or DynamoDB without losing information. Where he holds firm is latency: sub-millisecond fraud detection should not run virtually, while most other queries survive a couple of seconds. Sankar granted the premise and pressed on the conclusion: “I’m convinced AI needs relationships. Graph theory, AI needs 100%. Do we need to manifest all of this in a graph database?” Model the relationships. Move the data only where latency forces you to.
Does the ontology come first? Talisman argues most teams enter at the wrong end, reaching for an ontology while their own terms are still undefined. “An ontology doesn’t magically fix things.” (source) Mohan argues the opposite, that waiting for one is how projects die. Both are describing real failure modes. Skipping the definitional work hardens your mess; waiting for it means nothing ships.
Do AI Agents Need a Graph Database? Emil Eifrem and Prukalpa Sankar
Emil Eifrem and Prukalpa Sankar disagree on whether relationships have to live in a graph database.
Source: WTF is the Context Layer? Ep. 05
Is any of this settled vocabulary? Not yet, and the person who named graph databases is the bluntest about it. “There used to be cloud washing back in the days. I feel like we have some version of ontology washing going on.” (source)
Who owns the context layer
The pattern among organizations furthest along is a central platform with ownership federated by domain, the same shape data platforms settled into a decade ago. In practice, whoever owns the AI architecture inherits the layer.
Day to day that means narrow authority over shared artifacts and wide authority over local ones. One owner holds the core definitions every downstream agent inherits. A domain team keeps a scoped version for its own use cases, connected back to the root, so a local decision never silently rewrites a global one.
Centralized standards at the root, scoped context owned by each domain.
Source: Atlan
The ownership question has its own detail: who owns the context layer, whether it sits with data or AI teams, and how centralized and federated teams split the work.
Keep going
The definitions on this page come from WTF is the Context Layer?, a six-part series where the people building this layer argue it out on the record: an analyst who watched the category form at Gartner, the CTO of an independent semantic layer, an ontologist with twenty-five years in the discipline, the person who coined the term graph database, and the investor who named the context graph. Every quote above links back to the episode it came from.
If you would rather start from the parts than the argument, the context layer glossary defines the vocabulary, and why AI agents need an enterprise context layer makes the case in one place.
What Does Atlan Do? The Context Layer for Enterprise AI
A short walkthrough of how metadata from across a data stack becomes context an agent can actually use.
Source: Atlan on YouTube
---
FAQs about context layers
1. What is a context layer in simple terms?
It is the infrastructure between your data systems and your AI that supplies the business meaning a model cannot infer from raw tables: what each metric means, which records describe the same entity, where a number came from, and what an agent is permitted to do with it.
2. What is the difference between a context layer and a semantic layer?
A semantic layer standardizes and computes metric definitions inside an analytics domain, which keeps numbers deterministic. A context layer includes that and adds cross-system identity, procedural knowledge, governance, and decision history. The clearest test is that a semantic layer delivers deterministic results and does not decide which query to run.
3. Who needs a context layer?
Teams putting AI agents in front of business data that more than one system describes. The trigger is usually the second agent platform, the second team asking the same question differently, or the first regulated use case where you have to prove what the agent saw.
4. What are the components of a context layer?
Six: metadata, semantics, taxonomy, ontology, knowledge graph, and decision context. Metadata and semantics are enough to start a first use case. The rest add rigor as complexity grows.
5. Do I need a graph database to build a context layer?
No. Because the common database models are isomorphic, a knowledge graph can sit in Postgres, Mongo, or DynamoDB without losing information. A graph database earns its place when traversal latency is the binding constraint, as in sub-millisecond fraud detection.
6. Do I need an ontology before building a context layer?
Not to start. Do the vocabulary work on a narrow slice first, since an ontology built on undefined terms makes the ambiguity permanent. Write the competency questions the layer must answer before you build it.
7. How is a context layer different from RAG?
RAG retrieves passages from a corpus at query time. A context layer governs the structured meaning, identity, and policy that exist before retrieval happens. The two are complementary: a governed layer gives retrieval something trustworthy to draw from.
8. How long does it take to build a context layer?
A first working use case takes weeks, not quarters, if you scope it to one business question and the metadata and semantics behind it. Treating all six components as a prerequisite is the most common reason nothing ships at all.
Sources
- WTF is the Context Layer? EP 01, with Prukalpa Sankar
- EP 02, Why is Context so Important? with Sanjeev Mohan, SanjMo
- EP 03, Is Semantic Layer = Context Layer? with David Mariani, AtScale
- EP 04, From Ontology to Context Layer, with Jessica Talisman
- EP 05, How Do Graph Databases and the Context Layer Fit Together? with Emil Eifrem, Neo4j
- EP 06, Can a Context Graph Alone Make AI Reliable? with Jaya Gupta, Foundation Capital
- SKOS Simple Knowledge Organization System Reference, W3C
- Lack of AI-Ready Data Puts AI Projects at Risk, Gartner
- Gartner Top Data and Analytics Predictions