A semantic layer and a data catalog solve different halves of the enterprise data problem, and the line between them just moved. dbt Labs open sourced MetricFlow, the engine behind the dbt Semantic Layer, under Apache 2.0 on 14 October 2025, arguing that metric math should be deterministic, not guessed fresh by an LLM on every prompt.[1] Cube made the same move from a different angle: its semantic layer now exposes certified metrics through its own MCP server, alongside the APIs it already had.[2] Both categories speak the same protocol now. What still separates them is what each hands back: a semantic layer returns a metric’s logic, a data catalog for AI returns whether the agent is cleared to see the table behind it.
Neither answer makes the other optional. An agent requesting “Monthly Recurring Revenue” from a governed semantic model still needs to know whether it can query the subscriptions table underneath, and a catalog with airtight lineage still can’t say whether MRR excludes trials. Atlan’s Enterprise Data Graph connects both answers into one governed response, instead of leaving an agent to reconcile two systems alone.
| Dimension | Semantic Layer | Data Catalog |
|---|---|---|
| What it is | A modeling layer compiling business metrics into SQL (dbt, Cube, LookML) | A metadata system indexing and governing data assets |
| What it answers | “How is Revenue calculated?” | “Where does Revenue live, and can this agent see it?” |
| Delivery to AI agents | MCP server, REST, and GraphQL exposing certified metrics | MCP server exposing governed metadata and access policy |
| Governs access control | No | Yes, integrated with IAM and data policy |
| Governs metric logic | Yes, versioned and reusable | No, only descriptive mapping |
| Failure mode without it | Every tool calculates Revenue differently | An agent queries data it was never cleared to see |
| Best for | Standardizing metric math across BI and AI | Standardizing what’s discoverable, owned, and safe to query |
Semantic layer vs data catalog: what’s the difference?
The distinction holds even as the tooling changes: a semantic layer is executable logic, a data catalog is governance metadata. Cube frames its own category the same way: the layer standing between raw tables and every consumer of a number, dashboard or agent, that needs it calculated the same way twice.[3] Atlan plays the equivalent role on the governance side, connecting catalog metadata to the same certified metrics an agent retrieves from dbt or Cube.
Confusion persists because both now expose an MCP endpoint, and MCP is agnostic to whatever sits behind it. An agent calling a semantic layer’s MCP server gets a metric definition; one calling a catalog’s MCP server gets a permission decision and a lineage trail. Mistaking one for the other is how AI agent hallucination about business numbers actually shows up in production, not a model guessing at arithmetic, but an agent never learning that “Revenue” already meant something specific.
What is a semantic layer?
A semantic layer sits between the warehouse and every tool that needs a business metric, compiling dimensions, filters, and joins into SQL at query time rather than hand-written per consumer. What is a semantic layer covers the concept in full; the short version is that it turns “Revenue” from a table column into one definition every consumer shares.
dbt Labs’ argument for open sourcing MetricFlow: an agent requesting a metric by name and getting back proven SQL is right more often than one guessing the calculation from scratch.[4] dbt Labs’ own replication study backs this up: AI answered 83% of addressable natural-language questions correctly through the dbt Semantic Layer, with several answered at 100% accuracy.[1] Cube makes the same case from a different angle, over 400 companies run its semantic layer, and its MCP server lets an agent list certified measures before writing a query.[2] LookML established the same pattern years earlier, inside one BI tool.[5]
Two agents asking the same layer for “Q1 Revenue” get the same SQL and the same number, a guarantee prompt engineering alone can’t produce. Text-to-SQL vs semantic layer covers why ad hoc SQL reintroduces the inconsistency a semantic layer exists to remove; vendors converge on the same shape from different starting points, Unity Catalog’s semantic layer, Honeydew, and Lightdash each trade governance for modeling flexibility differently, a full rundown lives in best semantic layer tools.
Core components of a semantic layer
- Metric definitions: versioned SQL logic for a business term like Revenue or MRR, shared across every consumer instead of rewritten per tool
- Dimension and join logic: how a metric slices by region, product, or time without every consumer re-deriving the same joins
- API surface: REST, GraphQL, and SQL, plus increasingly an MCP server built for agent consumption rather than dashboards alone
- A governed metric catalog: the dbt Semantic Layer, Cube, LookML, and Atlan’s Semantic View Generator each maintain one, though only Atlan’s ties directly into catalog governance
What is a data catalog?
A data catalog indexes what data exists, tables, dashboards, pipelines, models, and attaches what an agent needs before touching it: ownership, provenance, freshness, permission to query. Its job here is narrow: decide access, not calculate metrics.
Gartner’s metadata management research names the same shift underway: from searchable index toward the system deciding what’s safe to hand an agent at query time.[6] Five governance dimensions fall to the catalog, not the semantic layer:
- Lineage: end-to-end tracking from source system through transformation to the response that used it
- Ownership: which team is accountable when a number looks wrong
- Quality: freshness and completeness scores an agent can check before trusting a table
- Classification: PII tags and sensitivity labels deciding what an agent may expose
- Access control: IAM-integrated policy deciding which agent, not just which human, can query which asset
The glossary overlaps with the semantic layer, partially: an entry says “Revenue lives in the transactions table, owned by Finance,” a label, not a calculation. Context layer vs semantic layer draws this line further, data catalog vs context layer covers what a catalog alone still misses, and AI agents for data catalog treats the glossary as connective tissue between definition and retrieval.
Core components of a data catalog
- Business glossary: a descriptive mapping from a term to the asset it lives in, not an executable definition
- Lineage graph: source-to-destination tracking across every transformation a number passes through
- Access policy: governance rules an MCP server enforces before an agent ever sees a row
- Quality and classification scores: freshness, completeness, and sensitivity metadata attached to every asset
Inside Atlan AI Labs: The 5x Accuracy Factor
The research behind why governed context, not model size, is the biggest lever on AI agent accuracy.
Get the Accuracy EbookSemantic layer vs data catalog: head-to-head
The sharpest differences show up in who owns the failure when an agent gets something wrong. A metric definition problem belongs to whoever owns the semantic layer. An access or freshness problem belongs to whoever owns the catalog.
| Dimension | Semantic Layer | Data Catalog |
|---|---|---|
| Primary focus | Metric logic and calculation consistency | Asset governance and access control |
| Key stakeholder | Analytics engineering, BI teams | Data governance, platform engineering |
| Measurement approach | Query consistency across every consumer | Governed-asset coverage, policy compliance |
| Implementation scope | Model metrics on top of the warehouse | Connect and classify assets across the estate |
| Time to value | Fast for one team’s metric set | Longer setup; value compounds as more systems connect |
| Tooling requirements | dbt, Cube, or LookML modeling skills | Catalog administration, glossary owners, IAM integration |
| Failure mode | Correct access, wrong number | Right number, an agent that never should have seen it |
| Delivery to AI agents | MCP server exposing certified metrics | MCP server exposing metadata and policy |
Example: a shared vocabulary that only holds if both travel together. Workday’s data team describes building toward a shared language that AI can draw on directly, not a spreadsheet of definitions one analyst keeps current by hand. That shared language only holds if the calculation and the access decision move together. A semantic layer alone, or a catalog alone, leaves exactly that gap open, and why AI agents fail in production traces more of these gaps back to the same root cause.
Do you need both a semantic layer and a data catalog?
Most enterprise AI programs need both, not one instead of the other. The semantic layer supplies the calculation; the catalog supplies the permission and the context around it.
How they work together
When an agent needs to answer “What was Q1 Revenue?”, the two systems resolve different halves of the question in sequence, not in competition:
- Catalog step: confirm the Revenue table exists, verify the agent can access it, and check that Q1 data is actually complete.
- Semantic layer step: retrieve the certified Revenue definition, the exact SQL, filters, and grain the organization already agreed on.
- Execution step: run the validated query against the governed, verified asset and return one answer, not two conflicting ones from two tools.
Skip the catalog step and an agent may query restricted or stale data with full confidence. Skip the semantic layer step and it may calculate Revenue on its own, quietly wrong. Atlan’s MCP server carries both halves at once, the same pattern behind MCP-connected data catalog work more broadly, so an agent queries one endpoint instead of reconciling the two itself.
When should you add a data catalog on top of your semantic layer?
The right answer depends on how many teams and agents share the same metrics, and how much is riding on getting access right.
Stay with a semantic layer alone when one team owns the metric set, informal access conventions are enough, and no agent outside that team queries the underlying tables directly.
Add a catalog when more than one agent needs the same metrics, access has to be policy-scoped rather than credential-scoped, or PII sits close enough to the data that a classification mistake would matter.
Invest in both from day one when building a multi-agent program or operating under an audit requirement that already assumes governance exists. Why AI agents need an enterprise context layer makes the case that this is cheaper to plan for upfront than to retrofit later, a conclusion AI agent governance reaches from the compliance side of the same problem.
Check your context readiness
Run a quick assessment on how governed your current metrics and metadata estate is before scaling agent workloads on top of it.
Take the AssessmentHow Atlan approaches semantic layers and data catalogs
Teams that wire agents straight to a semantic layer, or straight to a catalog, tend to end up calculating consistently but retrieving carelessly, or the reverse. Either failure mode is avoidable once both share one governed source of truth, the same argument context layer vs data catalog vs semantic layer makes across all three layers at once.
Atlan ingests metric definitions from dbt, Cube, and LookML rather than re-deriving them, which means the semantic layer stays the system of record for the calculation while Atlan binds lineage, ownership, and access policy to each metric on top. Atlan’s Semantic View Generator can also build governed metric views directly from catalog-registered assets, served to agents through the same MCP server that already handles catalog governance. The result functions as a Context Product: one versioned bundle carrying both the governance metadata an agent needs to validate access and the semantic definition it needs to calculate correctly, queryable from a single endpoint instead of two.
In Atlan’s AI Labs benchmark, adding this combined context improved AI’s text-to-SQL accuracy by 38%, a first-party result independent of which semantic layer or warehouse a team runs underneath. This is what Workday’s data team means by a shared language AI can draw on through Atlan’s MCP server, and it’s the same integration pattern behind data catalog as LLM knowledge base: governance and calculation logic served together, not reconstructed by the agent at query time.
Real stories from real customers: Semantic context at enterprise scale
"Atlan captures Workday's shared language to be leveraged by AI via its MCP server. As part of Atlan's AI labs, we're co-building the semantic layer that AI needs."
— Joe DosSantos, VP Enterprise Data & Analytics, Workday
"AI initiatives require more context than ever. Atlan's metadata lakehouse is configurable, intuitive, and able to scale to hundreds of millions of assets. As we're doing this, we're making life easier for data scientists and speeding up innovation."
— Andrew Reiskind, Chief Data Officer, Mastercard
Neither quote name-checks a specific semantic layer tool. Workday and Mastercard both arrived at the same conclusion this page argues for: the calculation engine underneath was never the whole problem, a shared vocabulary delivered through governed access was the other half.
See if your agents already know too much, or too little
Run through the checklist enterprise teams use to find gaps before scaling AI agents on top of any semantic layer or catalog.
Check Your ReadinessThe calculation question a data catalog can’t answer on its own
A data catalog answers the governance questions an agent must resolve before it queries anything: does this asset exist, is it accessible, how fresh is it, who’s accountable for it. It cannot tell the agent what Revenue means, whether refunds count against it, or which filters apply, that’s the semantic layer’s job, and no lineage or access control substitutes for it. Enterprises that answer only one half end up with agents that retrieve responsibly and calculate inconsistently, or the reverse, exactly the gap semantic layer for BI vs AI agents covers from the modeling side and why AI agents need an enterprise context layer covers from the governance side.
Atlan’s Context Layer and Context Graph architecture connects both halves. By ingesting dbt, Cube, and LookML semantic definitions and binding catalog governance metadata to them, Atlan creates Context Products, versioned, governed bundles that serve a metric’s logic and its access decision through one MCP endpoint, so an agent never has to stitch catalog and semantic layer together itself. The semantic layer for AI agents is the discipline behind building those definitions explicitly for agent consumption, not just for BI tools.
FAQs about semantic layers and data catalogs
1. What is the difference between a semantic layer and a data catalog?
A semantic layer defines business metrics as reusable SQL logic, turning raw schemas into consistent definitions like Revenue or Monthly Active Users. A data catalog tracks assets: lineage, ownership, quality, access. The catalog decides which assets an agent may query; the semantic layer decides how a metric gets calculated from them.
2. Does a data catalog replace a semantic layer?
No. A catalog manages metadata governance: ownership, provenance, access. A semantic layer manages metric consistency: what Revenue means in SQL terms. AI agents need both.
3. Do dbt, Cube, and data catalogs all use MCP now?
Largely, yes. dbt open sourced MetricFlow in October 2025 and serves it through an MCP-compatible API; Cube exposes its metric layer the same way. A governed catalog does this for access and lineage. All three speak the protocol; each hands back a different answer.
4. What is the difference between a semantic layer and a data mesh?
A data mesh is an organizational architecture: domains own their data products. A semantic layer is a technical layer standardizing metric definitions across tools. They’re complementary; a catalog governs both.
5. Can Atlan act as a semantic layer?
Atlan integrates with dbt, Cube, and LookML, ingesting their metric definitions and binding governance, lineage, ownership, and access policy to them. Its Semantic View Generator can also build governed metric views served to agents through MCP.
6. What tools make up a semantic layer?
Common tools include the dbt Semantic Layer, built on the open-source MetricFlow engine, Cube, LookML, and Atlan’s Semantic View Generator. Each defines metrics as versioned SQL logic multiple consumers can query consistently.
Sources
- Announcing Open Source MetricFlow, dbt Labs
- Semantic Layer for AI Agents (2026), Cube
- What Is a Semantic Layer?, Cube
- About MetricFlow, dbt Developer Hub
- LookML Terms and Concepts, Google Cloud
- Gartner Magic Quadrant for Metadata Management Solutions, Gartner
- dbt Semantic Layer FAQs, dbt Developer Hub