Google Cloud Knowledge Catalog is authoritative for what Google Cloud can observe and enforce on your BigQuery estate. An enterprise context layer, the category Atlan builds, is authoritative for what your business agreed. Google renamed the product on April 10, 2026, and according to a16z (2025), 37% of enterprises now run five or more models in production.
Four things changed in 2026. Dataplex Universal Catalog became Knowledge Catalog on April 10, legacy Data Catalog shut down on June 1, federated support for third-party Iceberg REST catalogs arrived in Preview, and Google shipped a remote MCP server that non-Google agents can call. Together they retire the old question, whether your estate is single-cloud or multi-cloud. What decides now is whose agents are asking, and which layer answers when both systems hold the same definition.
| Dimension | Google Cloud Knowledge Catalog | Enterprise context layer |
|---|---|---|
| What it is | Google’s Gemini-powered context engine inside Google Cloud | A neutral layer holding the definitions the business agreed |
| Where it is authoritative | What Google Cloud can observe and enforce | What was agreed, approved, and contracted |
| What it observes automatically | BigQuery, Looker, Spanner, Cloud Storage, Vertex AI | Systems inside and outside any single cloud |
| Who it enforces for | Principals and policies inside Google Cloud IAM | Named owners and approvers across the estate |
| Agent interface | Remote MCP server (Preview) | Hosted MCP server (GA), API, SQL |
| Cost model | DCU-hour metering, Standard and Premium tiers | Platform subscription |
| Best fit | Context Google Cloud produces about its own assets | One definition resolving identically per agent vendor |
What is the difference between Google Cloud Knowledge Catalog and an enterprise context layer?
Permalink to “What is the difference between Google Cloud Knowledge Catalog and an enterprise context layer?”What separates the two is authority, not the amount of ground each covers: an agent has to know which system’s answer counts.
The two are easy to conflate: both hold glossary terms, both hold data products, both ship an MCP server, and both use the word context. Several systems of semantics running at once each behave as the definitive one, which is why the distinction between a data catalog and a context layer decides what an agent receives. What the business agreed is the harder half: an agreement reached in a ticket or a review meeting is business context for AI no query engine observes.
The failure mode is quiet. Two agents on two vendors return two numbers for revenue, neither flagged as wrong, and the one that reaches the board deck is the one that got asked first. Only authority would have settled it.
What does Google Cloud Knowledge Catalog do on BigQuery today?
Permalink to “What does Google Cloud Knowledge Catalog do on BigQuery today?”Google Cloud Knowledge Catalog is Google’s context engine for a Google Cloud estate, and on BigQuery it reads your tables, views, and query history continuously rather than on request.
Dataplex Universal Catalog has been called Knowledge Catalog since April 10, 2026, with the API, CLI, and IAM names unchanged and the endpoint still dataplex.googleapis.com, per Google Cloud’s Knowledge Catalog overview. Its deprecation notice records legacy Data Catalog shutting down on June 1, 2026.
Google calls it “an always-on context engine” unifying structured, unstructured, and SaaS data into agent-ready truth, and it ingests automatically from BigQuery, Dataform, Iceberg REST Catalog tables, Vertex AI, Looker, Spanner, and Cloud Storage. Gemini Enterprise, previously Google Agentspace, consumes the result directly.
William Anderson, CTO of Bloomberg Media, in Google’s launch post: “By unifying Bloomberg Media’s enterprise metadata and business context through the Knowledge Catalog, we successfully launched our Data Access AI Agent.”
What Knowledge Catalog does natively on Google Cloud
Permalink to “What Knowledge Catalog does natively on Google Cloud”- Business glossary (GA). Terms, definitions, and term-to-column links.
- Data products (GA). BigQuery tables, views, and storage in a semantic wrapper.
- Quality and profiling scans. Column-level results, policy checks, anomaly detection.
- IAM and policy tags. Enforcement where the data physically sits.
- Gemini-derived descriptions. Generated from metadata Google Cloud already holds.
- Semantic search. A hybrid semantic search implementation at sub-second latency, which also produces the lineage an AI agent needs for Google Cloud jobs.
Which capabilities are generally available and which are in Preview
Permalink to “Which capabilities are generally available and which are in Preview”| Capability | Status | What it covers |
|---|---|---|
| Metadata aggregation | GA | Google Cloud data and AI services |
| Data products | GA, April 2026 | BigQuery assets in a semantic wrapper |
| Semantic search | GA | Hybrid retrieval, sub-second latency |
| Access-control-aware search | GA | Results filtered to the calling principal |
| Context federation | Preview | SAP, Salesforce Data360, Workday, Palantir, ServiceNow |
| BigQuery measures | Preview | Metric definitions against BigQuery tables |
| Deep multimodal extraction | Preview | Semantics from unstructured assets |
| Remote MCP server | Preview, “as is” | Agent access to search, schemas, glossaries |
That status column is the most useful line in any evaluation here, and it is Google’s own published position on which classes of context Google Cloud is authoritative for today. What it leaves open is which layer answers when the same definition also exists somewhere Google Cloud does not run.
What is an enterprise context layer authoritative for?
Permalink to “What is an enterprise context layer authoritative for?”An enterprise context layer is the neutral system holding the definition the business agreed, the approval that happened outside any warehouse, the contract attached to it, and the graph an agent traverses to get the same answer whichever vendor is asking.
Its authority comes from a different source. Google Cloud earns authority by running the engine and observing what happens inside it. The neutral layer earns authority by being where cross-system agreement is recorded and approved, a business event rather than a platform event, and a context layer reference architecture sets out how the two sit together.
Agents are why this now matters operationally. Asked what a metric means, an agent needs one resolution, not a per-system answer that is only locally correct. That comes from traversing an agent context graph rather than reading whichever table description sat nearest the query.
What an enterprise context layer holds
Permalink to “What an enterprise context layer holds”- The reconciled metric definition. The version two or more teams argued their way to.
- Approval and ownership state. Who certified this, when, and who is accountable if it is wrong.
- Cross-system lineage. Provenance spanning systems no cloud provider can observe.
- Data contracts. Data contracts for AI stating shape, freshness, and owner.
- The Enterprise Data Graph. The structure resolving one definition per query, so a Gemini agent and a Claude agent answer the same question the same way.
Authority here is a property of neutrality rather than capability, which makes a context layer complementary to a platform-native context engine rather than parallel to one.
What is the context layer, in plain terms?
The ebook walks through what an enterprise context layer holds, why agents need it as a distinct layer, and how it sits alongside the context your cloud platform already produces.
Get the Context Layer EbookWhere does federation stop and semantic extraction begin?
Permalink to “Where does federation stop and semantic extraction begin?”Federation registers an asset so it can be found and referenced. Semantic extraction produces the descriptions, relationships, and meanings an agent reasons over, and Google documents the difference itself.
Start with what Google shipped, because it is more than most comparisons credit. Google Cloud’s documentation lists federated support for Databricks Unity IRC, AWS Glue Data Catalog IRC, and Snowflake Horizon IRC, all in Preview, plus context federation for SAP, Salesforce Data360, Workday, Palantir, and ServiceNow. “Native means single-cloud” stopped being a true sentence in April 2026.
Then read the caveat Google publishes in the same overview:
“While Knowledge Catalog supports major third-party systems, certain automated semantic extractions might be limited to built-in Google Cloud services.”
That is the decision in one sentence. Our reading of it: a registered asset an agent can see is not yet a described asset an agent can reason over. Federation gets a table into the entry list; extraction produces the column meanings and relationships. According to Atolio (2026), an enterprise-search company with no stake in the catalog market, Knowledge Catalog’s “deepest integrations are with Google Cloud services,” which is a statement about where configuration effort lands rather than about capability.
The practical consequence is a division of labor you can write down.
| Class of context | Authoritative layer | Why | What breaks if you invert it |
|---|---|---|---|
| BigQuery schema and column semantics | Knowledge Catalog | Google Cloud reads the schema as jobs run | Your neutral record is stale the moment a table changes |
| Lineage inside Google Cloud | Knowledge Catalog | Emitted by the services that ran the jobs | Lineage is inferred, not observed |
| IAM and policy enforcement | Knowledge Catalog | Enforcement happens where the data sits | A rule exists on paper and nothing stops the query |
| Gemini-derived descriptions | Knowledge Catalog | Generated from metadata Google Cloud holds | You hand-write what a service produces |
| The reconciled metric definition | Enterprise context layer | Reconciliation happened between teams | Each system keeps its own version of one number |
| Approval and ownership state | Enterprise context layer | Approval is a business event with a named owner | An agent cannot tell an approved definition from a draft |
| Context beyond Google Cloud’s observation | Enterprise context layer | The record must exist where observation stops | Anything past the boundary is absent, not wrong |
| One definition per agent vendor | Enterprise context layer | Neutrality is the property being bought | Each agent inherits whatever its own stack holds |
Every row states what a layer is authoritative for. That is why a multi-cloud context layer argument built on cloud count has aged badly while the same argument built on context portability has not. What a semantic layer for AI agents has to answer is not where the asset is registered, but whose meaning the agent gets when both systems hold one.
Whose agents are asking? Why a BigQuery-only estate is often a multi-vendor agent estate
Permalink to “Whose agents are asking? Why a BigQuery-only estate is often a multi-vendor agent estate”The old axis was where your data lives. The current axis is how many agent vendors need the same definition to mean the same thing, and those two numbers stopped moving together.
According to a16z’s AI Enterprise 2025 study of 100 enterprise CIOs, 37% of enterprises now run five or more models in production, up from 29% the year before. One warehouse, five model vendors, and the model-agnostic context layer is the part of that stack nobody rebuilds per vendor. The choice between Google ADK, LangGraph, and AutoGen gets made per team, not per company.
Grounding those agents in agreed semantics changes their answers measurably, with a qualifier that matters. According to dbt Labs (2026), on modeled data, semantic-layer grounding took Claude Sonnet 4.6 from 90.0% to 98.2% and GPT-5.3-Codex from 84.1% to 100%. Across all questions in the same benchmark the figures are 64.5% for text-to-SQL and 72.7% with the semantic layer, a different measurement rather than the same one restated. Grounding pays off most once the modeling is done, so the layer and the modeling are one investment.
The same benchmark argues the other side: unaided text-to-SQL nearly doubled between 2023 and 2026, from 32.7% to 64.5%. Models are getting better at this alone, so the claim is narrower than the headline. What improves is the model’s guess, not the enterprise’s agreement about what a number means, and context management across multi-agent systems is an authority problem before it is a retrieval problem.
Which layer should serve context to your agents over MCP?
Permalink to “Which layer should serve context to your agents over MCP?”Both layers expose an MCP server, so an agent can reach the context either way. The open question is which layer should answer.
Google’s remote MCP server is real and it serves non-Google agents. Per Google’s documentation, the endpoint is dataplex.googleapis.com/mcp, authentication is OAuth 2.0 plus IAM with the roles/mcp.toolUser and roles/dataplex.catalogAdmin roles, and the named clients are Gemini CLI, ChatGPT, Claude, and custom applications. It exposes entry search, schemas, data products, glossaries, and lineage. Google’s FAQ states that agents “can discover and adaptively use Knowledge Catalog tools through a local or remote MCP server.” Status: Preview, “available as is and might have limited support.”
Agents can therefore reach what Google Cloud extracted, which is the honest 2026 position and a useful one. Why MCP matters for AI agents was never that it grants access to something otherwise unreachable. What a protocol standardizes, as an MCP architecture deep dive makes clear, is the shape of the request, not the authority of the response. Atlan’s hosted MCP server is GA and serves ADK and Gemini CLI agents today, so an estate running both has two servers to route between.
Make the judgement before you wire the second. Asked what a metric means, Google’s server is authoritative for what Google Cloud extracted, the neutral layer for what was agreed. That is a routing decision, and MCP delivers business context only once someone has decided which server owns which class of question.
Are your agents ready for the context you have?
Run the AI Agent Context Readiness Checklist to see which classes of context your agents can actually resolve today, and which are still answered by whichever system replies first.
Run the Readiness CheckWhat happens when both systems define the same business term?
Permalink to “What happens when both systems define the same business term?”Both systems can hold the same glossary term, data product, and quality signal, and without an explicit precedence rule the divergence is silent. Nothing alerts. The two records simply stop agreeing.
| Object both can hold | Held natively in | Precedence rule to set | How drift shows up | What to check |
|---|---|---|---|---|
| A glossary term | Knowledge Catalog glossary (GA); the context layer’s glossary | The layer holding the approved definition answers; the other points to it | Text diverges after a console edit no crawl collected | Definition text and named owner match, same date |
| A data product | Knowledge Catalog data products (GA); the context layer’s contract | Google Cloud for packaged assets, the context layer for contract and approver | Assets added in the console with no contract change | Asset list and contract name one owner and scope |
| A quality signal | Knowledge Catalog quality and profiling scans | Google Cloud produced the measurement, the context layer holds the agreed threshold | A scan passes while the agreed threshold moved | Scan threshold and contract threshold are one number |
| Asset ownership | IAM principals; the accountable owner in the context layer | IAM for access, the context layer for accountability | The principal is a service account with no human owner | Every certified asset resolves to a named person |
| A column-level description | Gemini-derived descriptions; Aspect field values written back | Schema authority with Google Cloud, value authority with the context layer | A generated description overwrites an approved one | Which system last wrote the field, and whether it is in scope |
One documented constraint settles most of the last row, and it is a mechanic rather than a preference. Atlan’s Aspects Reverse-Sync writes Aspect field values only: Aspect Type schemas are read-only through Atlan and deletion is unsupported. Schema authority therefore sits with Google Cloud by design, and value authority can sit with the neutral layer.
State the scope honestly in the same breath. Atlan’s Knowledge Catalog connector is in beta and BigQuery-only for enrichment; Cloud Storage, Spanner, and AlloyDB are roadmap. Anyone promising parity between a beta connector and a GA platform surface is selling something.
Three ways drift shows up
Permalink to “Three ways drift shows up”A term edited in the Google Cloud console does not reach the neutral layer until the next crawl, so for a window both systems are confidently correct and different. A definition approved in the neutral layer reaches the console only through the Aspect field values reverse-sync supports, so an approval can be real and invisible where the platform team is looking. Anything outside BigQuery has no write path today, which is the shape context drift takes when nothing is watching, and why context drift detection belongs in the operating routine rather than the incident review. That window is also where tribal knowledge re-enters: someone knows which record is right, and no agent can ask them.
Set the precedence rule before you connect the second MCP server
Permalink to “Set the precedence rule before you connect the second MCP server”- Do the two glossaries agree on the definition text and the named owner, checked the same day?
- Does the agent’s MCP path resolve to the layer holding the approved definition, not the one that answers fastest?
- Do the data product and the contract name the same owner and scope?
- When they disagree, which record would you defend to an auditor?
- Is there a reconciliation cadence with someone’s name on it, or only an integration?
Coexistence is an operational discipline, not a static state. Two systems holding live semantics over the same BigQuery assets will not stay consistent on their own; one becomes the de facto source of truth whether or not anyone chose it. Every shared term needs an explicit precedence rule or a periodic reconciliation job, and a shared context layer glossary is only as authoritative as the last reconciliation. Skip that and metric drift in enterprise text-to-SQL arrives through a second door.
When does Knowledge Catalog cover what your agents need on its own?
Permalink to “When does Knowledge Catalog cover what your agents need on its own?”There are real conditions under which a second layer is not yet earning its place.
Knowledge Catalog covers what your agents need on its own when your data genuinely sits in BigQuery and Looker, your agent surface is Gemini and Gemini Enterprise only, and no definition has to hold identically on a system Google Cloud does not observe. One more condition is your team: built on Dataplex, it is an infrastructure product that according to Atolio (2026) “typically requires data-engineering resources to configure well.”
Cost belongs in the judgement. Google’s pricing meters Knowledge Catalog in Data Compute Unit hours across Standard and Premium tiers, pay as you go, which Atolio notes “can be hard to plan for.” That is the model, not a comparison, and it is what a context layer TCO assessment across build, buy, and bundle puts to every option, bundling included, alongside the full-stack platform and best-of-breed context layer decision. For a single-vendor estate the bundled answer is defensible.
The neutral layer’s job starts on a specific day: when a second agent vendor asks the same question, or a definition has to hold somewhere Google Cloud does not observe.
How Knowledge Catalog and the Atlan context layer work together on BigQuery
Permalink to “How Knowledge Catalog and the Atlan context layer work together on BigQuery”The coexistence architecture is described most clearly by Google. Chaitanya Pydimukkala, Product Leader at Google Cloud and co-author of its Knowledge Catalog launch post:
“Through our integration with Atlan, we’re building the context layer that connects meaning and quality across even the most complex environments. Knowledge Catalog provides context for structured and unstructured data assets on Google Cloud, and Atlan augments it with machine-readable context about definitions, users, data, and semantics from across the data estate. Together, we’re expanding universal context to multi-cloud and hybrid cloud systems, so teams can maximize value from their data and AI.”
Chaitanya Pydimukkala, Product Leader, Google Cloud
The mechanics, with status on each. Atlan’s Knowledge Catalog connector is in beta, BigQuery-first: it crawls Aspects metadata, ingests Data Quality and Data Profiling scan results at column level, and auto-discovers new Google Cloud projects and Aspect Type definitions. Aspects Reverse-Sync writes context decisions back as Aspect field values, visible in the Google Cloud console. PSC-native transit keeps metadata inside the customer VPC under VPC Service Controls. Atlan’s hosted MCP server is GA, with LookML extraction and field-level lineage production-grade.
Google’s April 2026 launch post names Atlan among the third-party catalogs it supports, so the pattern is documented on both sides rather than asserted on one. Atlan runs the same division of labor alongside other platform-native context surfaces, written up for Snowflake Horizon Context and Genie Ontology, and the sequence in how to implement an enterprise context layer for AI holds here: decide the authority line first, connect second.
Real stories from real customers: context on Google Cloud and across the estate
Permalink to “Real stories from real customers: context on Google Cloud and across the estate”"Context is the differentiator. Atlan gave our teams the shared vocabulary and lineage to move from reactive data management to proactive AI enablement."
— Kiran Panja, Managing Director, Cloud and Data Engineering, CME Group
"Atlan captures Workday's shared language to be leveraged by AI via its MCP server. As part of Atlan's AI labs, we're co-building the semantic layer that AI needs."
— Joe DosSantos, VP Enterprise Data & Analytics, Workday
CME Group runs BigQuery and Looker with on-premises systems in one lineage graph, the estate shape this question is about. Atlan’s public CME Group story puts it at more than 18 million assets, over 1,300 glossary terms, and more than 100 active users in the first year. Workday’s version comes from the agent side: the shared language already existed, and what changed was making it resolvable by an agent over MCP.
See the context layer running alongside a cloud platform
The live demo series walks through the Enterprise Data Graph, reverse-sync into a cloud console, and how an agent resolves one definition across vendors.
Watch a Live DemoCoexistence is an operational discipline
Permalink to “Coexistence is an operational discipline”The axis moved. It used to run from single-cloud to multi-cloud, and Google answered that in April 2026 with federation, third-party aggregation, and an MCP server that ChatGPT and Claude can call. What remains is which layer is authoritative when both hold the same definition.
Both stay true only with work: a precedence rule, a reconciliation cadence, and someone’s name on it. Decide the authority line before you wire the agent harness, because otherwise the first thing you debug is the definition rather than the agent, and that is why context layer evaluation criteria start with authority rather than connector counts.
When four agents on three vendors ask what revenue means, which layer do you want to have answered, and did you decide that or did the tool ordering?
FAQs about Knowledge Catalog and an enterprise context layer for BigQuery
Permalink to “FAQs about Knowledge Catalog and an enterprise context layer for BigQuery”1. What is the Google Cloud Knowledge Catalog?
Permalink to “1. What is the Google Cloud Knowledge Catalog?”Google Cloud Knowledge Catalog is Google’s Gemini-powered context engine for a Google Cloud data estate, which Google describes as an always-on engine unifying structured, unstructured, and SaaS data into agent-ready context. It was renamed from Dataplex Universal Catalog on April 10, 2026.
2. Is Google Cloud Data Catalog still available?
Permalink to “2. Is Google Cloud Data Catalog still available?”No. Legacy Data Catalog was deprecated on February 3, 2025 and shut down on June 1, 2026, and its Business Glossary shut down the same day. Knowledge Catalog is the successor, and the API, CLI, and IAM names carried over unchanged.
3. How much does Google Knowledge Catalog cost?
Permalink to “3. How much does Google Knowledge Catalog cost?”Knowledge Catalog is metered in Data Compute Unit hours across Standard and Premium tiers, pay as you go, with no published flat per-asset price. Because the meter follows usage rather than a fixed inventory, third-party reviewers note the bill is hard to plan for.
4. Can Knowledge Catalog catalog non-Google sources?
Permalink to “4. Can Knowledge Catalog catalog non-Google sources?”Yes, with a status caveat. Google Cloud’s documentation lists federated support for Databricks Unity IRC, AWS Glue Data Catalog IRC, and Snowflake Horizon IRC, plus context federation for five SaaS platforms including SAP and ServiceNow. All are in Preview, and Google notes certain automated semantic extractions might be limited to built-in services.
5. Which parts of Knowledge Catalog are generally available and which are in Preview?
Permalink to “5. Which parts of Knowledge Catalog are generally available and which are in Preview?”Generally available: metadata aggregation, data products, high-precision semantic search, and access-control-aware search. In Preview: context federation, BigQuery measures, deep multimodal extraction, automated context curation, verified queries and semantic guardrails, and the remote MCP server, which Google offers as is with limited support.
6. Does Knowledge Catalog have an MCP server, and which agents can use it?
Permalink to “6. Does Knowledge Catalog have an MCP server, and which agents can use it?”Yes. Google publishes a remote MCP server at dataplex.googleapis.com/mcp, authenticated with OAuth 2.0 and IAM, naming Gemini CLI, ChatGPT, Claude, and custom applications as supported clients. It exposes entry search, schemas, data products, glossaries, and lineage, and its status is Preview.
7. Do you still need an enterprise context layer if you use Knowledge Catalog?
Permalink to “7. Do you still need an enterprise context layer if you use Knowledge Catalog?”It turns on one thing: whether any definition has to hold identically outside what Google Cloud can observe. If your data is BigQuery and Looker and your only agent surface is Gemini, Knowledge Catalog covers it. Once a second agent vendor asks, the neutral layer makes the answer the same either way.
8. What happens when Knowledge Catalog and an enterprise context layer both define the same business term?
Permalink to “8. What happens when Knowledge Catalog and an enterprise context layer both define the same business term?”Without an explicit precedence rule they diverge quietly and nothing alerts. Set the rule per object: Google Cloud is authoritative for schema, lineage, and enforcement, the neutral layer for the approved definition, the contract, and the accountable owner. Then reconcile on a cadence.
Sources
Permalink to “Sources”-
Introducing the Google Cloud Knowledge Catalog, Google Cloud Blog (2026)
-
Use Knowledge Catalog with BigQuery, Google Cloud documentation
-
Use the Knowledge Catalog remote MCP server, Google Cloud documentation
-
What Is Google Knowledge Catalog and How Does It Work?, Atolio (2026)
-
Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update, dbt Labs (2026)
-
How 100 Enterprise CIOs Are Building and Buying Gen AI, a16z (2025)