---
title: "AWS Glue and DataZone vs a Cross-Cloud Context Layer: What Survives the AWS Boundary"
url: "https://atlan.com/know/ai-agent/aws/aws-datazone-vs-cross-cloud-data-governance-platform/"
description: "AWS shipped three documented exits from its native governance stack since November 2025. See what crosses the AWS boundary to a context layer and what does not."
author: "Emily Winks"
author_role: "Data Governance Expert"
published: "2026-08-18"
updated: "2026-08-18T00:00:00.000Z"
---

---

The AWS-native governance stack is four products that behave as one: AWS Glue Data Catalog, AWS Lake Formation, Amazon DataZone and Amazon SageMaker Catalog. Since November 2025 it has had three documented exits from AWS. The third, in AWS's own SageMaker Unified Studio user guide, is bidirectional metadata synchronization with a catalog that is not AWS's: Atlan, with real-time reverse sync. Each exit moves the query boundary outward. No AWS-native context surface carries the glossary term, the owner or the certification state across with it, which is why the third exit is a catalog rather than a feature.

AWS documents both halves: a page for [third-party business data catalog integrations (AWS docs, 2026)](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/third-party-catalog-integrations.html) and a companion page for [the Atlan integration, with scheduled and reverse sync](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/atlan-integration.html). Documenting a partner integration is ordinary enterprise behavior, and it is still the useful fact, because it shows where AWS drew the line between what it builds and what it connects to. [The difference between a data catalog and a context layer](https://atlan.com/know/data-catalog-vs-context-layer/) sits underneath, and a [multi-cloud context layer](https://atlan.com/know/ai-agent/context-layer/multi-cloud-context-layer/) is the category the alternative belongs to.

- Amazon DataZone is still generally available and standalone as of August 2026; SageMaker Catalog is built on it.
- Glue catalog federation reaches Iceberg tables stored in Amazon S3. That qualifier carries more weight than the feature.
- The query boundary moved outward in 2025. The context boundary did not.

| Dimension | AWS-native governance stack | Cross-cloud context layer |
|---|---|---|
| What it is | Four AWS services at four layers, billed separately | One layer above every platform holding governed assets |
| What it governs | AWS-cataloged assets plus federated table pointers | Assets on any platform, plus their agreed meaning |
| Query boundary | Outside AWS since federation shipped, 2025 | Wherever each platform's engine reaches |
| Context boundary | Inside AWS on every shipped or announced surface | At the graph, independent of any platform |
| Identity and policy | Lake Formation and IAM decide who sees what | Reconciles with each native model, owns none |
| Non-AWS assets | Query yes, when Iceberg-on-S3; context no | First-class, same glossary and ownership rules |
| Cost model | Bundled into Glue usage, DataZone free tier | A line item you approve on purpose |
| Best for | AWS-only estates, engineering users | Governed assets on more than one platform |

---

## What does the AWS-native governance stack actually include?

Four products at four layers make up the AWS-native governance stack, plus two that are not yet generally available. Glue Data Catalog is a regional technical metastore, read by Athena, Redshift Spectrum, EMR and Glue ETL. Lake Formation is the permission engine beneath it. Amazon DataZone adds the business layer, domains, projects, glossaries and access requests, per Region per domain. SageMaker Catalog surfaces that same capability inside SageMaker Unified Studio, the studio announced at re:Invent 2024. Not GA: AWS Context, announced June 2026, and Glue's own business-context and semantic-search capability, in preview in four Regions.

One naming correction, because the search results get it wrong. Amazon DataZone was not renamed and not retired. Its governance model was absorbed into SageMaker Unified Studio, where the same capability appears as SageMaker Catalog, and [AWS shipped a domain-upgrade path on 2 June 2025 (AWS What's New, 2025)](https://aws.amazon.com/about-aws/whats-new/2025/06/amazon-datazone-upgrade-domain-sagemaker/) so existing domains can move across. AWS documents [SageMaker Catalog as built on Amazon DataZone (AWS docs, 2026)](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-datazone.html), so the two names describe one governance model, not two competing ones.

Which of the two AWS pieces you need is a separate decision with its own page: [AWS DataZone vs Glue Data Catalog](https://atlan.com/know/ai-agent/aws/aws-datazone-vs-glue-data-catalog/) works through it. The same folder covers [the Amazon Q, QuickSight and Glue bundle](https://atlan.com/know/ai-agent/aws/what-is-amazon-q-quicksight-glue-bundle/), and AWS Context sits near [Amazon Neptune](https://atlan.com/know/ai-agent/knowledge-graph/amazon-neptune-graph-database/) in the AWS graph story.

### Why the four names get confused

Practitioners describe the layering better than the product pages do. One AWS re:Post answer: *"At its core, it operates on AWS Datazone, which is also a solid choice, and this, in turn, is built on Lakeformation."* That shared substrate is why these four are best evaluated as one unit, and why [Lake Formation's IAM-scoped access model](https://atlan.com/know/zero-trust-data-governance/) decides what the layers above it can assert.

---

## Can AWS-native governance reach data outside AWS?

Yes. Since November 2025, AWS has shipped three documented paths to data outside AWS, so any evaluation claiming the stack stops at the AWS edge is out of date. Each carries a different payload.

### Exit 1: Glue Data Catalog federation to remote Iceberg catalogs

According to AWS's Big Data Blog (2025), [catalog federation lets AWS engines query remote Iceberg tables without copying them](https://aws.amazon.com/blogs/big-data/introducing-catalog-federation-for-apache-iceberg-tables-in-the-aws-glue-data-catalog/): *"you can query remote Iceberg tables, stored in Amazon Simple Storage Service (Amazon S3) and cataloged in remote Iceberg catalogs, using AWS analytics engines and without moving or duplicating tables."* Dated 26 November 2025, it names Snowflake Polaris Catalog, Databricks Unity Catalog and custom catalogs supporting the Iceberg REST specification, with Lake Formation vending scoped credentials.

The qualifier does the work: *stored in Amazon S3*. A Snowflake-native table in FDN storage, or a Delta table that is not [Apache Iceberg](https://atlan.com/know/snowflake/apache-iceberg-v3/) on S3, is out of reach. What federates is the table, schema and location, not glossary terms, ownership or certification state. Same limit you hit asking [what a remote catalog actually carries](https://atlan.com/know/ai-agent/semantic-layer/unity-catalog-semantic-layer/) on the [Databricks Unity Catalog](https://atlan.com/know/ai-agent/databricks/unity-catalog-metrics/) side.

### Exit 2: Amazon DataZone's Snowflake data source

This one is more capable than most comparisons admit. According to AWS's DataZone user guide (2026), [DataZone supports Snowflake as a third-party data source](https://docs.aws.amazon.com/datazone/latest/userguide/snowflake-data-source.html) and can *"index your Snowflake databases and schemas, capture column-level lineage from Snowflake's query history"*, then query those tables through federation. The doc names four grant categories: metadata sync via `GRANT REFERENCES`, marked *"no data access"*; lineage sync reading `QUERY_HISTORY` and `ACCESS_HISTORY`; query access; and warehouse usage. Column-level lineage out of another cloud's query history is real, and anyone weighing [a Snowflake estate](https://atlan.com/know/context-layer-for-snowflake/) should credit it.

### Exit 3: AWS's own answer is a third-party catalog

AWS publishes a page for exactly this problem: *"Amazon SageMaker Unified Studio supports metadata synchronization with third-party business data catalog platforms."* On payload its user guide is specific: *"you can synchronize key metadata elements such as projects, assets, descriptions, glossary terms, and their hierarchies."*

Business context crosses when something outside AWS holds it and reconciles it back, which is what a bidirectional sync does. It does not cross on an AWS-native surface, because none of them is scoped past AWS. The partner path is no proof the native stack is broken; it is the documented shape of working past its edge.

| Exit | Mechanism | What crosses | What does not cross | Source, date |
|---|---|---|---|---|
| Glue catalog federation | Remote Iceberg REST catalog, Lake Formation credential vending | Schema and location for Iceberg-on-S3 tables | Glossary terms, ownership, certification; non-Iceberg storage | AWS Big Data Blog, 2025-11-26 |
| DataZone Snowflake source | `aws datazone create-connection`, four grant categories | Databases, schemas, column lineage from query history | Definitions authored elsewhere; other non-AWS platforms | AWS DataZone user guide, 2026 |
| Third-party catalog sync | Bidirectional sync with a partner catalog | Projects, assets, descriptions, glossary terms, hierarchies | Nothing, but the catalog holding it is not an AWS service | SageMaker Unified Studio guide, 2026 |

---

## When is AWS-native governance the right answer?

For an AWS-only estate whose data users are engineers, the AWS-native stack is the right answer and adding a layer above it is overhead. Glue and Lake Formation already share one catalog, so DataZone or SageMaker Catalog adds business governance with no second metadata store to reconcile and no second place for a definition to go stale.

Lineage is the strongest single reason not to reach for something else. AWS states it directly: *"Data lineage in Amazon DataZone is an OpenLineage-compatible feature that can help you to capture and visualize lineage events."* Per [AWS's DataZone lineage documentation (2026)](https://docs.aws.amazon.com/datazone/latest/userguide/datazone-data-lineage.html), capture can run automatically for AWS Glue and Amazon Redshift databases added to a domain, and Glue Spark jobs on version 5.0 and above can be configured to send lineage events into a domain. An open standard, wired in, at no integration cost.

The economics are bundled and the comparison is not neutral. According to AWS's Amazon DataZone pricing page (2026), [each account gets 20 MB of free metadata storage, 4,000 free API requests and 0.2 free compute units per month](https://aws.amazon.com/datazone/pricing/), with compute at $1.776 per unit beyond that and core APIs including `CreateDomain` and `Search` excluded from the count. Glue pricing folds into Glue usage you already have, while a cross-cloud layer is a line item somebody approves, which is the honest framing of [build, buy or bundle](https://atlan.com/know/ai-agent/context-layer/context-layer-tco-build-vs-buy-vs-bundle/) here. DataZone administers cross-account sharing for you, and if the modeling work lives in SageMaker, SageMaker Unified Studio is an integrated place to do it, including for [Bedrock-based enterprise agents](https://atlan.com/know/ai-agent/ai-agent-applications/aws-bedrock-for-enterprise-agents/).

Breadth is the standard objection to any neutral layer. Atlan's AWS-side coverage is deep rather than nominal, with documented connectors for [AWS Glue Data Catalog](https://docs.atlan.com/apps/connectors/etl-tools/aws-glue), Athena, Redshift, Lake Formation, Amazon S3, SageMaker and SageMaker Unified Studio, and EMR. Three conditions change the answer, and only these three: the estate stops being AWS-only, business users need to participate in governing assets rather than only reading them, or context has to reach assets AWS does not catalog. Until one is true, use what is in the account rather than a [DIY context layer](https://atlan.com/know/ai-agent/context-layer/diy-context-layer/) stitched together by hand.

  The AI context stack, layer by layer
  A four-layer blueprint for deciding where each piece of context should live when governed assets sit on more than one cloud.
  Get the AI Context Stack

---

## Why does a federated table arrive without business context?

Because federation moves the query boundary and nothing else. A Snowflake table reachable through Glue arrives with a schema and a location, and with no glossary term, no owner and no certification state, because none of those is part of what an Iceberg REST catalog exchange transmits.

Every context surface AWS has shipped or announced for agents reads AWS. Glue's business-context and semantic-search capability reads Glue, and as of June 2026 remains a [preview in four Regions (AWS What's New, 2026)](https://aws.amazon.com/about-aws/whats-new/2026/06/aws-glue-data-catalog/). [AWS Context's integration list (AWS ML Blog, 2026)](https://aws.amazon.com/blogs/machine-learning/context-intelligence-for-your-data-and-ai-agents-at-scale/) is Glue Data Catalog, SageMaker Unified Studio and Lake Formation, all AWS services, and as of August 2026 a [third-party review describes it as coming soon (Caylent, 2026)](https://caylent.com/blog/aws-context-aws-automated-knowledge-graph-for-ai-agents). That is a snapshot of a service in flight, not a verdict on where it lands. The domain is itself a regional object: [DataZone quotas are stated per Region per domain (AWS General Reference, 2026)](https://docs.aws.amazon.com/general/latest/gr/datazone.html), and [cross-Region table access requires resource links (AWS docs, 2026)](https://docs.aws.amazon.com/lake-formation/latest/dg/cross-region-access.html).

| AWS surface | Reaches non-AWS assets? | Carries glossary, owner, certification? | Region scope |
|---|---|---|---|
| Glue catalog federation | Yes, Iceberg tables on Amazon S3 | No | Per Region, per account |
| DataZone Snowflake source | Yes, Snowflake only | Structure and lineage, not definitions | Per Region, per domain |
| Glue business context, semantic search | No | Yes, for Glue assets | Preview, 4 Regions |
| AWS Context | No, integration list is AWS services | Announced, not generally available | Not published |
| Lake Formation and IAM | Permissions on AWS-cataloged data | No, it is an access decision | Cross-Region via resource links |

Now the counter. AWS could extend the same federation pattern to glossary terms and owners without changing anything architectural, which makes this look like implementation debt rather than a boundary. The answer has nothing to do with roadmaps. A governance domain is scoped to the identity and policy model that owns it, and AWS's is Lake Formation and IAM. [Business context](https://atlan.com/know/business-context-for-ai/) is an assertion *about* an asset made by a person with authority over it, so the surface holding it inherits the scope of that authority model. Federating a table pointer requires agreement on a format. Federating a certification state requires agreement on [who is allowed to certify](https://atlan.com/know/ai-agent-access-control/). And a regional object cannot be the home of an ontology that spans an enterprise, let alone several clouds.

You can query a Snowflake table through Glue today and hand an agent zero business context on it, the exact state agents fail in, and why [the types of metadata an agent needs](https://atlan.com/know/ai-agent/data-for-ai/types-of-metadata-for-ai-agents/) is a longer list than schema. According to Gartner analyst Andrés García-Rodeja, speaking at the Gartner Data & Analytics Summit in March 2026, by 2028, 60% of agentic analytics projects relying solely on the Model Context Protocol will fail for lack of a consistent semantic layer. Glue's skill assets make it [an MCP-connected catalog](https://atlan.com/know/mcp-connected-data-catalog/) for the assets Glue holds, and that scoping is what AWS Context blurs when read as a [context layer vs knowledge graph](https://atlan.com/know/ai-agent/context-layer/context-layer-vs-knowledge-graph/) question. Nothing here says business context cannot leave AWS. It says no AWS-native context surface carries it across.

---

## What is a cross-cloud context layer?

A cross-cloud context layer is the surface where business context about an asset lives independently of the platform the asset sits on. Its unit of work is not a table but an assertion: this dataset means this thing, this person owns it, this version is certified, this is the chain behind it. It serves the people making those assertions and the agents reading them, which is why an [enterprise context layer](https://atlan.com/know/what-is-the-enterprise-context-layer/) is defined by what travels rather than where it is hosted.

Why now is less strategic than most vendors admit. According to the Flexera 2026 State of the Cloud Report, [multi-cloud footprints are frequently inherited unintentionally](https://www.flexera.com/blog/finops/flexera-2026-state-of-the-cloud-report-the-convergence-of-cloud-and-value/), through acquisitions and through teams that chose independently. Nobody designed the estate that now has to be governed as one.

One honest limit on the open-format argument. The [Apache Iceberg REST Catalog specification](https://iceberg.apache.org/rest-catalog-spec/) is what makes AWS's first exit possible, and it is a real standard doing real work. An open table format is not open context: Iceberg standardizes how a table is described, not who is allowed to say what it means.

### Core components of a cross-cloud context layer

A graph that spans platforms, so one asset has one identity whichever engine reads it. Glossary terms, ownership and certification that travel with the asset instead of with the console. Bidirectional synchronization with each platform's native catalog, which stays the system of record for technical metadata. Lineage stitched across boundaries rather than restarting at each one. And a retrieval surface an agent can reach, the part an [agent context layer](https://atlan.com/know/agent-context-layer/) makes or breaks. Those five hold together as a [reference architecture](https://atlan.com/know/ai-agent/context-layer/context-layer-reference-architecture/) because dropping any one puts the ontology back inside somebody's platform, and a [semantic layer for AI agents](https://atlan.com/know/ai-agent/semantic-layer-for-ai-agents/) built that way inherits whichever boundary its host draws.

---

## AWS-native stack vs cross-cloud context layer: head-to-head

The two options converge on the same underlying assets and diverge on where context is allowed to live. Read the table on that axis rather than as a feature count: most rows follow from one choice, whether the authority model governing an asset is the same one storing it.

| Dimension | AWS-native stack | Cross-cloud context layer |
|---|---|---|
| Primary scope | AWS-cataloged assets plus federated Iceberg pointers | Every platform holding governed assets |
| Metadata stores | One, already in the account | Two, native catalog stays system of record |
| Non-AWS assets | Query yes, context no | Context first-class, query stays with the platform |
| Region model | Regional domains, resource links to cross Regions | Enterprise-wide graph, no regional partition |
| Users | Engineers, plus business users inside DataZone | Business users, stewards, engineers, agents |
| Setup path | Console and blueprints; Snowflake source CLI-only | IAM role via CloudFormation, then connectors |
| Cost model | Bundled into Glue usage, DataZone free tier | A line item, sized to platform count |
| Failure mode | Green sync, empty lineage, no context off-AWS | Drift from a native catalog if sync lapses |
| Time to value | Fast inside AWS, stops at the first non-AWS platform | Slower to stand up, no platform edge to stop at |

Take one scenario. A Snowflake Iceberg table on S3, federated into Glue, queried by Athena, reached by an agent over MCP. AWS contributes the query path, the permission decision and credential vending, and does all three well. The cross-cloud layer contributes the glossary term that tells the agent what the table means, the owner to escalate to, the certification state deciding whether the number is safe to quote, and the lineage back to the Snowflake query history. Remove the AWS half and the query fails. Remove the context half and the query succeeds, and the answer is unattributable, which is worse, because nothing surfaces to say so.

### Five questions that settle this, if the table does not

How many platforms will hold governed assets eighteen months from now? Who is allowed to certify an asset, and does that authority exist outside one cloud's IAM? Do business users need to participate in governing assets, or only engineers? Does context need to survive a platform inherited through an acquisition? And what happens to your ontology when a platform is retired? Those five are the working version of [context layer evaluation criteria](https://atlan.com/know/ai-agent/context-layer/context-layer-evaluation-criteria/), and they decide whether you pick [a full-stack platform or a best-of-breed layer](https://atlan.com/know/ai-agent/context-layer/full-stack-ai-platform-vs-best-of-breed-context-layer/) on purpose rather than by default. The field version is one question you can ask in a meeting: what does this do for your Snowflake, GCP or Azure assets? Answering all five is most of an [AI governance framework](https://atlan.com/know/ai-readiness/ai-governance-framework/), the working definition of [cross-cloud data governance](https://atlan.com/know/data-governance/multi-cloud-data-governance/), and the shape of the [context graphs for AI agents](https://atlan.com/know/context-graphs-for-ai-agents/) an agent reads.

  What does a second context store actually cost you?
  Model the bundled option against a cross-cloud layer using your own platform count, user mix and rework hours.
  Open the ROI Calculator

---

## What goes wrong when metadata crosses the AWS boundary?

The crossings that do exist fail in five documented ways, four of them documented by AWS itself. All describe behavior current as of August 2026.

| Failure mode | What you see | Documented cause |
|---|---|---|
| Lineage fails silently | Sync reports success, lineage graph is empty | `GRANT IMPORTED PRIVILEGES ON DATABASE SNOWFLAKE` missing |
| Identifier case is normalized | Names arrive re-cased, exact-match lookups miss | `CATALOG_CASING_FILTER` defaults to `UPPERCASE_ONLY` |
| One connection, one schema | Many connections to cover one wide estate | `DATABASE` and `SCHEMA` properties are singular |
| Setup is CLI-only | No console path to hand a steward | Documented flow is `aws datazone create-connection` |
| Federation is Iceberg-on-S3 shaped | Tables in other storage are simply absent | Federation covers remote Iceberg catalogs over S3 |

The first deserves emphasis for how it presents. AWS flags it itself: without `GRANT IMPORTED PRIVILEGES ON DATABASE SNOWFLAKE`, *"the metadata sync succeeds but no lineage records are captured."* A green sync with an empty lineage graph is the worst state to hand an agent, and why [column-level lineage an agent can rely on](https://atlan.com/know/ai-agent/data-for-ai/data-lineage-for-ai/) has to be verified rather than assumed. Glue-side capture has its own edges: a crawler data-source run covers Amazon S3, DynamoDB, Delta Lake, Iceberg and Hudi tables on S3, JDBC and DocumentDB and MongoDB are unsupported, and AWS notes the lineage run fails past 100 tables.

Case normalization is the quiet one. `CATALOG_CASING_FILTER` accepts `UPPERCASE_ONLY` or `LOWERCASE_ONLY`, and AWS states that *"If omitted, defaults to UPPERCASE_ONLY."* An agent matching on a name it was told is exact will miss, a retrieval failure that looks like an empty result rather than an error, and grasping [why MCP matters for AI agents](https://atlan.com/know/mcp/why-mcp-matters-for-ai-agents/) shows how far that empty result travels before anyone notices. Add the quota shape, per Region per domain: 1,000,000 assets, 10,000 business glossary terms, 25 data source runs per data source per day. None of these is a defect. Each is a documented property that belongs in an [AI risk management](https://atlan.com/know/ai-readiness/ai-risk-management/) register before an agent depends on the crossing.

---

## How Atlan approaches context across AWS and non-AWS estates

Atlan sits on top of Glue, DataZone and SageMaker Unified Studio rather than replacing any of them, and AWS documents the integration in its own user guide. The starting problem is stated by AWS engineers, not by us. Leonardo Gomez, Principal Analytics Specialist Solutions Architect at AWS, on the AWS Big Data Blog in December 2025: *"Without tight integration between these systems, metadata becomes fragmented. A single asset can appear under different names, documentation might drift out of sync, and governance signals can become inconsistent across systems."*

The strongest illustration is Amazon's own. According to [AWS's May 2026 account of consolidating Amazon's internal catalogs](https://aws.amazon.com/blogs/big-data/how-amazon-is-moving-to-integrate-catalogs-to-improve-data-discovery-with-amazon-sagemaker/), Amazon teams were creating datasets and non-tabular assets outside the enterprise catalog, leaving users to search several platforms to find one thing. Fragmentation is not a vendor talking point when Amazon publishes its own case of it.

On mechanism, AWS's user guide is specific: *"On-demand and scheduled bidirectional metadata synchronization"*, *"Synchronization of glossary terms and descriptions, including parent-child relationships"*, and *"Real-time reverse sync of metadata updates from Atlan back to Amazon SageMaker Catalog."* AWS documents the connection too, an IAM role deployed by CloudFormation template following least privilege. Underneath, the Enterprise Data Graph spans platforms, and **Context Agents** and **Context Engineering Studio** are how the context on it gets built and kept current, which is [context engineering for AI agents](https://atlan.com/know/context-engineering-for-ai-agents/) applied to an estate rather than a prompt.

The pattern repeats wherever a platform ships its own context surface, which is why [Snowflake Horizon](https://atlan.com/know/snowflake/snowflake-horizon-context-and-atlan-context-layer/) and [the same seam on Databricks](https://atlan.com/know/ai-agent/databricks/genie-ontology-and-atlan-context-layer/) read as versions of one argument. Teams commonly start with AWS-native metadata because it is there and it works, then add a cross-cloud layer when one of the three conditions above turns true.

  Is your context ready for an agent that crosses clouds?
  A checklist for testing whether the glossary terms, owners and certification states an agent needs reach every platform it queries.
  Run the Readiness Check

---

## Why the context boundary sits where the authority model sits

The AWS-native stack and a cross-cloud context layer answer two different questions, and AWS's own documentation now separates them: what can this engine reach, and who is allowed to say what it means. AWS will keep pushing the first boundary outward, because federation is a format problem and formats get standardized. The second keeps sitting wherever the identity and policy model sits, because certifying an asset is an act of authority and authority does not federate on a spec. None of this makes the native stack broken. It makes it normal, with a documented edge and a documented way past it, which is the read to hold while [implementing an enterprise context layer for AI](https://atlan.com/know/how-to/implement-enterprise-context-layer-for-ai/) and the constraint that surfaces first in [the agent harness that consumes this context](https://atlan.com/know/how-to-build-ai-agent-harness/), usually as an answer nobody can trace.

  Book a Demo

---

## FAQs about AWS-native governance vs a cross-cloud context layer

### 1. Is Amazon DataZone the same as SageMaker Catalog?

No. One governance model, two surfaces. Amazon DataZone is the standalone service with its own console and API; SageMaker Catalog is the same capability inside SageMaker Unified Studio, built on DataZone. A June 2025 upgrade path moves existing domains across.

### 2. Can AWS Glue Data Catalog see data in Snowflake or Databricks?

Yes, with one qualifier. Glue catalog federation, announced November 2025, lets AWS analytics engines query remote Iceberg tables cataloged in Snowflake Polaris Catalog, Databricks Unity Catalog or any Iceberg REST catalog, but only where those tables sit in Amazon S3.

### 3. Is Amazon DataZone deprecated?

No. As of August 2026 it is generally available and still standalone, with its own user guide, console and quotas. Its governance model was absorbed into SageMaker Unified Studio and surfaced there as SageMaker Catalog, but neither that nor the upgrade path retired it.

### 4. Does Amazon DataZone support non-AWS data sources?

Yes, Snowflake. DataZone can index Snowflake databases and schemas, capture column-level lineage from Snowflake's query history, and support federated query. AWS splits the required grants into four categories: metadata sync, structure only with no data access, lineage sync, query access, and warehouse usage.

### 5. Can AI agents get business context on data outside AWS through AWS-native services?

Not today. No AWS-native context surface carries glossary terms, ownership or certification state across the AWS boundary. Glue's business-context capability reads Glue and remains in preview in four Regions, and AWS Context integrates only with AWS services. AWS documents third-party catalog sync as the mechanism.

### 6. What are the limitations of AWS-native data governance for a multi-cloud estate?

Three structural ones. Governance domains are regional objects, with quotas stated per Region per domain. Cross-platform reach is query-shaped, so a federated table arrives with schema and location but no business meaning. And the authority model is Lake Formation and IAM, so assertions about outside assets have nowhere native to live.

### 7. What is cross-cloud data governance?

Cross-cloud data governance is applying one set of definitions, ownership rules and access policies to data sitting on more than one cloud platform. It matters because most multi-cloud estates were inherited rather than designed, so the governance boundary has to be drawn independently of any provider.

---

## Sources

1. [Catalog federation for Apache Iceberg tables, AWS Big Data Blog](https://aws.amazon.com/blogs/big-data/introducing-catalog-federation-for-apache-iceberg-tables-in-the-aws-glue-data-catalog/)
2. [Connect Snowflake as a data source in Amazon DataZone, AWS Documentation](https://docs.aws.amazon.com/datazone/latest/userguide/snowflake-data-source.html)
3. [Third-party business data catalog integrations, AWS Documentation](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/third-party-catalog-integrations.html)
4. [Atlan integration, AWS Documentation](https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/atlan-integration.html)
5. [Unifying governance and metadata across SageMaker Unified Studio and Atlan, AWS Big Data Blog](https://aws.amazon.com/blogs/big-data/unifying-governance-and-metadata-across-amazon-sagemaker-unified-studio-and-atlan/)
6. [How Amazon is integrating catalogs to improve data discovery, AWS Big Data Blog](https://aws.amazon.com/blogs/big-data/how-amazon-is-moving-to-integrate-catalogs-to-improve-data-discovery-with-amazon-sagemaker/)
7. [Data lineage in Amazon DataZone, AWS Documentation](https://docs.aws.amazon.com/datazone/latest/userguide/datazone-data-lineage.html)
8. [Amazon DataZone pricing, AWS](https://aws.amazon.com/datazone/pricing/)
9. [Amazon DataZone endpoints and quotas, AWS General Reference](https://docs.aws.amazon.com/general/latest/gr/datazone.html)
10. [SageMaker Catalog and Amazon DataZone, AWS Documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-datazone.html)
11. [Amazon DataZone upgrade domain to SageMaker, AWS What's New](https://aws.amazon.com/about-aws/whats-new/2025/06/amazon-datazone-upgrade-domain-sagemaker/)
12. [Glue Data Catalog business context and semantic search, AWS What's New](https://aws.amazon.com/about-aws/whats-new/2026/06/aws-glue-data-catalog/)
13. [Context intelligence for your data and AI agents at scale, AWS ML Blog](https://aws.amazon.com/blogs/machine-learning/context-intelligence-for-your-data-and-ai-agents-at-scale/)
14. [AWS Context, an automated knowledge graph for AI agents, Caylent](https://caylent.com/blog/aws-context-aws-automated-knowledge-graph-for-ai-agents)
15. [Cross-Region access in AWS Lake Formation, AWS Documentation](https://docs.aws.amazon.com/lake-formation/latest/dg/cross-region-access.html)
16. [Iceberg REST Catalog specification, Apache Iceberg](https://iceberg.apache.org/rest-catalog-spec/)
17. [2026 State of the Cloud Report, Flexera](https://www.flexera.com/blog/finops/flexera-2026-state-of-the-cloud-report-the-convergence-of-cloud-and-value/)
18. [Crawl AWS Glue, Atlan Documentation](https://docs.atlan.com/apps/connectors/etl-tools/aws-glue)
19. [Data & Analytics Summit, Andrés García-Rodeja, Gartner](https://www.gartner.com/en/conferences/na/data-analytics-us)