Data catalog examples show up wherever an AI agent or a human analyst needs verified business context before acting on a number. CME Group traced trading-system lineage in days instead of weeks, and Tide turned a 50-day GDPR tagging process into hours, both using the same underlying catalog capabilities described across a data catalog in general. Below are five industry examples, financial services, healthcare, retail, marketing, and cross-functional data mesh discovery, each run through one consistent frame: the question an agent or analyst gets asked, the context the catalog has to supply to answer it correctly, the capability that supplies that context, and what happens when it’s missing.
- The same arc, every industry. Question, then context needed, then capability, then what happens without it, repeated across finance, healthcare, retail, marketing, and data mesh teams.
- The catalog itself hasn’t changed jobs. Search, lineage, glossary, quality, and classification answer the same questions whether the asker is a person typing into a search box or an agent calling an API.
- What changed is who’s asking. An agent commits to an answer the moment it retrieves one. A wrong table selection isn’t a private confusion anymore; it’s a query, a report, or a decision made on stale data.
Below, we walk through financial services, healthcare, retail, marketing, and cross-functional data mesh discovery, then how the catalog types themselves have changed, and where Atlan fits into supplying that context to an agent.
Where do data catalog examples show up across industries?
The table below maps five industries to the specific question an agent or analyst asks, the context a catalog has to supply, and what breaks when that context is missing.
| Industry | The question an agent or analyst asks | Context the catalog must supply | What happens without it |
|---|---|---|---|
| Financial services | Which system is authoritative for this transaction, and can the number in this report be trusted? | Column-level lineage across branch, online, and mobile systems | Manual reconciliation across systems, discrepancies caught late, audit prep measured in weeks |
| Healthcare | Can this patient field go into an analysis, or does it need to be masked first? | Automatic PHI/PII classification and permission state at query time | Manual review bottlenecks, or a field surfaces where nothing flagged it |
| Retail | Which of these similarly named tables is the one still in production use? | Lineage and popularity metadata showing what’s actually queried | Duplicate effort, stale tables treated as current, wasted compute |
| Marketing | Which dataset has verified order-to-customer lineage I can actually use for this campaign? | Data quality scores, column descriptions, and lineage connecting orders to profiles | The analyst guesses, and the campaign is built on the wrong join |
| Cross-functional / data mesh | What does “churn” mean in the billing domain’s data product, and can I trust it here? | Cross-domain ownership, definitions, and metadata attached to the data product itself | Every team re-derives its own definition, and duplicate metrics multiply |
Anchoring every industry to the specific question an agent or analyst is actually asking, not just the feature that happens to answer it, is what separates a useful example from a generic feature list.
How has the data catalog’s job changed?
Data catalogs used to be measured by how much they covered: how many assets were documented, how many terms were defined, how many people logged in. That coverage frame made sense when a catalog’s job was to be browsed by people. It breaks down once an AI agent is the one asking, because coverage says nothing about whether the agent got the right answer, only whether a page existed somewhere. Three shifts describe how the job itself changed, not just who’s doing it.
| Dimension | Before: manual, siloed, closed | Recent: automated, collaborative, extensible | Now: autonomous, conversational, self-improving |
|---|---|---|---|
| Metadata supply | Humans curate everything by hand; the documentation backlog never closes | Playbooks and AI assistance draft descriptions for stewards to review | Context agents generate and maintain metadata directly from evidence, such as query history and pipeline code |
| Metadata consumption | Context sits in the catalog; the actual work happens everywhere else | Metadata is embedded in the tools people already use, personalized by role | An AI agent is a first-class consumer, retrieving context through MCP, semantic search, and a governed context store |
| Flexibility and learning | What users couldn’t find was never known; the system stayed closed | Open APIs, webhooks, and a marketplace extend what the catalog connects to | Usage traces flow back into the system, so the catalog learns from what it got wrong |
None of this is an argument that catalogs are becoming obsolete. If anything, the opposite case holds: the catalog is more load-bearing now than it was in the coverage era, because an agent depends on it at the exact moment it reasons, not sometime after. Every example that follows sits in that third column. None of them are hypothetical: they’re catalogs already answering the question an agent or analyst asked, not just cataloging the fact that data existed.
Financial services: what an agent needs to trace a transaction for compliance
Before compliance teams, or an AI agent, can trust the number in a report, someone has to trace a transaction’s lineage across systems, and financial institutions can’t afford to make that a manual reconciliation project every time a number gets questioned.
The question an analyst, or increasingly an agent, asks is direct: which system is authoritative for this customer’s transaction history, and can the number in this report be trusted? Answering it requires column-level lineage that follows a deposit from the branch system through online banking, mobile apps, and call center tools, a queryable trail an agent can retrieve at the moment it needs it, not just a diagram of how systems connect.
The capability behind that trail is automated lineage tracking: the catalog watches how data actually moves and keeps that map current as pipelines change, instead of relying on someone to redraw it after every migration. Without it, reconciliation becomes a manual, cross-system exercise, discrepancies get caught weeks after they start compounding, and audit prep turns into a scramble instead of a query.
CME Group automated onboarding and lineage across its cloud and legacy trading systems with Atlan, cutting implementation cycles from weeks to days. The same column-level lineage that AI agents query is what an agent would check today before trusting a transaction figure, whether the person asking is an auditor or an AI assistant working in financial services.
An agent that can’t trace a number back to its source doesn’t just slow down; it guesses, and a compliance team finds out only after the number has already shipped in a report.
Healthcare: how a catalog keeps an AI agent from surfacing protected health data
Healthcare catalogs have to classify sensitive fields automatically, so neither a person building a report nor an AI agent answering a question can surface protected health information without the right permission in place.
The question shows up at the point of use: can this patient field go into an analysis, or does it need to be masked first? Answering it correctly requires the catalog to already know, before anyone asks, which columns hold PHI or PII, and what permission state applies to the person or agent making the request.
The capability is automated sensitive-field classification paired with access alerts: the catalog scans for patterns that indicate protected data, tags them, and notifies data stewards when someone attempts to reach restricted fields, closing the gap between HIPAA compliance policy and what actually happens at query time. Without it, teams either build a manual review bottleneck into every request, or worse, an agent surfaces protected data because nothing was there to flag it.
Tide, a UK digital bank serving nearly 500,000 customers, used Atlan to automate the identification and tagging of personally identifiable information for GDPR compliance, reducing a 50-day manual process to a matter of hours through rule-based automation.
The stakes are the same whether the requester is a compliance analyst or a healthcare AI agent working from the same context layer built for healthcare: a catalog that can’t classify sensitive data at the moment of the query isn’t protecting anyone, it’s just documenting the exposure after the fact.
Retail demand forecasting: connecting sales and inventory data
Forecasting demand accurately depends on one trusted view across point-of-sale, online, and inventory systems, for a retail analyst and, increasingly, for the agent generating that forecast. The question behind that view sounds trivial and isn’t: which of these similarly named tables is the one still in production use? Answering it needs lineage and popularity metadata, not just a list of tables but a signal for which ones are genuinely queried and which are abandoned copies nobody remembered to deprecate.
The capability is automated lineage combined with usage and popularity metrics, surfacing the datasets people, and agents, actually rely on instead of every table that was ever created. Without that signal, teams duplicate effort rebuilding analyses that already exist, treat stale tables as current, and burn compute on assets nobody uses.
Takealot, a South African eCommerce company, used Atlan’s automated lineage and popularity metrics to speed up root-cause analysis across its data estate, and deprecating the unused BigQuery assets it identified through that visibility saved the company $6,000 a year in warehouse costs.
That saving looks small next to a bad forecast, but knowing what’s actually in use before you build on it is the same thing a retail AI agent needs through its own context layer for retail before recommending a reorder quantity. An agent that can’t tell an active table from an abandoned one forecasts confidently against the wrong data, and nobody finds out until the numbers stop matching reality.
Marketing: finding the right customer dataset without a ticket queue
Marketing analysts need customer and order-history data fast enough to act inside a campaign window, not after a data team clears a multi-day request queue.
The question an analyst asks the catalog is specific: which dataset has verified order-history-to-customer-profile lineage I can actually use for this campaign? Answering it requires data quality scores, column-level descriptions, and lineage connecting orders back to customer profiles, all visible at the point of search, not buried in a wiki page written two reorganizations ago.
The capability is catalog search that surfaces quality and lineage signals alongside the search result itself, so the decision about which dataset to trust happens in seconds instead of a Slack thread. Without it, the analyst either guesses, and the campaign runs on the wrong join, or waits on the data team, and the campaign window closes before the data arrives.
A global marketing team needing customer data for a cross-sell campaign found the right order-history dataset in minutes once quality scores and lineage were visible directly in search results, reviewed governance on the same screen, and shared findings the same day, the kind of self-service a context layer for data analytics teams is meant to enable.
That speed only holds if quality and lineage stay current automatically. A search result with stale trust signals is worse than none at all: it tells the analyst, or the agent generating campaign copy, to trust a dataset that no longer deserves it.
Cross-functional discovery: how data mesh teams find data without a central bottleneck
A data mesh model only works if domain teams can discover and trust each other’s data products without routing every request through a central team.
The question that comes up constantly in a mesh is definitional: what does “churn” mean in the billing domain’s data product, and can I trust it for my analysis? Answering it needs cross-domain metadata, ownership, and definitions attached to the data product itself, not just to individual tables, so a consumer in a different domain can tell what they’re actually looking at.
The capability is the catalog acting as a shared discovery layer across domains: definitions, ownership, and usage travel with the data product wherever it’s consumed. Without it, every domain re-derives its own version of “churn” or “active user,” and duplicate metrics quietly multiply until nobody can say which number is correct.
Autodesk implemented Atlan to support its data mesh strategy, giving 60 business domain teams full visibility into how their data products were being consumed and enabling 45 use cases built in two years through self-service discovery, the same pattern a context layer for data governance teams and a context layer for data engineering teams both depend on.
The same problem shows up outside a formal mesh. A telecommunications company analyzing customer churn pulled service interactions, billing, and usage patterns from across departments into one integrated view, using the catalog’s metadata to understand formats and relationships first, the kind of AI-ready context for telecom a churn model needs.
A catalog that can’t carry a definition across domain boundaries doesn’t remove the bottleneck, it just moves it from a ticket queue to a Slack thread, and an agent inherits whichever version of the definition it happens to retrieve first.
What types of data catalogs exist today?
The catalog landscape splits into four broad types. Getting the vendor names right matters now: several of the biggest cloud catalogs changed names or scope in 2026 alone.
Single-cloud, cloud-native catalogs still integrate most tightly with their own provider’s services. Google Cloud’s catalog product is now called Knowledge Catalog, not Data Catalog: Google renamed Dataplex Universal Catalog to Knowledge Catalog in April 2026, and the legacy standalone Data Catalog service begins a phased shutdown starting June 2026. Google now positions Knowledge Catalog explicitly around grounding AI agents in enterprise context through MCP, not just human search. Similarly, “Azure Purview” hasn’t been the correct name since April 2022; Microsoft renamed it Microsoft Purview, and the current catalog capability sits inside Microsoft Purview Unified Catalog. AWS Glue Data Catalog keeps its name and recently added a Preview feature, business context and semantic search, that lets AI agents retrieve domain context such as query patterns and usage rules directly from catalog assets.
Tool-specific, embedded catalogs scan only what was created inside one platform. Tableau Catalog still fits that description: it indexes workbooks, data sources, and flows authored in Tableau, and little else. Databricks Unity Catalog has moved out of that category; it now governs tables, views, volumes, functions, models, and MCP services as a single governance layer for data and AI together, a materially broader scope than a tool-specific catalog.
Custom-built and open-source catalogs, a category DataHub and OpenMetadata anchor, trade a licensing fee for engineering investment: full customization, at the cost of building and maintaining search, lineage, and governance in-house.
| Catalog type | Example platform | What it covers | Where the gap remains |
|---|---|---|---|
| Single-cloud, cloud-native | Google Cloud Knowledge Catalog, Microsoft Purview, AWS Glue Data Catalog | Metadata native to that provider’s own services | Fragmented view once data spans more than one cloud |
| Tool-specific, embedded | Tableau Catalog | Content authored inside that specific BI tool | No visibility outside that tool’s own environment |
| Unified data-and-AI governance | Databricks Unity Catalog | Tables, views, volumes, functions, models, and MCP services in one layer | Scoped to assets that live on that platform |
| Custom-built, open-source | Community-maintained catalogs (DataHub, OpenMetadata) | Fully customizable metadata handling | Ongoing engineering investment to build, secure, and maintain |
None of these types is wrong for the job it was built for. The gap shows up at the seams: a single-cloud or single-tool catalog gives a clean answer inside its own boundary and none once a question crosses into a different cloud, tool, or domain, the seam covered in semantic layer vs data catalog, knowledge graph vs data catalog, data catalog vs context layer, and AWS DataZone vs Glue Data Catalog for teams standardized on one cloud.
According to Gartner (2026), by 2030 universal semantic layers will be treated as critical infrastructure alongside data platforms and cybersecurity, the same seam a single-cloud or single-tool catalog can’t close.
How Atlan turns these examples into context an agent can use
Every example above depends on the same underlying capabilities, search, lineage, glossary, quality, and classification, becoming machine-readable and queryable by an agent at runtime, not just visible in a UI. This is the same argument behind why a data catalog needs to serve agents, not just people: none of it answers whether an agent that asked a real question got the right one back, and if not, why not.
That shift rests on four pieces working together. Context Agents create context directly from evidence, descriptions, lineage, and quality signals, instead of waiting for someone to write it up. The Context Lakehouse stores and vectorizes that context for retrieval through vector search, semantic search, hybrid search, or graph traversal, whichever mode actually answers the question asked, not one keyword box for everything. MCP and conversational AI serve that context to agents and people, one governed connection instead of a page to browse. A learning loop routes what gets missed back to where it can be fixed, so the next answer improves instead of repeating the same gap.
| Catalog capability | What it gives an agent |
|---|---|
| Search and discovery | Which tool or dataset to select for the task |
| Glossary and definitions | The right business term to generate a correct query |
| Lineage | A way to verify where an answer’s numbers actually came from |
| Quality and freshness | Whether to use a dataset or reject it |
| Classification and policy | Permission enforcement at the moment of the query, not after |
| Ownership and certification | Which source to prefer when two datasets disagree |
| Usage history | Which join is likely valid, based on what’s actually been queried before |
The financial services example above needs graph traversal to trace a transaction without loading the whole enterprise data graph; the marketing example needs semantic search that understands “customer order history” even when nobody typed those exact words. Retrieving a scoped sub-graph instead of the whole graph matters for accuracy too: Liu et al. (2023) found language model accuracy drops when relevant information is buried in the middle of a large context window, the same risk agent memory built on a catalog has to retrieve around selectively.
In Atlan’s AI Labs benchmark, adding governed context lifted text-to-SQL accuracy from 16.1% to 22.2%, a 38% relative gain across 174 unique queries and 522 evaluations, the same mechanism behind every example above: an agent with verified context answers more often than one guessing from training data alone. Prukalpa Sankar’s keynote below covers why the catalog’s job changed from cataloging what exists to answering what gets asked.
None of this requires an agent to be smarter. It requires the catalog to have already done the work of becoming machine-readable, which is the actual line between a data catalog and one an AI agent can use.
Real stories from real customers: context that scales across industries
"AI initiatives require more context than ever. Atlan's metadata lakehouse is configurable, intuitive, and able to scale to hundreds of millions of assets. As we're doing this, we're making life easier for data scientists and speeding up innovation."
— Andrew Reiskind, Chief Data Officer, Mastercard
"Atlan captures Workday's shared language to be leveraged by AI via its MCP server. As part of Atlan's AI labs, we're co-building the semantic layer that AI needs."
— Joe DosSantos, VP, Enterprise Data and Analytics, Workday
Getting started with data catalog examples in your organization
Turning these examples into your own catalog starts with the same first step regardless of industry: an inventory of what people, and any agents already in use, actually ask of your data today.
Start there instead of with a tool. Audit the real questions your team fields weekly, then pick the highest-value dataset from one of the five examples above, not the entire estate, as the first target. Automate discovery and lineage before anything else, since classification and glossary work compound on top of a lineage map that already exists, not the other way around, per a full implementation plan. Once core coverage is in place, connect the catalog to wherever agents already query, an MCP server or an internal AI assistant, so the context a person already trusts becomes the context an agent retrieves too.
Two mistakes show up across almost every failed rollout, regardless of industry: buying tooling before anyone has written down the question it needs to answer, and over-ingesting the entire estate on day one instead of scoping to the flows that actually matter. Both are avoidable, and both are more about sequencing than budget.
The catalogs above stopped being measured by what they covered and started being measured by what they answered. That’s the actual shift underneath every industry here, and it’s the same shift a context layer built for AI agents and an enterprise context layer both have to make: not more documentation, but a system that knows when a question came in, whether it got answered correctly, and what to fix when it didn’t.
FAQs about data catalog examples
1. What’s the difference between a data catalog and a data dictionary?
A data catalog provides a broad, business-friendly inventory of data assets with context about ownership, quality, and lineage, while a data dictionary lists technical specifications for developers, such as data types and constraints. Catalogs serve people and agents searching for datasets; dictionaries describe the structure of what they find. Most modern catalogs include dictionary functionality as one component, not a replacement for it.
2. How long does it take to implement a data catalog?
Timelines depend on scope, not total data volume. Teams that automate discovery and start with a small set of critical, high-value datasets can launch a working catalog in six to eight weeks. Comprehensive enterprise rollouts covering thousands of sources typically run three to six months, and starting with the highest-value flows shortens that timeline more reliably than adding headcount.
3. What are the most common use cases for data catalogs?
The most common use cases are accelerating data discovery, enforcing governance and compliance policies, improving data quality through visibility, supporting cross-functional collaboration between technical and business teams, mapping dependencies for migrations, and enabling self-service analytics. A newer use case, common across the examples here, is supplying verified context to an AI agent before it answers a question.
4. Can a data catalog work with multiple cloud platforms?
Enterprise-grade catalogs connect across providers, including Google Cloud, Microsoft Azure, and AWS, plus on-premises environments, avoiding the vendor lock-in that comes with a single cloud’s native catalog. Before choosing one, confirm it supports the specific sources your estate actually uses and can handle the scale of your data, since coverage claims vary significantly between a single-cloud catalog and a genuinely multi-cloud one.
5. How do data catalogs handle sensitive data and compliance?
Catalogs support compliance through automated classification of sensitive fields, access controls tied to that classification, audit trails showing who accessed what, and lineage tracking to support data subject access requests. Organizations use this to quickly locate every location containing PII, financial data, or protected health information, which shortens the time it takes to respond to a regulatory request.
6. How does Atlan help teams adopt data catalogs quickly?
Atlan combines automated discovery with AI-generated context, so metadata and lineage populate without manual entry while stewards review and certify rather than write from scratch. Teams typically launch with a scoped set of core datasets in weeks, then expand coverage incrementally, using built-in templates to automate governance workflows that would otherwise slow adoption down.
7. How are organizations using data catalogs as AI agent infrastructure?
Organizations connect their catalog to AI agents as a source of runtime context, so an agent queries the catalog before generating a response instead of guessing from training data. The catalog identifies the authoritative dataset, verifies quality and ownership, and confirms permissions, the same capabilities in every example above, now serving an agent’s query instead of only a person’s search.
Sources
- Transition from Data Catalog to Knowledge Catalog, Google Cloud
- Knowledge Catalog for AI Agents, Google Cloud
- Azure Purview Is Now Microsoft Purview, Microsoft Azure Blog
- Learn About Microsoft Purview Unified Catalog, Microsoft
- Data Discovery and Cataloging in AWS Glue, AWS
- Unity Catalog Overview, Databricks
- About Tableau Catalog, Tableau
- Enhanced Metadata Improves AI Query Accuracy, Atlan AI Labs
- Lost in the Middle: How Language Models Use Long Contexts, arXiv
- Gartner Announces Top Predictions for Data and Analytics in 2026, Gartner