A data catalog and master data management fail at opposite ends of the same pipe. Master data management settles which version of a record is authoritative at write time and hands back a golden record; a data catalog answers what exists, what is current, and who may use it, after the record exists. According to the IBM Institute for Business Value (2025), 43% of COOs rank data quality as their most significant data priority, and which discipline fixes that depends on where your failure sits. Each answers a plurality problem, and whatever you skip is a gap in the context layer your agents query.
Most enterprises buy the wrong one first because they diagnose the symptom instead of the failure point. The symptom is always the same sentence: nobody trusts the number. The enterprise context layer above them carries only the answer one of them produced, a separate question from data catalog vs context layer.
- Buy master data management when one entity is authored in more than one operational system.
- Buy a data catalog when the records are correct and nobody, human or agent, can find, date, or clear them.
- Buy neither yet when you run one system of record per domain. Without a plurality, both become shelfware.
- Neither fixes a business that never agreed what a customer is.
| Dimension | Data catalog | Master data management |
|---|---|---|
| What it is | Searchable inventory of data assets and their context | Decides which version of an entity wins |
| The question it answers | Which asset should I use, and may I? | Which record is the real one? |
| Who complains first | Analysts and agents who cannot find or date what they need | Finance, operations, customers billed twice |
| Who owns it | Data platform and analytics teams | Operations, finance, or a master data function |
| Buy it when | Readers outnumber the people you can brief in person | One entity is authored in more than one system |
| Skip it when | Everyone who needs the data knows where it is | You have one system of record per domain |
What is a data catalog, and what does it actually fix?
Permalink to “What is a data catalog, and what does it actually fix?”A data catalog is a searchable inventory of data assets that answers read-time questions: what exists, where it is, whether it is current, and whether you may use it. It collects metadata from source systems into one view of everything a person or an agent can reach, and every question it answers is asked after the record already exists.
Its scope is wide: tables, files, dashboards, models, and the terms attached to them. That is the surface a data catalog for AI has to expose, which makes the types of metadata AI agents need a design question rather than a documentation chore, and lineage for AI agents a retrieval problem.
Core components of a data catalog
Permalink to “Core components of a data catalog”Five capabilities do the work: metadata management (source, type, owner, freshness, kept current as the estate changes), discovery and search, data profiling for null rates and anomalies, collaboration on the asset itself, and integration breadth across warehouses, lakes, BI tools, and operational systems.
Its value is conditional on plural readers: it pays when more people and agents query the data than can be told in person, and does nothing for a record written wrong, the boundary drawn in semantic layer vs data catalog.
What is master data management, and what does it actually fix?
Permalink to “What is master data management, and what does it actually fix?”Master data management decides which version of a shared business entity is authoritative, at write time, before a transaction commits. Its subject is non-transactional master data: customer, product, supplier, employee. Its deliverable is a golden record. Records that name the same entity are matched, merged under survivorship rules, and whatever the rules cannot settle goes to a steward.
These platforms synchronize that data across the systems that author it, and they run on stewardship as an operational role rather than a committee, because the accountability is the mechanism. The line against cataloging is sequence, not scope, which is why systems of record frames the write side.
Practitioners ask first whether the category is still alive. It is. According to Gartner (2026), the firm retired its Magic Quadrant for Master Data Management Solutions in 2022 and published a new edition on 6 April 2026 evaluating roughly 20 vendors. An analyst firm does not un-retire a Magic Quadrant for a dead category.
Core components of master data management
Permalink to “Core components of master data management”Five again: data integration into the systems that author entities, standardization and deduplication, stewardship workflow so a person resolves what the rules cannot, data quality management on mastered attributes (what data quality for AI agents applies downstream), and hierarchy management for product categories and legal entity trees, where most implementation time goes.
It pays when writers are plural. One system of record per domain and it has nothing to adjudicate, the honest reason a share of these programs disappoint, and why master data management is not metadata management.
How does a data catalog differ from master data management?
Permalink to “How does a data catalog differ from master data management?”The two disciplines fail at opposite ends of the same pipe, a sharper line than scope or posture. A data catalog provides visibility across a wide array of data assets; master data management maintains the consistency of specific key entities. Correct, and where every other page stops.
Master data management earns its budget where records are written: one real-world entity authored into more than one operational system, authority settled before a transaction commits. A catalog earns it where records are read. The records are fine. Nobody, human or agent, can find them, tell which of four similar tables is current, or know whether they may use it.
Enterprises buy wrong because the symptom is identical and the failure point is not. According to the IBM Institute for Business Value (2025), more than 25% of organizations lose over $5 million a year to poor data quality and 7% lose over $25 million. That funds either purchase, which is why the same confusion runs through active metadata vs context layer and systems of semantics. On a thread on r/databricks titled “Databricks MDM / Data governance,” a participant asks how master data management differs from catalog solutions, naming Atlan among them. It is also the failure behind most wrong-join complaints in text-to-SQL for enterprise: the SQL is valid, the join key is not, and nothing objects.
| Dimension | Data catalog | Master data management |
|---|---|---|
| Primary purpose | An inventory of datasets, so users discover and understand data | Accuracy and uniformity of key data entities |
| Key functionalities | Metadata management, discovery, profiling, collaboration, integration | Integration, deduplication, stewardship workflow, data quality, hierarchies |
| Main benefits | Faster discovery, shared understanding, policy at point of use | One authoritative version, audit defensibility, fewer reconciliations |
| Use case examples | Finding purchase history without asking IT; settling a dataset dispute | Collapsing duplicate patient records; one customer view |
| Point in the data lifecycle | Read time, after the record exists | Write time, before the transaction commits |
| Unit of work | An asset: a table, a file, a dashboard, a term | An entity instance: this customer, this supplier |
| Output artifact | An entry with lineage, ownership, freshness, certification | A golden record with match rules and survivorship history |
| Enforcement point | The query, the access request, the agent call | The write path, the stewardship queue |
| Time to value | Weeks to a first useful domain, then continuous | 3 to 6 months cloud, 9 to 18 months on-premises |
| Failure mode when done badly | A complete inventory nobody opens, bought for optics | An 18-month program that mastered undisputed attributes |
| What it silently cannot fix | A duplicate customer record: the catalog sits downstream of it | Four similar tables nobody can date: never in scope |
Which failure do you actually have? A five-question diagnostic
Permalink to “Which failure do you actually have? A five-question diagnostic”Five questions, answered against your own systems rather than a category description, tell you which discipline you are buying. No other page on this result publishes a test, which is why sequencing gets settled by whichever vendor got the meeting first.
- Is the same entity created in more than one system? Not copied into, created in. A customer typed into a CRM and also into billing is a write-time problem; one created once and replicated is not.
- Does somebody have to decide which version wins before a transaction commits? If the wrong answer produces a duplicate invoice rather than a wrong chart, the failure is upstream of the reader.
- Is the dispute about which record is right, or which table to use? The first is write time, the second read time, and they get confused in the same meeting.
- How many people and agents read this data without being able to ask its owner? Past the size of a room, discovery stops working through relationships.
- Do you have one system of record per domain? If yes, you have neither plurality.
Answering honestly is also the cheapest way to write a data contract somebody holds, and the conversation a context layer for data governance teams starts from rather than finishes with.
The table in the next section runs eleven symptoms through this same test and prints the verdict beside each one.
Buy neither yet deserves the same weight as the other two verdicts. Each discipline answers plurality, plural writers or plural readers, and with neither present the software has nothing to do and the program produces documents instead of decisions.
Data catalog and master data management use cases, side by side
Permalink to “Data catalog and master data management use cases, side by side”Run the write-or-read question against real scenarios and eight familiar enterprise complaints sort themselves, four to each side. In the read-time four the records are individually correct and nobody can find, date, or clear them, which is what data governance evidence assembled from an inventory can prove and mastering cannot. In the write-time four the records exist and disagree, which is why document intelligence for enterprise treats entity resolution, cataloging, and access policy as one governed substrate rather than three purchases, and why metadata management for AI has to know which side of the pipe a signal came from. The table sorts those eight alongside the diagnostic’s own symptoms. The sorting is the whole decision, one case at a time.
| Scenario, as you actually see it | Failure point | Which discipline answers it, and what to buy | What will not help, and why |
|---|---|---|---|
| Marketer hunting customer purchase history, analyst asking weekly where a dataset lives | Read time | A data catalog | Mastering the entity does not locate the table |
| Four similarly named tables nobody can date | Read time | A data catalog | Mastering entities never in dispute |
| Financial institution proving who accessed what and why | Read time | A data catalog | A golden record carries no access history |
| An agent that cannot tell whether it may use a table | Read time | A data catalog | Deduplicating the rows inside it |
| Manufacturer whose departments read one dataset differently | Read time | A data catalog | The records agree; the readings do not |
| Non-profit staff who cannot interpret their own fields | Read time | A data catalog | There is no entity conflict to resolve |
| Patient records duplicated and conflicting across clinical systems | Write time | Master data management | Cataloging duplicates documents them; lineage on the conflicting tables does not remove them |
| Customer split across storefront, CRM, and support, then billed twice under two account IDs | Write time | Master data management | Lineage shows three sources, not which identity is real, and the record stays findable and wrong |
| Duplicate vendor records across regional offices, one supplier onboarded three times | Write time | Master data management | An inventory of the duplicates, however well documented, is still duplicates |
| Insurer needing audit-defensible current records | Write time | Master data management | Discovery does not establish authority |
| One system of record per domain, one disputed metric | Neither yet | Neither. Agree the metric definition first | Both. There is no plurality |
When does master data management earn its budget?
Permalink to “When does master data management earn its budget?”Master data management earns its budget when more than one operational system authors the same entity and the cost of getting authority wrong is a transaction rather than a report. The triggers are easy to check: duplicate customers billed twice, a supplier onboarded three times, a SKU meaning two things in two ERPs. If none has happened, the business case is theoretical.
The cost shape is what nobody on this topic publishes. According to OvalEdge (2026), a vendor publication rather than an analyst firm, cloud-native implementations typically go live in 3 to 6 months while on-premises and hybrid deployments run 9 to 18 months, close to the roughly 9-month average reported for SAP Master Data Governance in user review data. Budget stewardship headcount separately, because the exceptions queue is permanent, and read the duration against what a context layer ROI model assumes about first value.
The strongest opposing case belongs to Malcolm Hawker, Chief Data Officer at Profisee and host of the CDO Matters podcast. His position, paraphrased rather than quoted because the available text is a truncated search snippet: business users never ask for a catalog, they complain about duplicate customer records, reports that disagree, and reconciliation that takes days. Cataloging is oriented inward, toward the data function’s efficiency; master data management is oriented outward, toward the accuracy of what the business receives. So spend there first. He sells master data management software, worth saying plainly, and it does not make him wrong.
For write-time failures he is right: if one entity is authored in three systems, priority belongs to the discipline that rules on authority, and an inventory of three conflicting records only describes the problem more precisely. The framing stops working where the records are correct and nobody can find, date, or clear them. That is not data-team efficiency, that is a business user waiting. If your answers point this way, the next question is which platform, and enterprise alternatives to SAP Master Data Governance compares five on architecture and price.
When does a data catalog earn its budget?
Permalink to “When does a data catalog earn its budget?”A data catalog earns its budget when the number of people and agents reading the data exceeds the number who can be told about it in person. That threshold arrives twice: when analyst headcount outgrows tribal knowledge, and again when agents query the same estate with no relationship to trade on.
The triggers are as checkable as the write-time ones. Four similarly named tables and no way to tell which is current. The same question about where a dataset lives, every week, in the same Slack channel. A published number two teams reproduce differently because they picked different sources. An agent that cannot tell whether it may use a table, the gap AI agents for data catalog work addresses and why the agent memory layer and the data catalog question comes up at all.
The honest charge against catalogs is that they get built and then go unused, and it is often true. The condition is specific: bought for compliance optics, scoped to complete coverage rather than a named read-time failure, staffed by nobody once the rollout ends. One that starts with a single domain where readers are demonstrably stuck does not fail this way, which is most of what how to prepare enterprise data for AI agents is about. The line runs the other way too: no amount of cataloging fixes a duplicate customer record, because its job is to describe assets accurately, including the inaccurate ones.
What tools do data catalogs and master data management need?
Permalink to “What tools do data catalogs and master data management need?”The tooling requirement splits the same way the failure does: catalog tooling is built for retrieval, master data management tooling for adjudication. That is why the two rarely converge into one product even when a single vendor sells both.
Data catalog tools create, maintain, and search an inventory of datasets: metadata management, discovery and search, data profiling, collaboration, and integration breadth across databases, warehouses, lakes, and cloud storage. Data quality metrics surface here as signals attached to assets rather than as enforcement, and the questions worth asking are collected in context layer evaluation criteria.
Master data management tools standardize and synchronize master data across systems: integration into the systems that author entities, standardization and deduplication, stewardship workflow for the exceptions rules cannot resolve, data quality management on mastered attributes, and hierarchy management. Examples include SAP Master Data Governance, IBM InfoSphere MDM, and TIBCO EBX.
What separates a working deployment from shelfware is not the feature list. It is whether the tool enforces where your failure happens: the write path and the stewardship queue for one, the query and the access request for the other, the distinction underneath context graph vs context store.
How do a data catalog and master data management work together?
Permalink to “How do a data catalog and master data management work together?”They are not converging into one product. The golden record stays where it is produced. What changed in 2026 is that it now has to be legible to something that cannot ask a follow-up question, and that relationship has a direction: master data management is becoming an input.
The evidence is third-party, not an Atlan assertion. On 17 June 2026 Gartner published How Master Data Management Boosts the Context Layer for AI Agents (document 8011469), whose own summary reads: “D&A leaders face AI initiative failures because agents lack reliable master data context, which often leads to amplified data quality issues. Embed master data management into your context layer to build a trusted foundation, which boosts agent accuracy and helps deliver AI agent success.” The title carries as much evidence as the sentence: an analyst firm applying the category word to this pairing, and prescribing embedding rather than choosing. According to Forrester (2026), Jayesh Chaurasia, Senior Analyst at Forrester Research, describes the same movement in a post on a master data management vendor repositioning around context, writing that “the next enterprise AI bottleneck isn’t model choice or orchestration but shared context.”
Three patterns carry the mechanics:
- Match rules become stable entity keys. One side contributes resolved identity, the other the map of where it appears, and a cross-system join becomes a join rather than a guess.
- The catalog wraps the golden record in evidence. Lineage, ownership, freshness, and certification state are what let a reader tell the golden record is the golden record.
- The agent query path needs both at once. Authority from one, permission and currency from the other, which is why the core components of a context layer list both, and the argument what a context layer is makes from the other end.
Integration is not convergence. Convergence would mean one product doing both jobs; these patterns describe two products with a contract between them, which is also why context layer vs knowledge graph is worth drawing carefully. The write-time discipline keeps its owner, its software, and its exception queue, and the read-time layer stops pretending it can adjudicate.
Why an AI agent hits both failures in the same query
Permalink to “Why an AI agent hits both failures in the same query”A single agent question touches the authority of a record and its discoverability at once, which is why the 2026 answer to this comparison differs from the 2023 answer. Every other page treats agents as a closing paragraph. They are the reason the line got sharper, not blurrier.
Take one request: show me all customer churn risk. Answering it means joining CRM account data with support ticket history and billing events, and no single system holds that mapping. Without identity resolution the agent returns partial answers or silently joins on the wrong keys, the failure agent context layer work exists to close. “Silently” is load-bearing: a wrong-key join produces a confident number and no error. The agent needs both which customer record is authoritative, which master data management produces, and which table holds churn signals, how current it is, and whether it may read it, which a catalog produces. Each discipline was designed to answer only the other’s question, the problem behind a knowledge graph for AI agents.
The stakes are documented. According to Gartner (2025), through 2026 organizations will abandon 60% of AI projects unsupported by AI-ready data, and over 40% of agentic AI projects will be canceled by the end of 2027 on escalating costs, unclear value, or inadequate risk controls. Rita Sallam, Distinguished VP Analyst at Gartner, put the mechanism plainly: “Agentic AI outcomes depend on context including semantic representations of data.” That puts AI agent accuracy and a semantic layer for AI agents on one critical path. Agents do not change which failure you have. They raise the cost of guessing, because a human who cannot tell which table is current asks somebody, and an agent answers anyway. Ask for a supplier’s total exposure, or which products underperform in EMEA, and the same pair appears: one supplier identity rather than three regional variants, one SKU definition across both ERPs, and on the read side, which contract or sales table is certified rather than superseded.
Where Atlan sits, and what it does not do
Permalink to “Where Atlan sits, and what it does not do”Atlan does not create golden records, does not run match and merge, and does not approve which version of a record is authoritative. Those are write-time jobs owned by a master data management platform and the stewards who operate it, and claiming them would be the overclaim this page argues against everywhere else.
What Atlan is, is the Context Layer for AI. The golden record stays in whichever system produces it; Atlan carries the lineage, ownership, freshness, and policy context that lets a person or an agent know which record to trust and whether they may use it. The Enterprise Data Graph holds typed relationships and cross-system identity mapping across 80+ sources, so a query traverses real connections rather than inferring them.
The rest of the surface aims at the read-time failure. Context Agents and the Context Engineering Studio are where teams author and test the definitions agents will use, the Context Lakehouse is the Iceberg-native storage layer underneath, and the MCP Server is the agent-facing interface, so context arrives through a protocol rather than a scrape. Column-level lineage, certification state, and policy evaluated at traversal time are what turn an inventory into something an agent can act on, work that how to implement an enterprise context layer for AI and context engineering for AI agents sequence.
The boundary is the useful part. If your failure is write time, this is not the software you need: a layer above a record nobody has adjudicated will faithfully deliver whichever version it found.
Real stories from real customers: Context at enterprise scale
Permalink to “Real stories from real customers: Context at enterprise scale”Neither story below is a master data management story, and neither customer is described here as having replaced or added master data management. They are evidence of the read-time problem at a scale where relationships stop working.
"AI initiatives require more context than ever. Atlan's metadata lakehouse is configurable, intuitive, and able to scale to hundreds of millions of assets. As we're doing this, we're making life easier for data scientists and speeding up innovation."
— Andrew Reiskind, Chief Data Officer, Mastercard
"Atlan captures Workday's shared language to be leveraged by AI via its MCP server. As part of Atlan's AI labs, we're co-building the semantic layer that AI needs."
— Joe DosSantos, VP Enterprise Data & Analytics, Workday
AI Agent Context Readiness Checklist
Score whether your estate can hand an agent trustworthy context, including the cross-system identity question behind most wrong-join failures.
Take the ChecklistWhat nobody in this comparison can tell you yet
Permalink to “What nobody in this comparison can tell you yet”The decision is not which category is better. It is which failure you have, and that failure has a location in the pipe: before the record commits, or after it exists. Two things are genuinely unresolved.
First, nobody publishes how many enterprises run both disciplines. No analyst firm reports it and no survey establishes it, so if you want to know whether running both is the norm or the exception, the public record does not say and any figure you are quoted is invented.
Second, the line this page is drawn on is moving. Peeters, Steiner and Bizer found the best large language models need no or only a few training examples to match models fine-tuned on thousands, and they hold up better on entities they have not seen; later work extends it, from fine-tuning for entity matching to up to 150% higher accuracy and up to five times fewer API calls from in-context clustering. If matching gets cheap enough, the economics of a dedicated write-time platform change. The counterweight is that extraction is the cheap part, and knowledge graph construction for AI is where the expensive part shows up: survivorship rules, stewardship approval, and write-back are not model problems at all. Which half wins is open, and anyone who tells you the answer is selling one of the two things.
FAQs about data catalogs vs master data management
Permalink to “FAQs about data catalogs vs master data management”1. What is the difference between a data catalog and master data management?
Permalink to “1. What is the difference between a data catalog and master data management?”A data catalog inventories data assets so people and agents can find, date, and clear them for use. Master data management decides which version of a shared entity, a customer or a supplier, is authoritative across the systems that author it. One works at read time, the other at write time.
2. Is master data management still relevant in 2026?
Permalink to “2. Is master data management still relevant in 2026?”Yes. Gartner retired the Magic Quadrant for Master Data Management Solutions in 2022 and published a new edition on 6 April 2026 evaluating roughly 20 vendors. An analyst firm does not un-retire a Magic Quadrant for a dead category. What changed is the job: the golden record now has to be legible to software.
3. Which should you implement first, a data catalog or master data management?
Permalink to “3. Which should you implement first, a data catalog or master data management?”Whichever matches the failure you can name. If the same entity is created in more than one system and somebody has to rule on authority before a transaction commits, master data management goes first. If the records are correct and nobody can find, date, or clear them, the catalog goes first.
4. Do you need both a data catalog and master data management?
Permalink to “4. Do you need both a data catalog and master data management?”Only if you have two pluralities. Plural writers, the same entity authored into several operational systems, is what master data management is for. Plural readers, more people and agents than can be briefed in person, is what a catalog is for. Many enterprises have one and not the other.
5. Is master data management part of data governance?
Permalink to “5. Is master data management part of data governance?”Master data management is usually run as a governance program and staffed by data stewards, but it is a distinct discipline with its own software category, Magic Quadrant, and deliverable. Governance sets the policy for who decides; master data management applies it to a record before it commits.
6. What are the four types of master data management?
Permalink to “6. What are the four types of master data management?”The four commonly cited implementation styles are registry, consolidation, coexistence, and centralized or transactional hub. Registry indexes records in place. Consolidation builds a golden record for reporting only. Coexistence writes it back to source systems. A centralized hub authors the master record directly, the only style enforcing authority pre-commit.
7. Is master data management (MDM) an ETL tool?
Permalink to “7. Is master data management (MDM) an ETL tool?”No. These platforms move data and ship integration connectors, but the work that defines them is adjudication: matching records that refer to the same entity, merging them under survivorship rules, and routing exceptions to a steward. ETL reshapes data without ruling on which version wins.
8. Can AI agents do entity resolution instead of a master data management platform?
Permalink to “8. Can AI agents do entity resolution instead of a master data management platform?”Partly, and the gap is narrowing. Peeters, Steiner and Bizer found the best large language models need no or few training examples to match models fine-tuned on thousands, and hold up better on unseen entities. What models do not supply is the operational half: survivorship rules, stewardship approval, and write-back.
Sources
Permalink to “Sources”- How Master Data Management Boosts the Context Layer for AI Agents (doc 8011469), Gartner
- Lack of AI-Ready Data Puts AI Projects at Risk, Gartner
- Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, Gartner
- Lack of Semantics Causes Inaccurate AI Agents and Wasted Spending, Gartner
- The True Cost of Poor Data Quality, IBM Institute for Business Value
- Context, Not Models, Is The Real AI Bottleneck, Forrester
- Entity Matching using Large Language Models, arXiv
- Fine-tuning Large Language Models for Entity Matching, arXiv
- In-context Clustering-based Entity Resolution with Large Language Models, arXiv
- Master Data Management Tools, OvalEdge
- Why prioritize MDM over data catalogs, Malcolm Hawker on LinkedIn
- Databricks MDM / Data governance, r/databricks