Skip to main content

What Is OpenMetadata Used For?

Emily Winks, Data Governance Expert, Atlan
Data Governance Expert
Updated:
|
Published:
19 min read

Key takeaways

  • OpenAI's internal data agent, Kepler, runs on OpenMetadata for 3,500+ internal users across 70,000 datasets.
  • OpenMetadata is free and Apache 2.0 licensed; Collate sells a managed version with a free tier.
  • OpenMetadata 2.0 shipped in August 2026, branded the open context layer for AI; 2.0.2 landed 16 September.
  • OpenMetadata now ships its own MCP server, but one catalog's graph is narrower than an estate's context.

What is OpenMetadata used for?

OpenMetadata is an open-source metadata platform, built and commercially backed by Collate, used for data discovery, lineage, quality testing, and governance-workflow automation across an organization's data estate. Its newest use case is serving cataloged metadata directly to AI agents through its own MCP server, most visibly in OpenAI's internal data agent, Kepler. OpenMetadata 2.0, shipped in August 2026, brands the project "the open context layer for AI", so the buying question is no longer whether it targets agents but how much of an estate one catalog's graph covers.

Where OpenMetadata is actually used today:

  • Discovery and lineage: search, browse, and trace data across every connected source
  • Data quality and governance: automated test suites and workflow-enforced data contracts
  • Agent access: a built-in MCP server exposes the catalog directly to AI agents
  • The scope question: OpenAI's Kepler runs on it, and one catalog still covers one slice of an estate

See if your catalog is AI-agent-ready

Check Agent Readiness

OpenMetadata is the open-source metadata platform that OpenAI’s internal data agent, Kepler, runs on to serve context across 70,000 datasets to more than 3,500 internal users, a build Atlan tracks closely as one read on what “AI-agent-ready” actually requires. Built by the team now running Collate, its commercial backer, OpenMetadata calls itself “The Open Context Layer for Data and AI” on its own GitHub page, shipped 2.0 in August 2026 under that banner, and ships an MCP server for direct agent access. This page covers what it’s actually used for: discovery, lineage, quality, governance, agent access, and where one catalog’s graph stops short of an estate’s context.

Context Conference session: Why invest in context? With Cindy Hoots, Leigh-Ann Russell and Prakash Kota. Virtual, Oct 28, 11 AM ET.
The conference for teams teaching AI their business
October 28 / 11 AM ET / Virtual. Join the leaders and builders defining the context layer, the missing infrastructure of the AI era.
Save Your Spot →

Most of what gets written about OpenMetadata treats it as a feature list: discovery, lineage, quality checks, governance workflows. Atlan’s own architecture walkthrough of OpenMetadata already covers that ground. The harder question is the one a team asks once an agent needs to touch the catalog: does cataloging your metadata make it usable by an agent, or just easier for a person to browse?

  • OpenAI’s own deployment is the sharpest current proof point. It built Kepler, its internal data agent, on OpenMetadata as the underlying catalog.
  • The project has moved onto context ground deliberately. OpenMetadata 2.0 shipped in August 2026 branded “the open context layer for AI”, and its case-study index now runs to 36 named companies with OpenAI leading.
  • The agent-facing layer is shipping, not announced. The MCP server exposes the catalog to agents without a person querying the UI first, and 2.0.2 hardened its tooling in September 2026.
  • Scope is the question a label does not settle. A catalog’s graph reaches the sources connected to it, which is a different surface from everything an agent needs context from.

Below: who built OpenMetadata, its sharpest real-world use case, how it works, the full taxonomy of what it’s used for, whether that adds up to AI-agent readiness, and how it compares to Collate.

Quick facts

Field Detail
What it is Open-source metadata platform, self-described on GitHub as “The Open Context Layer for Data and AI”
Built and backed by Collate, the commercial company founded by OpenMetadata’s creators, including Suresh Srinivas
License and community Apache 2.0, verified from the LICENSE file; 15,256 GitHub stars as of 19 September 2026
Current release 2.0 shipped 24 August 2026; 2.0.2, a maintenance release, on 16 September 2026
Connectors 130+ data connectors, per OpenMetadata’s own homepage
Headline proof point OpenAI’s internal data agent, Kepler, runs on it, serving 3,500+ internal users across 70,000 datasets
Core use cases Data discovery, lineage, data quality testing, governance-workflow automation, MCP-enabled agent access
Deployment model Self-hosted by default; Collate offers a managed service with a free tier

What is OpenMetadata, and who’s behind it?

OpenMetadata is an open-source metadata platform, built and commercially backed by Collate, that catalogs, connects, and increasingly serves as the layer behind data discovery, governance, and AI-agent access across an organization’s data estate. It’s an Apache 2.0-licensed metadata store with connectors, a business glossary, a lineage graph, and, more recently, its own MCP server, not just a search UI sitting on top of a database.

The project was created by the team that now runs Collate, led by Suresh Srinivas, Co-Founder and CEO of Collate. OpenMetadata states its own lineage plainly: “Built by the founders of Apache Hadoop, Atlas, and Uber Databook.”

OpenMetadata’s own GitHub repository carries 15,256 stars as of 19 September 2026, with releases shipping most weeks rather than monthly. Its README describes the project as “The Open Context Layer for Data and AI,” and its homepage now leads with “Introducing OpenMetadata 2.0, the open context layer for AI.” Two of its own first-party pages disagree on scale: the homepage cites “4,000+ enterprise deployments,” while its dedicated case-studies page cites “3,000+.” Treat both as directional evidence of a fast-growing project, not a precise count; the two pages haven’t been reconciled publicly.

For the full architecture walkthrough, Atlan’s dedicated deep-dive on OpenMetadata’s design covers how the components fit together. This page instead focuses on what they’re actually used for, starting with the single most concrete example available right now.


What is OpenMetadata used for right now?

The single most concrete, current answer is that OpenAI built its internal AI data agent, Kepler, on top of OpenMetadata as the underlying catalog. Collate’s own case study on the build puts numbers on it: 3,500+ internal users served by Kepler, a data platform spanning 15 tools, and 580+ petabytes processed daily across 70,000 datasets. Note the phrasing in the source is internal users, not employees, and the case study is adapted from a Summit '26 keynote.

That case study is the sharpest “used for” proof point available for OpenMetadata. Kepler gives OpenAI’s internal users a self-service way to ask questions of internal data without waiting on a data team, with OpenMetadata supplying the catalog, lineage, and glossary context the agent reasons over. Atlan’s analysis of what OpenAI’s data agent reveals about enterprise AI goes deeper into the six context layers behind that build.

A second, differently-scoped deployment: Gorgias cataloged more than 45,000 data assets and built 1,400+ dbt models daily across five projects, processing more than five petabytes daily, according to OpenMetadata’s own Gorgias case study. Antoine Balliet, Senior Data Engineer at Gorgias: “OpenMetadata offers us a strong layer of discovery. It saves our data analysts a lot of time and has helped us clean up our warehouse by identifying unused models and tables.”

One quote from a third deployment is worth sitting with. Angelita Frozza Sanches, Head of Core Data Platform at Scout24, told OpenMetadata’s own case-studies team: “We all talk about AI, but the real product is not AI. Context is the real product.” She was describing her experience with a catalog tool, not making a claim about agent-readiness specifically, but a practitioner reaching for that language unprompted is a data point worth noting as this page builds toward the same distinction.

The Kepler build specifically is analyzed further in Atlan’s piece on how to build an AI agent harness; the fuller argument about what separates a catalog from AI-agent-ready context comes later in this page.


How does OpenMetadata work?

Underneath the “used for” answers above, OpenMetadata works by ingesting metadata through connectors, storing it in a searchable graph alongside a business glossary and lineage, and now exposing that graph to agents through its own MCP server.

Ingestion runs on API-based connectors that pull technical metadata from databases, warehouses, BI tools, and pipelines. OpenMetadata publishes 130+ of them, and its 2.0.2 release in September 2026 was a maintenance release focused on connector reliability across Oracle, Databricks, Snowflake, BigQuery, Clickhouse, KafkaConnect, Grafana, and StarRocks. Once ingested, that metadata sits in a graph enriched by a business glossary, where human context, definitions, ownership, tags, attaches to technical assets. Data lineage is tracked at the column level, so a change upstream can be traced to everything it touches downstream.

The newest layer is the MCP server, a recent addition exposing cataloged metadata directly to AI agents instead of requiring a person to query the UI first. Atlan’s explainer on why MCP matters for AI agents covers the broader shift: catalogs built for human search are being retrofitted with agent-facing access, one MCP server at a time, and OpenMetadata’s is a real instance of that pattern, not a marketing claim.

How OpenMetadata moves from raw metadata to agent access Connectors API-based ingestion from sources Metadata graph + business glossary + column-level lineage MCP server new layer agent-facing query layer AI agents and human users Each stage is real and shipping. What it doesn't show: whether the metadata graph is unified with everything else an agent needs, or governed per agent and per user at the point context is requested, not just retrieved.

OpenMetadata's pipeline runs from connector to agent. What it evaluates and governs at each hop is the open question the rest of this page works through.

A well-organized pipeline answers “how does OpenMetadata work.” It does not answer how far the resulting graph reaches across an estate, which is the argument the next two sections take up.


Where does your context strategy actually stand?

Get the practical breakdown of what a real AI context stack needs beyond a catalog, in one brief.

Get the AI Context Stack

What are OpenMetadata’s core use cases?

OpenMetadata’s real-world use cases sort into five buckets, discovery, lineage, data quality, governance automation, and now agent access, and each one lands somewhere different on the distance between “cataloged” and “AI-agent-ready.”

  • Discovery: search and browse across every connected source, the use case most of OpenMetadata’s case studies describe first. Atlan’s own piece on enterprise search with AI covers what changes once that search interface has to serve an agent, not just a person typing a query.
  • Lineage: a column-level graph tracing where data comes from and where it goes, the same mechanism behind data observability for AI pipelines.
  • Data quality: automated test suites and alerting. OpenMetadata’s data-quality features have their own dedicated deep-dive already; Atlan’s complete guide to OpenMetadata data quality goes further than this page needs to.
  • Governance automation: data contracts and workflow enforcement, an actively maintained use case, not a legacy one. OpenMetadata’s release notes record steady work on this layer: 1.13.3 in July 2026, then the 2.0 line from August 2026, with 2.0.2 shipping 16 September 2026. Atlan’s own take on data contracts for AI and data mesh and AI frame why this use case matters beyond a single team’s workflow.
  • Agent and MCP access: the newest bucket, a live MCP server exposing the catalog to agents directly, covered in the previous section.

The table below is this page’s real payoff: not another feature list, but each use case measured against what AI-agent-ready context actually requires, the same lens Atlan’s metadata management for AI framework applies more generally.

Use case What OpenMetadata gives you out of the box What AI-agent-ready context still needs
Discovery Search and browse across connected sources Ranking and relevance tuned for agent queries, not just human search
Lineage Column-level lineage graph Lineage exposed in a form an agent can reason over at query time, not just a diagram a person reads
Data quality Automated test suites and alerting Quality signals attached to context an agent consumes before acting, not just a dashboard a person checks
Governance automation Data contracts, workflow enforcement Policy enforced at the point an agent requests context, per agent and per user, not just a workflow a person approves
Agent and MCP access A live MCP server exposing the catalog A semantics and policy layer unifying this catalog with everything else in the estate an agent also needs

Each row tells the same story: OpenMetadata gets a use case to “cataloged and queryable.” Getting the rest of the way to “an agent can safely act on this” is different work, made concrete in the next section. Teams cataloging both structured and unstructured data for AI, or connecting sources as part of a broader plan to prepare enterprise data for AI agents, run into this same gap once the catalog itself is done.


Does OpenMetadata make your data AI-agent-ready, or just easier to browse?

The project is moving hard at this question, and a buyer should read the movement rather than a snapshot. OpenMetadata 2.0 shipped in August 2026 under the banner “the open context layer for AI”. The case-studies index has grown to 36 named companies with OpenAI at the front, and the FDNY testimonial talks about connecting metadata to Claude. An argument that OpenMetadata is not aimed at agents has expired.

So the honest question is narrower, and it is about scope rather than intent.

The operational cost is real and worth naming first, because it is inherent rather than a shortcoming: self-hosted OpenMetadata means owning ingestion, uptime, upgrades, and any extension work yourself. OpenMetadata’s own answer to that is Collate, which runs the same technology as a managed service with a free tier. That tradeoff is a resourcing decision, not a product gap, and it should be priced as one.

Past the ops question sits the thing a label cannot settle. OpenMetadata gives you a well-organized graph over the sources you connect to it, and an MCP server to expose that graph. An agent working a real enterprise question often needs more than one system’s worth of context: the warehouse, the BI layer, the ticketing system, the documents, and the policies that say who may see which. Whether a catalog’s graph covers that set is an estate-by-estate answer, and it is the question worth asking about any vendor using context-layer language, including OpenMetadata and including Atlan. The related distinction, a queryable catalog versus governed context delivered per agent and per user, is where the evaluation should land.


Not sure how big your context gap is?

Run the calculator that scores the distance between a cataloged data estate and one an AI agent can safely act on.

Run the Context Gap Calculator

Is OpenMetadata free, and what’s the difference between OpenMetadata and Collate?

OpenMetadata itself is free and open source under Apache 2.0. Collate is the commercial company, founded by OpenMetadata’s own creators, that sells a managed version so teams don’t have to run the open-source deployment themselves.

Running OpenMetadata costs nothing in license fees; what costs money is the operational effort of running it yourself, or Collate’s managed service if a team would rather not. Collate markets that service as “Get started free” at cloud.getcollate.io, so the entry point is a free tier rather than a paywall. Suresh Srinivas, Co-Founder and CEO of Collate, frames the split directly in his own writing on why teams choose one path over the other: the open-source project stays open, and Collate exists for teams that would rather pay for someone else to operate it.

For a head-to-head against a different open-source catalog project, that’s a separate question this page doesn’t re-litigate here; Atlan’s dedicated breakdown covers that ground directly. For teams past the “should I self-host” question and ready to deploy, Atlan’s OpenMetadata setup guide picks up where this page leaves off. Either path still runs into the same agent-readiness question this page has been building toward: AI agents acting on a data catalog, or using a catalog as an LLM’s knowledge base, need more than either OpenMetadata or Collate ships on its own.


How Atlan approaches AI-agent-ready context

A catalog that’s easy to browse isn’t automatically a catalog an agent can safely act on. The gap between “cataloged” and “AI-agent-ready” from the previous section is exactly where OpenMetadata’s own use cases, real and mature as they are, stop. Technical metadata, business context, and the meaning behind terms usually live in different places, so an agent, or a person, has to assemble context by hand before it’s trustworthy.

Atlan’s Enterprise Data Graph unifies technical metadata, business context, and glossary meaning into one place people and agents both query directly. Policy context is enforced at the moment context is delivered, per agent, per user, rather than as a workflow a person signs off on separately. Lineage and certification flow into that live, queryable layer instead of sitting in documentation someone has to go find, and Atlan’s MCP server serves this governed context to agents across the estate, not just the slice one catalog covers. Atlan’s own take on semantic layers for AI agents and metadata management’s role in enterprise AI covers why that unification matters beyond any single tool.

Atlan doesn’t replace a catalog like OpenMetadata. It sits above one, turning browsable metadata into something an agent can act on safely, the shift context engineering and a properly built AI agent harness both depend on. Teams evaluating AI-ready data, or working through Context Layer 101 for the first time, land on the same question: cataloged, or ready for context-aware AI agents to use unsupervised? A catalog that never captures the tribal knowledge in people’s heads answers that honestly. Atlan’s approach to implementing an enterprise context layer for AI, and the distinction between a context graph and a knowledge graph, both pick up from where OpenMetadata’s catalog stops.


See the context layer in action

Watch how a governed context layer picks up where a self-hosted catalog stops.

Watch the Demo Series

Cataloged is not the same as AI-agent-ready

OpenMetadata’s real “used for” answer is wider than a feature list. Discovery, lineage, data quality, and governance automation are all genuinely mature, and its newest use case, serving cataloged metadata directly to AI agents through its own MCP server, is the sharpest current proof point available, anchored by OpenAI’s Kepler running on it for 3,500+ internal users.

Two things follow. Self-hosting it means owning the operational cost yourself, which Collate’s managed service exists to absorb. And whether OpenMetadata is good at cataloging isn’t the open question; its 36 named case studies answer that plainly enough.

The open question is scope. A catalog’s graph covers the sources connected to it, and an agent answering a real enterprise question usually reaches past that boundary into systems, documents, and policies the catalog never ingested. Closing that distance is the work a governed context layer does, and it is the question to put to every vendor now using the phrase, this one included.

Check your context maturity

Score where your catalog stands on the distance from "cataloged" to "an agent can safely act on this."

Take the Maturity Assessment

FAQs about OpenMetadata

1. Is OpenMetadata free?


Yes. OpenMetadata is open source under the Apache 2.0 license, verified from the LICENSE file at the repository root, so there’s no cost to download, deploy, or modify it. What isn’t free is the operational effort: you own ingestion, uptime, upgrades, and any custom extension work. Collate, the commercial company behind OpenMetadata, runs it as a managed service with a free tier at cloud.getcollate.io for teams that would rather not run that operation themselves.

2. What is Collate, and how is it different from OpenMetadata?


Collate is the commercial company founded by OpenMetadata’s creators, including Co-Founder and CEO Suresh Srinivas. OpenMetadata describes its own origins as “built by the founders of Apache Hadoop, Atlas, and Uber Databook”. OpenMetadata is the open-source project itself, free to self-host; Collate packages and manages that same technology as a hosted service with a free tier, handling the infrastructure, upgrades, and support a self-hosted deployment otherwise requires you to run yourself.

3. What are the main limitations of OpenMetadata?


The biggest real cost is operational, not functional: self-hosting means you own ingestion, uptime, upgrades, and extension work. That is inherent to self-hosting rather than a gap in the product, and Collate’s managed service is OpenMetadata’s own answer to it. The second consideration is scope: OpenMetadata’s graph covers the sources you connect to it, which is a narrower surface than the full set of systems an agent may need context from.

4. Does OpenMetadata support AI agents and MCP?


Yes. OpenMetadata ships its own MCP (Model Context Protocol) server, letting AI agents query cataloged metadata directly rather than requiring a person to search the UI first, and the 2.0.2 release of 16 September 2026 covers MCP tool consolidation and hardening. The clearest proof point is OpenAI’s internal data agent, Kepler, which runs on OpenMetadata as its underlying catalog for more than 3,500 internal users across 70,000 datasets.

5. Is OpenMetadata the same as a data catalog?


OpenMetadata includes a data catalog, search, browse, and discovery across connected sources, but it’s broader than that: it also handles lineage, data quality testing, glossary and business context, and governance-workflow automation. Its own GitHub README calls it “The Open Context Layer for Data and AI,” and its homepage leads with “Introducing OpenMetadata 2.0, the open context layer for AI”. Measure that against your own use cases: the label describes intent, and what matters for an agent is how many of your systems the graph actually reaches.

6. Who created OpenMetadata?


OpenMetadata was created by the team that now runs Collate, its commercial backer, led by Co-Founder and CEO Suresh Srinivas. OpenMetadata describes its own origins as “built by the founders of Apache Hadoop, Atlas, and Uber Databook”. The project is open source under Apache 2.0 and has 15,256 GitHub stars as of September 2026, with releases shipping most weeks rather than monthly.


Sources

  1. OpenMetadata GitHub Repository and README: “The Open Context Layer for Data and AI,” GitHub
  2. OpenMetadata LICENSE (Apache License 2.0), GitHub
  3. OpenMetadata Homepage, open-metadata.org
  4. How OpenAI Built a Self-Service AI Data Agent on OpenMetadata, Collate/OpenMetadata Case Study
  5. Gorgias Case Study: 45,000+ Assets, 5+ Petabytes Processed Daily, Collate/OpenMetadata
  6. OpenMetadata Case Studies and Testimonials Index, open-metadata.org
  7. OpenMetadata Release Notes, OpenMetadata Docs
  8. OpenMetadata Releases (2.0.0, 2.0.1, 2.0.2), GitHub
  9. Why OpenMetadata Is the Right Choice for You, OpenMetadata Blog (Suresh Srinivas)

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI. It translates business knowledge, including data definitions, working procedures, and governance policies, into context AI can actually use. This knowledge lives in a single Enterprise Data Graph that every team and AI agent can reach.

In Atlan's AI Labs benchmark, adding this context improved AI's text-to-SQL accuracy by 38%.

Atlan is recognized as a Leader across multiple Gartner reports and Forrester Waves, and is trusted by over 400 enterprises representing $10T+ in market cap, including Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, and Elastic.

Bridge the context gap.
Ship AI that works.