Skip to main content

DataHub Data Contracts: How They Work and Where They Fit

Ayswarrya G, Contributing Writer, Atlan
Contributing Writer, Data Engineering & Metadata
Updated:
|
Published:
12 min read

Key takeaways

  • Contracts in DataHub state expectations. Enforcement depends on assertion runners, alerts and pipeline hooks.
  • Existing dbt tests and Great Expectations checkpoints can feed DataHub as the assertions behind a contract.
  • Contracts attach to datasets only, one per dataset, and deleting one leaves its assertions in place.
  • Native freshness, volume and schema monitors are DataHub Cloud features, not open-source ones.

What are DataHub data contracts?

DataHub data contracts are producer-owned agreements that promise consumers a specific level of quality for a physical dataset. Each contract groups verifiable assertions into three categories: schema, freshness and data quality. If any assertion in the contract fails, the whole contract is marked failing. Contracts are verifiable rather than declarative, because they sit on checks that run against the physical data rather than on metadata. One contract attaches to one dataset, owned by its producer, and a consumer reads a single contract status instead of scanning dozens of individual tests.

What defines a DataHub contract:

  • Verifiable checks schema, column-level data checks and operational SLAs, not metadata.
  • Assertions tests that run against the physical asset: schema, freshness, volume, column and custom.
  • Producer ownership one contract per dataset, owned by the team that produces it.

Where contracts meet meaning:


A data contract is a promise, and a promise is only as good as whatever checks it. DataHub gets the first half right: producers declare what consumers can rely on, and the declaration is backed by tests that run against real data rather than a wiki page. The second half is where teams get surprised. Atlan approaches the same problem from the consumer side, attaching YAML contracts to the assets they govern, running quality rules natively in Snowflake, Databricks and BigQuery, and serving the result to agents over MCP alongside the definitions and lineage a pass-or-fail status cannot carry on its own.

Score Your Governance Readiness


Answer across five dimensions. It returns a maturity band and a sequenced 90-day plan from your score. Read the skill.

Paste into a new chat

Use the skill at https://atlan.com/skills/governance-readiness-score.md to score your governance readiness. Ask me for whatever it needs.

Run once in a terminal

curl -fsSL --create-dirs \
  -o ~/.agents/skills/governance-readiness-score/SKILL.md \
  https://atlan.com/skills/governance-readiness-score.md

For an agent

curl -fsSL https://atlan.com/skills/governance-readiness-score.md


How do DataHub data contracts improve data quality?

DataHub’s data contracts documentation frames a contract as a public producer commitment, which shifts quality from an implicit expectation to a verifiable guarantee. If a pipeline violates an active contract, it can trigger alerts or block the flow entirely, so broken data never reaches dashboards or models.

The practical wins are narrow but real. DataHub puts an SLA on freshness, clarifies ownership, reuses tests you already maintain by ingesting dbt test results, and covers checks on the physical data such as schema and column values, one slice of the metadata an agent actually needs to answer a question. Rather than scanning dozens of individual tests, a consumer reads one contract status on the dataset profile. That consolidation is the point: fewer signals, each one meaning something. It is the same instinct that makes data contracts for AI worth writing.


How do DataHub data contracts work?

Contracts operate on a foundation of verifiable assertions. An assertion is a specific test: checking a schema structure, enforcing volume limits, validating staleness. The contract points to one dataset and to the assertions that define its guarantees. The lifecycle runs in four steps.

Step 1: Define the assertions


There are three ways to create them:

  1. Declare assertions in YAML and register them with the DataHub CLI for an infrastructure-as-code workflow.
  2. Create them programmatically through the APIs and SDKs.
  3. Build them manually in the DataHub UI.

Freshness, schema and data quality assertions must exist before you can build a contract from the interface, a stricter starting bar than most enterprise AI-readiness programmes set.

Edition matters. Native monitors such as freshness, volume and schema checks are DataHub Cloud features, so Core users usually source assertions from dbt tests or Great Expectations runs.

Step 2: Bundle assertions into a contract


In the UI you open the dataset profile, go to the Quality tab, select Data Contracts, and choose which assertions to include. Programmatically you can use GraphQL, the Python SDK, REST, YAML, or the Open Data Contract Standard.

Step 3: Run the assertions


A contract is only as current as its last assertion run, the same trap that lets an LLM knowledge base go stale without anyone noticing. There are two paths: schedule assertions inside DataHub Cloud Observe, or run checks in dbt, Great Expectations or another tool and publish results back through the API. The second path is where most open-source deployments live, and it is why contract freshness tends to track CI health rather than data health.

Step 4: Evaluate status and act on it


Each contract carries a state of ACTIVE or PENDING. Active contracts are meant to be enforced; pending contracts serve visibility and planning. DataHub’s entity documentation is explicit that contracts do not enforce themselves, so blocking a bad load means wiring checks into your pipelines and CI/CD, and closing that loop is the harder half of the work. The contract is a statement of intent with a status light attached.

Design limits to know early


A few design choices shape how far contracts go in a larger organization:

  • Datasets only: contracts attach to datasets, with support for data products and other entities listed as future work.
  • One contract per dataset: consumer-oriented contracts are not directly supported. Teams use tags, structured properties or custom workflows to track which assertions matter to each consumer, a workaround that gets expensive as the agent context layer grows past a handful of datasets.
  • Separate cleanup: deleting a contract leaves its assertions in place, so both have to be removed deliberately.

When is DataHub the right choice?

Built under Apache 2.0, DataHub appeals to organizations self-hosting their metadata stack or integrating deeply with Kafka-native, event-driven architectures. If your team prefers managing governance logic through GitOps, custom Python SDK scripts and integrations as code, the tooling is adequate, though the TCO of building, buying or bundling a context layer usually argues the other way at scale.

Contracts work well under specific conditions:

  • You have platform engineers to run it: self-hosting Core means operating Kubernetes, Kafka and Elasticsearch or OpenSearch, and owning every upgrade. The criteria for building or buying metadata tooling price that honestly.
  • Your checks already live in dbt or Great Expectations: contracts sit on top of tests you already maintain.
  • Your focus is technical and pipeline metadata: engineering teams tracking ingestion, schema change and impact analysis get strong value from the event-driven model, and can pair it with DataHub’s OpenLineage support on the lineage side.
  • Open source licensing is mandatory: Apache 2.0 meets strict OSS requirements, which is also why teams end up comparing it with OpenMetadata.
  • Producers and engineers are the audience: if business users and agents are not the primary consumers of contract status, the producer-oriented model is usually enough.

The model starts to strain once contract status has to travel beyond the data team. If business stewards need to act on quality signals, if contracts have to cover data products and semantic models rather than datasets alone, or if agents need quality tied to certified definitions and lineage at query time, a producer-oriented dataset contract is the starting point rather than the answer. Managing metadata and understanding what it means are different jobs, and the difference is sharper for agents than for people: a person seeing a failed check asks a colleague, and an agent does not, Treating context as a catalog concern rather than a pipeline one is what closes that gap.


How Atlan approaches quality and context governance

Contracts and quality checks tell you whether data met a threshold. Agents also need to know what the data means, who certified it, where it came from, and whether it is safe to use for a given question. Atlan treats data quality as one layer of the governed context serving humans and agents alike, which is why AI agents need an enterprise context layer and not just better tests.

Data contracts that live with the asset


Atlan lets producers and consumers write contracts once and keep them attached to the assets they govern, whether that is YAML templates, CI/CD and impact analysis, webhooks, or assets such as tables, views, materialized views and the output ports of data products. Coverage beyond the dataset is what starts to matter once data products enter the picture, and it is where metadata management for AI diverges from classic cataloguing.

Quality rules that run in your warehouse


Atlan’s Data Quality Studio executes rules with native compute, including Snowflake Data Metric Functions and Databricks and BigQuery stored procedures, the same push-down approach behind context for Snowflake and context for Databricks. Rules cover completeness, uniqueness, statistical checks, freshness, reconciliation and custom SQL. Trust signals in the form of warnings, trust scores and notifications appear wherever people find the asset, which stops quality becoming a dashboard nobody opens and keeps it inside the data catalog agents read. Data quality for AI agents raises the bar again, because poor data quality in LLMs surfaces as confident wrong answers rather than errors.

Governed context for AI agents


Quality results matter most when they reach the agent at the moment it answers. Context Engineering Studio bootstraps context from real SQL, pipelines and BI semantics, tests it with eval suites built from your own dashboards and queries, then ships it to runtimes over MCP. Evaluating context is not the same as testing data quality, and the Atlan MCP server is what carries the result to the agent.

Context Agents generate descriptions, metrics and quality rules so context keeps pace with change rather than aging into tribal knowledge. The Iceberg-native Context Lakehouse keeps context, lineage and quality history queryable with time travel for audits, so you can reconstruct what an agent saw when it answered. Those edges land in the Enterprise Data Graph next to lineage and the Active Ontology an agent reasons over.


Moving forward with DataHub data contracts

DataHub data contracts give producers a disciplined way to promise schema, freshness and quality on the datasets they own. For engineering-led teams already running dbt or Great Expectations, they add visibility with modest effort, and the reuse of existing tests is the cheapest governance win available.

Structural contracts alone are rarely enough once AI readiness enters the conversation. A contract tells an agent that a table passed its checks. It does not tell the agent which of two revenue definitions the table encodes, or that the definition changed in June, No assertion carries that business context. Connecting quality to meaning is a different layer of the stack, and it is where an enterprise context layer does the work a contract was never designed to do.

Start where you already have tests. The question worth answering before you scale contracts across an estate is who gets paged when one fails, and whether that person knows what the dataset is for.

Book a demo


FAQs about DataHub data contracts

1. What is a data contract in DataHub?


A data contract in DataHub is an agreement between a dataset’s producer and its consumers promising a certain level of data quality. It bundles selected assertions into schema, freshness and data quality categories, and the contract fails if any included assertion fails. DataHub’s documentation describes it as producer-oriented, with one contract per physical dataset.

2. What is the difference between Atlan and DataHub for contracts?


DataHub is an open source metadata platform under Apache 2.0 with a paid Cloud edition, built primarily for engineering and data platform teams. Atlan is a managed Context Layer for AI serving business users, stewards and AI agents alongside engineers. For contracts specifically, Atlan attaches YAML contracts to assets and enforces them with warehouse-native quality rules, while DataHub builds contracts from assertions and leaves enforcement to connected tools.

3. Is DataHub enough for ensuring the accuracy of AI agents?


DataHub data contracts can confirm that a dataset meets its schema, freshness and quality checks. Agent accuracy also depends on consistent business definitions, lineage, certification and access policies reaching the agent when it answers. Most teams pair quality checks with a governed context layer and pre-deployment evaluation to get reliable agent behavior.

4. What are assertions in DataHub data contracts?


Assertions are the specific data quality tests that form the building blocks of a contract. They evaluate whether a dataset complies with expected schemas, freshness SLAs, volume limits or column-level quality metrics. A contract points at one dataset and at the assertions that define its guarantees.

5. Can DataHub data contracts be automated?


Yes. Evaluation can be automated by scheduling assertions inside DataHub Cloud Observe, or by running checks in external tools such as dbt tests and Great Expectations within CI/CD workflows and publishing results back through the API.

6. Are data contracts available in DataHub open source?


Yes. DataHub’s Core versus Cloud comparison lists data contracts in both editions. Scheduled freshness, volume and schema monitoring, assertion notifications and pipeline circuit breakers are listed for DataHub Cloud only, so open source users typically run checks in dbt or Great Expectations and publish results back.

7. Do DataHub data contracts block bad data automatically?


No. Contracts define and track expectations, while enforcement depends on running assertions, configuring alerts and adding checks to pipelines or CI/CD. DataHub Cloud offers an API that pipelines can call as a circuit breaker before reading or writing data.


Sources

  1. DataHub Docs | Data contracts
  2. DataHub Docs | Assertions
  3. DataHub Docs | dbt ingestion source
  4. DataHub Docs | Great Expectations integration
  5. DataHub Docs | Open Data Contract Standard
  6. DataHub Docs | dataContract entity
  7. Atlan Docs | Data quality

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI. It translates business knowledge, including data definitions, working procedures, and governance policies, into context AI can actually use. This knowledge lives in a single Enterprise Data Graph that every team and AI agent can reach.

In Atlan's AI Labs benchmark, adding this context improved AI's text-to-SQL accuracy by 38%.

Atlan is recognized as a Leader across multiple Gartner reports and Forrester Waves, and is trusted by over 400 enterprises representing $10T+ in market cap, including Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, and Elastic.

Bridge the context gap.
Ship AI that works.