---
title: "Context Lakehouse - The World's First Context Store Built for AI"
url: "https://atlan.com/context-lakehouse/"
description: "The Context Lakehouse is the only knowledge architecture built for a world where AI is both the primary producer and consumer of context. Iceberg-native, open formats, MCP, A2A, SQL access."
keywords: "context lakehouse, ai context, iceberg, metadata lakehouse, atlan"
---

> Atlan is hosting Context Conference, bringing together the leaders and builders at the frontier of giving AI the context it needs to understand their business. It runs online on October 28, 2026, from 11:00 AM to 2:00 PM ET. Atlan co-founder Prukalpa Sankar opens and closes the day. Leaders from AstraZeneca, BNY and Verizon share why they invest in context and what they get from it. Registrants get early access to The AI Context Gap, a new study from MIT Technology Review Insights. Register: https://atlan.com/context-conference/

**Context Lakehouse: the world's first context store engineered natively for AI.** The Context Lakehouse is the only knowledge architecture built for a world where AI is both the primary producer and consumer of context.

[Book a demo](https://atlan.com/forms/talk-to-sales-contact/) / [Watch it live](https://atlan.com/activate/)

## Architecture at a glance

Consumers on top of the Context Lakehouse:

- **Custom agents:** Google, Snowflake, Databricks Notebook
- **Vertical agents:** Decagon, Sierra, Writer
- **General purpose agents:** Claude, OpenAI
- **Analytics engines:** Snowflake, Databricks, Apache Spark

The Context Lakehouse layers:

1. **Activation:** SDKs, Open APIs, MCP, A2A, Webhooks & Alerts, App Framework, Orchestration Engine, Design System
2. **Context intelligence:** Query Parser, Hybrid Search (keyword and semantic), Knowledge Graph, Vector Search; runs on your compute (Google Cloud, Databricks, Snowflake) and your models or LLMs (OpenAI, Anthropic, Snowflake)
3. **Iceberg-native metadata store:** Open & Extensible, Version-Controlled, Decentralized Compute, Technical REST Catalog, Time-Travel & Auditable, Object Registry; lives in your lake (Google Cloud, Databricks, Snowflake)
4. **Enterprise-ready foundation:** Single Tenant, Cloud Agnostic, Multi-Region Redundancy, RBAC / ABAC, Disaster Recovery, Audit Trails

## Trusted by AI-forward enterprises

- "AI initiatives require more context than ever. Atlan's metadata lakehouse is configurable, intuitive, and able to scale to hundreds of millions of assets." - Andrew Reiskind, Chief Data Officer, Mastercard
- "Atlan is our context operating system to cover every type of context in every system including our operational systems. For the first time we have a single source of truth for context." - Sridher Arumugham, Chief Data Analytics Officer, DigiKey
- "We have Atlan as the metadata plane across our tech stack, independent of where technology is. We have one place to define the data, understand what it means, and where it comes from." - Oliver Gomes, VP Analytics & Strategy, FOX
- "With Atlan we cataloged over 18 million assets and 1,300+ glossary terms in our first year, so teams can trust and reuse context across the exchange." - Kiran Panja, Managing Director, Cloud & Data Engineering, CME Group

## The problem: analytics needed data infrastructure, AI needs context infrastructure

The data world built lakes, warehouses, and pipelines to tame data at scale. AI demands the same architectural investment, but for context.

- **Context at machine speed.** Data needed fast pipelines. AI needs fast context. Agents query definitions, policies, and relationships thousands of times per hour; infrastructure built for human-speed access compounds into minutes of latency per run.
- **Context is big data.** Data grew until you needed a data lake. Context is growing the same way. Every agent interaction creates observations, quality signals, usage patterns. Infrastructure that only reads collapses when agents write back at scale.
- **Context is relational.** Data needed schemas. AI needs graphs. Agents traverse lineage chains, cross domain boundaries, and check governance policies in a single call. A flat index can't support that. You need a traversable map of the entire estate.
- **Context needs versioning.** Data needed lineage. Context needs time travel. When an agent makes a mistake, you need to reconstruct exactly what context it saw and when. Without it, you can observe the output but never explain the reasoning.

## How it works: one open store, every protocol AI agents speak

### Architecture - built for agents that read context, write it back, and traverse it at scale

A knowledge graph for relationships and meaning. Iceberg-native file storage for scale and portability. Together, they give AI agents the richest possible context in the most open possible format.

- **Knowledge graph.** Everything an agent needs to understand an asset (domains, lineage, policies, quality scores) connected in a single traversable graph.
- **Iceberg-native storage.** Apache Iceberg underneath: ACID transactions, schema evolution, time travel, and standard SQL from any compatible engine. Open and portable by default.
- **Vector-native AI search.** Every asset stored with its vector embedding, so agents retrieve context by meaning, not keywords, across billions of assets.
- **Time travel.** Every historical state is queryable, enabling point-in-time retrieval for compliance and full audit trails for GDPR, CCPA, and SOX.

### Agent access - speaks every protocol AI agents use and every protocol humans already know

From MCP for governed queries to SQL for analytics, Context Lakehouse meets your stack where it is. Works with Claude, Cursor, Windsurf, or any MCP client: context before action.

- **MCP (Model Context Protocol).** What an asset means, whether it's trusted, and what policies apply, checked in a single call before anything executes.
- **A2A (Agent-to-Agent).** Agents write observations, quality signals, and usage patterns back into the store; context compounds with every interaction.
- **SQL.** Standard SQL over Iceberg: query metadata exactly as you query your data, with every engine your team already runs.
- **REST APIs & Graph APIs.** Full programmatic access to the graph: build integrations, automate workflows, or pipe context into model training pipelines. Open by design.

## Industry recognition: validated by Gartner and Forrester

- **Leader in the 2026 Gartner Magic Quadrant for Data & Analytics Governance:** "The core architecture of Atlan's 'metadata lakehouse' is built on Apache Iceberg's open table format, which allows for scalable, unified metadata cataloging for structured, semistructured and unstructured data. This allows for trusted, query-and auditable metadata for automation, event-driven and real-time policy enforcement across diverse scenarios and systems." [Report](https://atlan.com/gartner-magic-quadrant-data-governance-2026/)
- **Leader in the 2025 Gartner Magic Quadrant for Metadata Management Solutions:** "Atlan's Metadata Lakehouse forms the core foundation, built on an open and highly performant architecture. It is designed to be Iceberg native and includes a knowledge graph and event stream engine." [Report](https://atlan.com/gartner-magic-quadrant-metadata-management-solutions-2025/)
- **Leader in The Forrester Wave: Data Governance Solutions, Q3 2025:** "Its knowledge graph and AI-powered automation support clear data ownership, surfacing policy-relevant context and automating governance workflows centered on a metadata lakehouse." [Report](https://atlan.com/know/forrester-wave-data-governance-2025/)

## Learn more about context infrastructure

- [The enterprise context layer - 53+ resources](https://atlan.com/know/enterprise-context-layer/) - what the context layer is, why AI agents need it, how to implement it, and why teams that have it are 5x more likely to reach production.
- [Context graphs: $1T opportunity. Four positions. Zero consensus.](https://atlan.com/great-data-debate-2026/context-graphs/) - Bob Muglia, Karthik Ravindran (Microsoft), Tony Gentilcore (Glean), and Prukalpa Sankar debate who should own the context graph.
- [84% invest in AI. 17% reach production. Here's the gap.](https://atlan.com/resources/ai-broke-the-data-stack-predictions-for-2026/) - 550+ data leaders on the 7 structural shifts forcing data teams to rebuild for an AI-first world.
- [Mastercard CDO on context by design](https://atlan.com/regovern-watch-center/mastercard-context-by-design/) - Andrew Reiskind on why AI initiatives require more context than ever, and how Atlan's architecture scales to hundreds of millions of assets.
- [Metadata lakehouse vs. data catalog](https://atlan.com/know/metadata-lakehouse-vs-data-catalog/) - a catalog stores metadata for humans to browse; a metadata lakehouse is active infrastructure: Iceberg-native, bidirectional, traversable at machine speed.

## FAQ

**What is the Context Lakehouse?**
Atlan's knowledge architecture for storing, managing, and serving the context AI agents need to operate accurately at enterprise scale. It combines a knowledge graph for relationships and meaning, Iceberg-native file storage for portability and ACID guarantees, vector-native search for semantic retrieval, and full time-travel for compliance and audit. It is the store that every Atlan product reads from and writes to, and that any external agent can access via MCP, A2A, SQL, or API.

**How is Context Lakehouse different from a data catalog?**
A data catalog stores metadata for humans to browse. The Context Lakehouse is an active knowledge architecture designed for machine-speed access. Iceberg-native storage means context is queryable with standard SQL from any engine. The knowledge graph means relationships are traversable at depth in under 100ms. Bidirectional writes mean agents improve context on every interaction. Vector-native search means retrieval is by meaning, not search bar. A catalog is a directory. A Context Lakehouse is infrastructure.

**What does "Iceberg-native" mean and why does it matter?**
The Context Lakehouse stores all metadata in Apache Iceberg table format, the same open standard your best data already lives in. This gives you ACID transaction guarantees, schema evolution without breaking consumers, time travel for any historical state, and compatibility with every SQL engine your team already runs (Spark, Trino, DuckDB, Snowflake, BigQuery, Flink). Your context is stored in open formats you own and can query independently of Atlan. Your context is your IP; Iceberg-native ensures you can always access it.

**Which AI protocols does Context Lakehouse support?**
Four, natively: MCP (Model Context Protocol) for governed, trust-checked context delivery to any AI agent; A2A (Agent-to-Agent) for bidirectional writes where agents post quality signals and observations back; SQL via Apache Iceberg for programmatic access from any compatible engine; and REST and Graph APIs with SDKs in Python, Java, Node.js, and Go for custom integrations.

**How does time travel support compliance requirements like GDPR, CCPA, and SOX?**
Because Context Lakehouse is built on Apache Iceberg, every version of every asset state is automatically preserved and queryable. GDPR: prove exactly what data classification applied to an asset on any past date, and demonstrate that deletion policies were enforced. CCPA: access and deletion audit trails are built into the storage layer. SOX: every change to any financial data asset (who changed it, when, and what the previous state was) is queryable as a table. Compliance is a query, not a manual process.

**Can I use my own storage infrastructure (BYOC)?**
Yes. Context Lakehouse is designed around open formats and bring-your-own-compute (BYOC) principles. Because context is stored in Apache Iceberg, it can live in your own cloud storage (S3, GCS, ADLS) and be queried by your own compute engines. You are not locked into Atlan's infrastructure. Your context files are portable, owned by you, and readable by any Iceberg-compatible tool.

Build AI on infrastructure that's built for AI: [Book a demo](https://atlan.com/forms/talk-to-sales-contact/)