---
title: "Context Graphs as AI Evaluation Infrastructure"
url: "https://atlan.com/context-and-chaos/issue/context-graphs-as-ai-evaluation-infrastructure/"
description: "Enterprise AI evaluation fails not because of models, but because evaluation is decoupled from governed context. Introducing context drift and the evaluation graph."
keywords: "context graphs, AI evaluation, context drift, enterprise AI, data governance, evaluation graph, LLMOps"
---

> Atlan is hosting Context Conference, bringing together the leaders and builders at the frontier of giving AI the context it needs to understand their business. It runs online on October 28, 2026, from 11:00 AM to 2:00 PM ET. Atlan co-founder Prukalpa Sankar opens and closes the day. Leaders from AstraZeneca, BNY and Verizon share why they invest in context and what they get from it. Registrants get early access to The AI Context Gap, a new study from MIT Technology Review Insights. Register: https://atlan.com/context-conference/

A Context & Chaos contributor deep dive by Sandipan Bhaumik (Tech Leader - Data & AI), published April 9, 2026 (18 min read), subtitled "The Governance Layer Enterprise AI Has Been Missing". It argues enterprise AI evaluation fails not mainly because of models but because evaluation infrastructure is decoupled from the governed context that defines a correct answer at a point in time. It names this **context drift**, separates it from model failure and policy gaps, maps the four layers of enterprise context, and introduces the **evaluation graph**: a typed, versioned property graph linking evaluation runs to governed context snapshots, with a minimum viable schema. Audience: data architects, AI/ML engineers and governance practitioners in regulated enterprises, especially those who have seen a compliance team fail to reproduce an evaluation result.

## A failure nobody has named

Model drift has vocabulary, detection and remediation. Context drift has none: no standard detection mechanism, common schema or remediation path.

**Context drift** is when the definitions, policies, data sources and permissions around an AI system change and the evaluation infrastructure doesn't know. Model, prompt and output look the same, but the reality underneath has shifted silently.

Example: "active customer" in financial services has several definitions, owned by different teams, encoded in different systems, revised without announcement. A credit risk chatbot answers against some version of it. Which version, as of when, approved by whom? If it changed three months ago and the eval pipeline doesn't know, the system is measured against a reality that no longer exists. "The score didn't change. The context did."

Diagnose every failure as one of three problems with different owners, remedies and compliance implications: **model failure**, **context failure**, or **policy gap**. Blaming the model by default is how teams spend months retraining systems that were never broken.

## What context actually is

Prompts and logs show what was asked and answered, not which version of reality was in force. Caveat: most enterprise catalogs aren't mature enough to serve as evaluation infrastructure without preparation (partial coverage, inconsistent glossary entries, lineage that stops at ingestion). Propagating incomplete governance into evaluation makes the problem harder to see.

Four layers, each with a technical home:

| Layer | Where it lives | What it records |
|---|---|---|
| Data context | Catalog and lineage | Source tables or APIs consulted, point-in-time snapshot, certified freshness SLA, owning data product team; table lineage, column tags, owner, certification status |
| Semantic context | Business glossary | Precise definition of every business term used, its version, effective date range, approving steward, exceptions ("net revenue", "days past due", "eligible counterparty" are versioned artefacts). Missing: a foreign key into the evaluation log |
| Policy context | Access control and regulatory layer | Entitlements active at query time, data classifications, governing regime (GDPR, SR 11-7, DORA), approval workflow state. What risk teams care about; absent from standard LLMOps tooling |
| User context | Identity and entitlement layer | Who asked, role, permissions, organisational context. Same question, same definitions, different users can get different answers via access controls, residency or role filtering |

A standard evaluation log holds input, output, model version, latency and a score: a record of what happened, not of the conditions under which it happened.

## Two worlds that have never met

Governance teams think in assets (a noun world); AI teams think in runs (a verb world). Governance has spent years, in many large financial institutions more than a decade, building lineage graphs, glossaries, catalogs, classification tags and policy registries. AI teams are unaware it exists, from ontological incompatibility rather than negligence.

- **Ontology vs evaluation graph:** structurally similar, functionally different. An ontology defines what concepts mean across a domain (shared meaning). An evaluation graph records what was true for this system at this moment under these conditions (defensible evidence). It consumes ontological inputs; it is not an ontology.
- **Context graph vs evaluation graph:** a context graph is the living representation of enterprise context (meaning, ownership, rules) and the source of truth. The evaluation graph consumes it, recording which version was in force for a run: the audit trail.

The integration is technically straightforward. It hasn't happened because governance measures asset coverage and stewardship health while AI teams measure model performance and deployment velocity, with no shared reason to converge, until something goes wrong that compliance cannot ignore.

## Reproducibility is a political problem

Technical benefits (coverage, faster feedback, regression detection) don't unlock budget in regulated institutions. This does: without a versioned context snapshot attached to every evaluation run, a compliance team cannot legally defend the result.

- In banking, insurance, healthcare, or anything touching personal data or consequential decisions, reproducing an AI result is a regulatory requirement. "Same prompt, same model" is not the same context once definitions, policies or source-table ownership change.
- SR 11-7, the Federal Reserve's model risk management guidance (which, per the article, most UK banks mirror), requires validating models under the conditions in which they operate, including data definitions and policy constraints.
- A definitional snapshot is the only artefact that makes an eval result legally defensible. The Chief Risk Officer cares whether she can explain to a regulator, with evidence, what the system was doing and why.

## The evaluation graph: schema and wiring

A typed, versioned property graph linking an evaluation run to the governed context in force at execution. Not a replacement for your eval framework: a context provenance layer alongside it, connected by foreign keys. Minimum viable schema: six node types, five edge types.

**Nodes**

- **Task:** the question, instruction or scenario. `task_id`, `text`, `domain`, `criticality_tier`, `created_at`.
- **ContextSnapshot:** point-in-time capture of governed artefacts in force. `snapshot_id`, `captured_at`, `snapshot_hash` (deterministic hash of all referenced artefact versions, for exact reproducibility checks).
- **Run:** execution record. `run_id`, `model_version`, `retrieval_config`, `tool_manifest`, `prompt_template_version`, `execution_timestamp`; foreign key to your eval platform's run record.
- **Score:** `score_id`, `metric_name`, `value`, `evaluator_type` (human / LLM-as-judge / deterministic), `evaluation_timestamp`. A Run can have several.
- **GovernanceArtefact:** subtyped `GlossaryTerm`, `PolicyVersion`, `DatasetRegistration`, `AccessEntitlement`; always `artefact_id`, `version`, `effective_from`, `effective_to`, `owner_team`, `steward_id`, `source_system`.
- **UserSnapshot:** `user_id`, `role`, `active_entitlements`, `data_residency_region`, `permission_scope`, `captured_at`; records who was asking and what they could see.

**Edges**

- Run USES ContextSnapshot (exactly one per run).
- ContextSnapshot REFERENCES GovernanceArtefact (one or more versioned artefacts).
- Run REQUESTED_BY UserSnapshot.
- Task EVALUATED_BY Run (many runs per task: different models, dates).
- Run SCORED_AS Score (one or more).

Store it in any property graph database, or as lakehouse tables, which is simpler and sufficient for most enterprise evaluation volumes.

### Example: credit risk chatbot at a UK retail bank

A relationship manager asks: "Is this customer eligible for the bridging facility?"

- **Task:** tagged domain `credit_risk`, `criticality_tier` high.
- **ContextSnapshot** (captured 09:14:32 on the run date) references: GlossaryTerm "eligible_customer" v2.3 (effective since the previous quarter's model governance review, owned by the Credit Risk Data Office); PolicyVersion "bridging_facility_eligibility_v4" (active under the current FCA product governance framework, approved by the Credit Policy Committee); DatasetRegistration for the customer master dataset (gold tier, 24-hour freshness SLA, recertified 18 hours before the run).
- **UserSnapshot:** "credit_rm" role, entitlements scoped to retail banking, UK data residency. An investment banking colleague asking the same question would have different entitlements and possibly a different answer, all captured.
- **Run:** model version, retrieval config pointing to the vector store built on the certified dataset, prompt template version.

Six months later, with "eligible_customer" at v2.4, a regulator asks about a March decision: the `snapshot_hash` confirms v2.3 was in force, the steward is named, the approval record is traceable, and the user's entitlements are recorded.

Failure taxonomy, now with somewhere to land:

- Wrong despite correct context: **model failure** (retrain, fine-tune or swap).
- Correct against a stale or incorrect glossary term: **context failure**.
- Contextually correct but violating a policy the system had no record of: **policy gap** (incomplete wiring from the policy registry).
- Correct for one user's entitlements, wrong for another's: **user context failure**.

## What each team does next

- **Governance steward:** find the AI use case with the highest regulatory exposure, check whether its catalog entries have versioned definitions with effective dates, named stewards and lineage to source. Where they don't, that is where context drift enters; close it before wiring anything to evals.
- **LLMOps engineer:** add a ContextSnapshot capture as a pre-run hook on the highest-stakes eval suite. A JSON record of relevant glossary term versions, policy versions and user entitlements, written to a table keyed on `run_id`, is enough to start. The habit cannot be retrofitted; establish it from the first instrumented run.
- **Risk and compliance:** stop accepting evaluation reports with no context provenance. "Scored 0.87 on groundedness" without stating which definition versions it was computed against is not auditable. Ask for the snapshot; inability to produce one is itself a material finding.

## Where this breaks down

Pick the first domain where all three hold at once: highest regulatory exposure, most advanced catalog maturity, and eval tooling already instrumented. One of three yields a prototype that never leaves proof of concept; two of three exposes gaps faster than teams can fill them; three of three produces evidence.

Then: attach one evaluation run to one governance artefact (a single glossary term with real version history is enough); define the minimal schema for that run with the six node types; treat the first reproducible result (replaying a run against an archived snapshot with a stable score) as proof of principle, not of scale.

Out of scope: helpfulness remains subjective and user satisfaction can't be derived from a policy version. The evaluation graph solves reproducibility and auditability, not what "good" means to the person asking.

## Conclusion

Most large organisations already have both layers: governance's lineage graphs, glossaries, catalogs and policy registries, and AI teams' evaluation pipelines logging every run. Neither is connected to the other. The evaluation graph is the connection: not a new platform or transformation programme, but a context provenance layer answering what the system was operating against when it produced a result, and whether you can prove it. Context drift is occurring in every production AI system in large organisations today. "The organisations that build this before an incident will explain their architecture. The ones that build it after will explain their incident first."

The author calls the evaluation graph a framework in progress, not a finished standard, and invites practitioners from regulated industries to share experiences via [LinkedIn](https://www.linkedin.com/in/sandipanbhaumik/) or the [AgentBuild newsletter](https://agentbuild.substack.com). Views are the contributor's; Context and Chaos states it is information-first, with no promotions, paid or otherwise.

## About the author

Sandipan Bhaumik has spent almost two decades building data and AI foundations. Through AgentBuild Weekly he shares how builders and founders can move beyond AI hype to create agentic systems that think, adapt and work.

## Related reads

- [Metadata Weekly Is Now Context and Chaos](https://atlan.com/context-and-chaos/issue/metadata-weekly-is-now-context-and-chaos/) (April 2026)
- [BI-Ready Is Not AI-Ready](https://atlan.com/context-and-chaos/issue/bi-ready-is-not-ai-ready/) (March 2026)
- [Gartner D&A 2026: Where the Context Layer Became a Budget Line Item](https://atlan.com/context-and-chaos/issue/gartner-danda-2026-where-the-context-layer-became-a-budget-line-item/) (March 2026)
- [Data Governance vs AI Governance: Why It's the Wrong Battle](https://atlan.com/context-and-chaos/issue/data-governance-vs-ai-governance-why-its-the-wrong-battle/) (March 2026)

Context & Chaos is Atlan's community newsletter on context engineering, governance, architecture and discovery. Browse all articles at https://atlan.com/context-and-chaos/.