---
title: "Enterprise AI's 200-Millisecond Problem"
url: "https://atlan.com/context-and-chaos/issue/enterprise-ais-200-millisecond-problem/"
description: "Your model will take 20,000 tokens. Choosing which 20,000 is a structure problem, and the unit worth modeling is the decision trace."
keywords: "Context Engineering, Knowledge Graphs, Retrieval, AI Agents"
---

> Atlan is hosting Context Conference, bringing together the leaders and builders at the frontier of giving AI the context it needs to understand their business. It runs online on October 28, 2026, from 11:00 AM to 2:00 PM ET. Atlan co-founder Prukalpa Sankar opens and closes the day. Leaders from AstraZeneca, BNY and Verizon share why they invest in context and what they get from it. Registrants get early access to The AI Context Gap, a new study from MIT Technology Review Insights. Register: https://atlan.com/context-conference/

A Context & Chaos essay by Dan McCreary, published August 27, 2026 (13 min read). A model accepts roughly 20,000 tokens and the retrieval path has about 200 milliseconds to pick them; McCreary argues that choosing which 20,000 is a structure problem, that the unit worth modeling is the decision trace, and he adds "the least glamorous part of the story: the math."

**Author.** Dan McCreary (Author, Context Graph and Token Efficiency; [dmccreary.github.io](https://dmccreary.github.io/)) has spent three decades, since Neal Stephenson's The Diamond Age in 1995, on how knowledge moves through complex systems. He builds open agent skills that generate intelligent textbooks whose learning graphs and xAPI event streams predict what a student has mastered.

The piece builds on three earlier Context & Chaos essays: [Jessica Talisman](https://contextandchaos.substack.com/p/ontologies-context-graphs-and-semantic) (the industry confused measurement with meaning), [Prukalpa](https://contextandchaos.substack.com/p/what-an-enterprise-context-layer) (the context substrate is AI-ready data, semantics and ontology, and skills, built, tested, reviewed and versioned like software), and [Juha Korpela](https://contextandchaos.substack.com/p/conceptual-modeling-is-the-context) (conceptual modeling is what AI agents now need most).

## The 200-millisecond problem

- The answer to an employee's question is spread across systems never designed to be read together. The model will accept [roughly 20,000 tokens](https://arxiv.org/abs/2307.03172); retrieval has about 200 milliseconds to choose them.
- It is selection under two budgets, latency and cost. Bigger windows, more embeddings, and cleverer prompts answer adjacent questions; the real one is which 20,000 tokens, in what order, how fast.
- Wrong selection: confident hallucination. Right but four seconds: nobody uses it. Right, fast, and expensive: "your CFO learns your name."
- Biology analogy: the spinal cord pulls your hand off a hot stove over wiring laid down in advance. The graph is wiring inside a context layer, not the context layer itself.

## Why a graph and not a really good search box

- Vector search answers "what looks like this?" The business runs on "what is connected to this, how, who approved it, and what happened the last three times?", which returns a different and usually much smaller set.
- Structure lets you allocate the token budget: personalised PageRank seeded on the question's entities ranks which of forty related records bear on this query; centrality shows which definitions everything depends on. Retrieve a bounded subgraph (the decision, its precedents, its approvers, their immediate neighborhood) instead of the top fifty similar chunks. Bounded retrieval means a bounded token count.
- The boundary also answers context poisoning: inconsistent answers across runs usually come from retrieval with no principled edge, not the embedding model. Structure gives a stopping rule.
- Caveat, agreeing with [Andrew Lentz](https://contextandchaos.substack.com/p/your-agent-doesnt-need-to-walk-the): authoring context and runtime querying are separate decisions; for bounded questions the agent needs definitions, join keys, grain, and the governing rule, not a graph to walk. Everything here is the authoring side; the agent never touches the graph.
- Diagram: the graph (relationships, dependencies, structure) is distilled into small, relevant context that feeds the agent. "The graph is upstream wiring. The agent reads the distillation."

## The atomic unit is the decision trace

Not the document, table, or metric definition: the trace of what happened, why, who approved it, which precedent justified it, and what was true at the time. Example: the CRM knows a discount was 22 percent, but not that the VP approved it because the customer's plant flooded, that three similar exceptions were granted in 2024, or that approval happened in a hallway and was entered a week later by someone else.

Four layers of a real trace, which most systems capture none of:

- **Exception logic.** The rule and the conditions under which it was set aside; only the record of what was done predicts next time.
- **Historical precedent.** Prior decisions that made this one defensible; invisible unless decisions link to each other.
- **Cross-system synthesis.** Most decisions draw on three or four systems with no shared key or vocabulary; the synthesis happened in someone's head.
- **Out-of-band approval.** The conversation that authorized the exception and never touched a system of record.

Why incumbents struggle: warehouses sit on the read path after the fact, agent platforms on the execution path without persistence, CRM and ERP are optimized for the transaction. Each captures the what, none the why. Prukalpa's market version (January): in a heterogeneous enterprise the integrator beats the vertical application. The architectural version: the why has no owner yet, so whoever captures it accumulates an asset nobody can buy later.

## Why the structure does the work

Compression follows from how traces are shaped, not from graphs in general. Three modeling choices:

1. **Make the dependency graph acyclic.** If edges point from a concept to what it depends on, the minimal context is the target's direct and transitive dependencies, and a topological sort orders it so nothing appears before what it rests on. Dependency DAGs are explored in retrieval research but underused for enterprise context. Andrew Lentz reached the same point from cost: deep, narrow, acyclic graphs are cheap to walk; shallow, dense, cyclic ones are ruinous by depth three. Cycles are the expensive thing, not depth.
2. **Type your edges and mean it.** PREREQUISITE_OF, DEPENDS_ON, and PART_OF each support a different traversal; a generic RELATED_TO cannot be traversed selectively.
3. **Choose the right grain.** A domain resolves to roughly one hundred to six hundred core concepts, one classification each. Finer is noise; coarser stops being minimal.

The retrieval question becomes "what does this concept depend on?", which has a small, computable answer. Diagram: similarity retrieval pulls a dense tangle; dependency retrieval follows one chain (decision, rule, definition, precedent). "Smaller context. Same decision."

## What the numbers say

- The objective is useful reasoning per token. Proposed metric: **Reasoning Density Score = F1 divided by tokens** ("miles per gallon for context").
- Benchmark co-authored with Daniel Yarmoluk (still in preprint, "hold it loosely"): a clinical trial corpus of 2.68 million tokens encoded as a compact knowledge graph of 2,614 tokens, roughly thousandfold compression. The graph holds concepts, dependencies, and classifications, not the corpus; prose is discarded.

## The unglamorous plumbing

No new infrastructure is required; a context graph is assembly work on twenty years of half-done data management.

- [Semantic layers](https://dmccreary.github.io/context-graph/chapters/02-semantic-layers/): shared meaning over the lakehouse. Metadata management: what is trustworthy.
- [Metadata registries](https://dmccreary.github.io/context-graph/chapters/03-metadata-management/) and ISO 11179: precise, non-circular definitions that state permissible values, "the highest-compression token in the building." Most glossaries fail on the first entry.
- [Process mining, lineage and provenance](https://dmccreary.github.io/context-graph/chapters/07-process-mining-lineage/): reconstruct decision traces from event logs every workflow, ticketing, and approval system already emits.
- [Bitemporal modeling](https://dmccreary.github.io/context-graph/chapters/13-graph-data-modeling/): separates when something was true from when it was recorded, to answer "what did we believe on March 3rd."

Caution: Prukalpa argued in February that semantic layers failed as abstractions over ungoverned data. Graph structure over ungoverned content reproduces that failure with better response times. The plumbing is the graph's substance.

## What it costs, and where this is not the answer

The compact graph cost about a tenth of a cent per query to consult and considerably more to build (concept extraction, dependency modeling, expert review, upkeep). Compute the break-even first. Source: the [Benchmarking Token Costs Paper](https://github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf).

Where the structure does not pay:

- **Single-hop factual lookup** ("what is the current list price"): retrieval is fine and cheaper.
- **Corpora without real dependency structure** (support tickets, customer conversations, news): forced edges make a slower search index.
- **Content that churns faster than you can rebuild**: monthly concept turnover never amortizes the build.

At a few thousand tokens the compact graph is a build artifact, not a system: small, versioned, rebuilt on a schedule, read by the retrieval path, and never queried at execution time.

## The wiring

Everyone rents the same models; the record of why a company decided what it decided cannot be rented, and most organizations discard it daily. Start narrow: take one decision agents keep getting wrong (a pricing exception, an eligibility call, an approval), write out what a complete trace holds (rule, exception, precedent, person), and check which of the four your systems store. It is usually one. That gap is the work and the asset. "Measure the tokens."

## A book from Dan

Dan McCreary writes free, open-source textbooks on building with LLMs, including [Context Graph](https://dmccreary.github.io/context-graph/) and [Token Efficiency](https://dmccreary.github.io/token-efficiency/); both informed this piece.

All Context & Chaos issues: [https://atlan.com/context-and-chaos/](https://atlan.com/context-and-chaos/).