---
title: "The Context Problem Nobody's Fixing in Talk to Data Initiatives (NL > SQL)"
url: "https://atlan.com/context-and-chaos/issue/the-context-problem-nobodys-fixing-in-talk-to-data-nl-to-sql/"
description: "What testing NL-to-SQL on an enterprise-scale schema - with multi-layer validation and statistical significance - revealed about production readiness."
keywords: "NL-to-SQL, Context Engineering, AI Agents, Data Quality"
---

> Atlan is hosting Context Conference, bringing together the leaders and builders at the frontier of giving AI the context it needs to understand their business. It runs online on October 28, 2026, from 11:00 AM to 2:00 PM ET. Atlan co-founder Prukalpa Sankar opens and closes the day. Leaders from AstraZeneca, BNY and Verizon share why they invest in context and what they get from it. Registrants get early access to The AI Context Gap, a new study from MIT Technology Review Insights. Register: https://atlan.com/context-conference/

A Context & Chaos deep dive by Manoj Shanmugasundaram (Data Engineering Leader), published February 19, 2026 (10 min read). A controlled NL-to-SQL experiment (522 evaluated runs) finds that the gap between "talk to data" demos and production is context, not model intelligence, and that concise, machine-usable context beats verbose documentation.

About the author: Manoj Shanmugasundaram specializes in cloud and data solution architecture, designing systems that address complex business challenges. He writes about what production-scale deployments reveal about AI readiness and the context gaps that matter most.

## Metadata vs context

Teams connect an LLM to the warehouse, feed it schema, catalog descriptions and glossary terms, and the first demo works. In production, joins get messier, domain language creeps in, and accuracy falls off a cliff. The instinct is to blame the model, provider, prompt or lack of fine-tuning.

- **Metadata** is raw material: business glossaries, column descriptions, catalog entries.
- **Context** is metadata shaped into machine-usable signals: concise domain rules, SQL patterns, guardrails and contrasts a model can act on.

## LLMs are brilliant generalists with zero institutional knowledge

A model can see every table and column but doesn't know what "eliminated" means, that "churned" differs between marketing and finance, why three columns are named "date", or that some joins are technically valid and semantically disastrous. It fills gaps with plausible guesses, and fails with confident, convincing wrong answers.

## We tested one variable: context

Setup: a Formula One dataset with 13 tables, 94 columns and 174 unique natural-language questions, each evaluated three times (522 runs). Same dataset, model and questions; only the context layer changed.

| Condition | What it contained | Size | Result |
|---|---|---|---|
| Bare schema | Table names, column names, data types | ~29 lines | 16.1% correct (about one in six) |
| High-signal context | Schema plus concise business definitions, SQL usage patterns, domain rules | ~64 lines | 22.2% correct: a 38% relative improvement, significant at p below 0.0001 |
| Human documentation context | Same information as verbose, catalog-style narrative | ~176 lines | 13.8% worse than the concise version, at 52% higher cost from larger prompts |

22% absolute accuracy is not production-ready, but the direction and magnitude show where the leverage is. The surprise: more documentation made things worse.

## Why "more context" can become noise

Humans skim prose for the one sentence that matters; models can't. Narrative dilutes the signals that matter: which tables to join, what terms mean, which columns are easy to confuse, what defines "active" or "eliminated". What works is higher signal density, not more text.

## Example: the "eliminated" problem

Question: "Which F1 drivers were eliminated in the first round?" Nothing in the schema is labeled "eliminated".

- Without domain context the model wrote `SELECT driver_name FROM results WHERE position IS NULL;` (valid, semantically wrong).
- With the rule encoded (eliminated in round one means the slowest drivers in Q1) it wrote `SELECT driver_name FROM qualifying ORDER BY q1_time DESC LIMIT 5;`

One piece of institutional knowledge separated a plausible lie from a useful answer.

## Where context matters most

| Question complexity | Share of set | Lift from high-signal context |
|---|---|---|
| Simple (e.g. "average age of drivers?") | ~70% | ~1.15x; schema alone is usually enough |
| Medium (joins, aggregations, domain-specific definitions) | ~27% | **2.15x** |
| Very complex (nested logic, multi-step reasoning) | ~4% | ~1.00x; needs tool planning, chain-of-thought prompting or human review |

The biggest return is in the middle band, the everyday analytics questions. Start investing there.

## Four principles of context engineering

1. **Show usage, not just meaning.** Instead of "the millisecond column contains lap time measurements", give a runnable pattern, e.g. fastest lap: `SELECT milliseconds FROM lap_times ORDER BY milliseconds ASC LIMIT 1;`. Models learn more from one correct query than paragraphs of prose.
2. **Encode guardrails.** Capture how not to use a field: "never join events directly to users_raw", "status = 'active' is unreliable before 2021", "country here is billing, not shipping". Making such rules explicit flipped wrong answers to correct ones.
3. **Disambiguate lookalikes side by side.** Put confusable fields together with explicit contrast, e.g. drivers.nationality is the driver's home country; circuits.country is the race location. Contrast is context.
4. **Speak production language.** Use fully qualified names (analytics.f1.drivers, not drivers), real column names and snippets from actual production queries, so the model isn't guessing between two dialects.

## The economics

Richer context cost roughly $4.02 per 1,000 queries incrementally. A conservative ROI scenario assumed each correctly answered self-service query avoids about $0.50 in support or rework; returns were significant. The real expense is one confidently wrong number in a high-stakes decision. Run your own numbers.

## What we didn't test

One model, one domain, one dataset scale. No comparison against RAG with curated query examples, few-shot libraries or semantic layer integration. The best condition still failed roughly 80% of queries. Context is necessary, not sufficient: production NL-to-SQL still needs validation layers, execution safeguards and human oversight for high-stakes work.

## How to start without boiling the ocean

1. Pick one domain where errors matter (revenue, risk, growth, operations).
2. Collect real medium-complexity questions from query logs, tickets, notebooks and meetings; focus on joins and aggregations with domain nuance.
3. Rewrite context for the tables those questions touch: example queries, guardrails, contrast notes. Aim for signal density, not coverage.
4. Measure accuracy before and after, including which error types disappear. Scale what works.

## The mindset shift

Metadata used to help humans browse data; it now has a second job, helping machines reason safely. The highest-leverage investment may not be a bigger model or a better prompt but turning existing metadata into context machines can use, one domain and one set of questions at a time.

Reference: the findings draw on a published study, "Research Shows How Enhanced Metadata Delivers 38% Better AI Accuracy for Certain Queries": https://atlan.com/know/enhanced-metadata-improves-query-accuracy/

## The Insight Index (weekly digest links)

- [Why keeping Data & Analytics and AI in IT keeps killing ROI](https://periksson.substack.com/p/why-keeping-data-and-analytics-and) - Patrick Eriksson
- [The Principles of Data Modeling (Facts & Dimensions)](https://mahboubyassine.medium.com/the-principles-of-data-modeling-facts-dimensions-0aff9512eaf0) - Yassine Mahboub
- [The Data Job isn't Dying Because the Trust Problem is Exploding](https://opinionatedintelligence.substack.com/p/the-data-job-isnt-dying-because-the) - Julie Zhuo
- [The struggle for good AI governance is real](https://www.cio.com/article/4128980/the-struggle-for-good-ai-governance-is-real.html) - Grant Gross
- [How to Build Your First Team of AI Agents](https://drphilippahardman.substack.com/p/how-to-build-your-first-team-of-ai) - Dr. Phillipa Hardman
- [The Human Elements of the AI Foundations](https://metadataweekly.substack.com/p/the-human-elements-of-the-ai-foundations) - Gaurav Ramesh

## About Context & Chaos

A community newsletter where practitioners, builders and thinkers share lessons on context engineering, governance, architecture, discovery and the human side of data and AI work. Originally published in the [Context & Chaos newsletter on Substack](https://metadataweekly.substack.com/p/the-context-problem-nobodys-fixing). Pitch a piece: https://metadataweekly.substack.com/p/write-for-metadata-weekly

Related reads:

- [BI-Ready Is Not AI-Ready](https://atlan.com/context-and-chaos/issue/bi-ready-is-not-ai-ready/)
- [Ontologies, Context Graphs, and Semantic Layers: What AI Actually Needs in 2026](https://atlan.com/context-and-chaos/issue/ontologies-context-graphs-and-semantic-layers-what-ai-needs-in-2026/)
- [Conceptual Modeling Is the Context Engineering Nobody Is Doing](https://atlan.com/context-and-chaos/issue/conceptual-modeling-is-the-context-engineering-nobody-is-doing/)
- [Context Graphs as AI Evaluation Infrastructure](https://atlan.com/context-and-chaos/issue/context-graphs-as-ai-evaluation-infrastructure/)
- [Browse all Context & Chaos articles](https://atlan.com/context-and-chaos/)