Retrieval and metadata cover different halves of the same agent problem. Amazon Bedrock Knowledge Bases finds the right passages in your documents, while a data catalog records what your tables, columns, and metrics mean and who owns them. Atlan’s context layer for AI agents spans both, with an Enterprise Data Graph for shared definitions and lineage, and an MCP server that lets an agent reach that context beside its knowledge base calls.
| What it is | A comparison of Amazon Bedrock Knowledge Bases, AWS’s managed retrieval feature, with an external data catalog, which stores metadata about data assets |
| Key trade-off | Retrieval over content (Knowledge Bases) vs. meaning, ownership, and lineage over data assets (catalog) |
| Best for | Knowledge Bases: answering from documents. Catalog: answering about tables, metrics, and who owns them |
| Where they overlap | Knowledge Bases can read Amazon Redshift and the AWS Glue Data Catalog for structured questions |
| What neither covers | Organizational knowledge that sits outside the systems each one connects to |
What does Amazon Bedrock Knowledge Bases do?
Amazon Bedrock Knowledge Bases is the AWS feature that gives an application or agent retrieval-augmented generation over private content. According to AWS documentation (2026), it takes over the pipeline work of RAG: you connect a data source, and Bedrock parses documents, splits them into chunks, converts the chunks to embeddings, and writes them to a vector index while keeping a mapping to the original document. A deeper walk-through of types, deployment, and limits sits in our guide to Amazon Bedrock Knowledge Bases, and the underlying pattern is the one described in what retrieval-augmented generation is.
The user guide lists the unstructured sources it connects to: Amazon S3, Confluence, Google Drive, Microsoft OneDrive, Microsoft SharePoint, a web crawler, and a custom source. It also lists the vector stores behind it: Amazon OpenSearch Serverless, OpenSearch Service managed clusters, Amazon Neptune, Amazon Aurora, Pinecone, Redis Enterprise Cloud, MongoDB Atlas, and Amazon S3 Vectors. Multimodal documents with tables, charts, and diagrams are in scope, which separates it from a text-only index.
Retrieval itself comes through four API operations: Retrieve returns the most relevant chunks, RetrieveAndGenerate adds a generated answer with citations, GenerateQuery converts a natural language question into a query for a structured store, and AgenticRetrieveStream has a model break a complex question into sub-queries and retrieve iteratively. The advanced RAG techniques that teams add by hand, such as reranking and hybrid search, are the same levers this service exposes as options.
AWS now offers two flavors. The GA announcement (June 17, 2026) describes Amazon Bedrock Managed Knowledge Base, where AWS handles ingestion, storage, and retrieval, with six native data source connectors: Amazon S3, SharePoint, Confluence, Google Drive, OneDrive, and a web crawler. With a customer-managed knowledge base, you provision and maintain the vector database yourself, and the native connectors are S3 and a custom source. Our comparison of managed Knowledge Base against a custom RAG pipeline covers that trade-off in detail, and the cost side of custom retrieval is a separate question again.
Two details matter for this comparison. First, Knowledge Bases is not limited to unstructured text. AWS documents that it connects to structured data stores through the Amazon Redshift query engine, supports Amazon Redshift and the AWS Glue Data Catalog (through AWS Lake Formation) as structured sources, and converts natural language queries into SQL using query patterns, query history, and schema metadata. Second, a GraphRAG option with Amazon Neptune Analytics combines vector search with graph traversal across the relationships it finds between entities in your documents. The documentation lists limits: GraphRAG supports only Amazon S3 as a data source, and each data source holds up to 1,000 files by default, with a request path to raise that to 10,000. If you work with graphs, the relationship between a vector store and a graph database and the broader knowledge graph for AI agents explain what GraphRAG adds.
What does an external data catalog do?
An external data catalog is a system of record for metadata, the information about your data rather than the data itself. AWS describes its own example, the AWS Glue Data Catalog, as a centralized repository that stores metadata about data sets: an index to the location, schema, and runtime metrics of your data sources, organized into databases and tables. Crawlers scan data sources, infer schemas, and populate the catalog, which can connect to sources inside and outside AWS. The same documentation lists data lineage for auditing and provenance, and integration with AWS Lake Formation for fine-grained access control.
The part that matters for agents is business meaning. Per AWS, business context and semantic search are in preview in the Glue Data Catalog, letting you attach glossary terms, custom metadata fields, and skill assets to assets, and search them by meaning through the Glue Search API. During the preview, AWS notes there is no asset-level access control and that access runs through IAM. That is an honest signal of where the AWS-native catalog sits today: strong on technical metadata, with business context still arriving. Our look at where Glue stops and the split between DataZone and Glue go deeper on the AWS-native options.
A catalog built for AI consumers shifts emphasis from search screens to machine-readable meaning. That is the argument behind the data catalog for AI and the checklist for what makes data AI-ready: an agent reading a column named rev_adj cannot guess which of three revenue fields finance reports, so the definition has to live next to the asset. That layer of definitions is business context for AI, and it is the layer teams most often discover is missing.
How do Bedrock Knowledge Bases and an external data catalog differ?
The table below compares them on the dimensions that decide architecture, using the AWS-native Glue Data Catalog as the catalog example. Each row is sourced to AWS documentation linked in the Sources section.
| Area | Bedrock Knowledge Bases | External data catalog (Glue Data Catalog as example) |
|---|---|---|
| Core purpose | Retrieval for RAG: returns relevant chunks, or a generated answer with citations | Central repository of metadata: location, schema, and runtime metrics of data sources |
| What it works on | Unstructured and multimodal documents from S3, Confluence, SharePoint, and more; structured stores through Redshift and the Glue Data Catalog | Metadata about data assets such as databases and tables, populated by crawlers or defined manually |
| Backend | A vector index in a vector store, or Neptune Analytics for GraphRAG | Metadata tables organized into databases and tables |
| Business meaning | Relevance by semantic similarity between the query and chunks | Glossary terms and custom metadata fields (business context is in preview) |
| How it updates | You sync a data source to ingest additions, changes, and deletions; some sources support direct ingestion | Crawlers scan sources and extract metadata; tables can also be defined manually |
| Access control | Managed Knowledge Base supports ACL-aware retrieval that filters results by document permissions | Integrates with Lake Formation for fine-grained access control |
Read the table as two systems with different units of work. Knowledge Bases works at the level of a passage of content. A catalog works at the level of an asset and its description. Because Knowledge Bases can read the Glue Data Catalog itself, the two already overlap on AWS, and the Neptune versus data catalog question shows a similar boundary with graph stores. The practical conclusion is that neither replaces the other, which is why the next section looks at how they combine.
How do the two work together on AWS?
There are two combinations, and AWS documents both.
The first is inside Knowledge Bases. When you connect a structured store, Knowledge Bases uses schema metadata to convert questions into SQL, so the quality of that metadata sets the quality of the generated queries. Catalog metadata is not an optional extra here, because the SQL generator reads it. A table with unclear column names and no descriptions produces SQL that is syntactically fine and semantically wrong.
The second is at the agent boundary. AgentCore Gateway converts APIs, Lambda functions, and existing services into MCP-compatible tools, and its MCP targets operate in aggregation mode, combining the capabilities of every attached MCP target into one virtual MCP server. A managed knowledge base can be added as a gateway target, exposing Retrieve and AgenticRetrieveStream as MCP tools. Gateway also accepts external MCP servers as targets, with OAuth, IAM, or API key authorization. An agent behind that one endpoint can therefore retrieve passages and also query a catalog or context layer that speaks MCP. The protocol background is in what the Model Context Protocol is, the reasons it spread are in why MCP matters for AI agents, and the pattern of MCP-connected data catalogs shows what the catalog side looks like. If you plan the agent layer too, compare Bedrock Agents with LangGraph and read how Bedrock fits enterprise agents.
The same shape appears on other platforms. The question of how a runtime and a neutral context layer divide the work on Databricks has the same structure as retrieval and metadata on AWS, and the walkthrough on connecting an AI coding agent to a data warehouse over MCP shows the metadata-over-MCP half in a hands-on setting.
Where do both stop?
Each system reaches only what it connects to, and that boundary is where agents start guessing.
Connector coverage. A managed knowledge base ingests from six native sources, per the GA announcement, and a customer-managed one from S3 and a custom source. A catalog crawler covers the stores it supports, which AWS lists as including Amazon S3, Amazon RDS, Amazon Redshift, and Apache Hive. Anything outside those lists, such as a runbook in a wiki you did not connect, a decision made in a ticket, or a metric definition that lives in one analyst’s head, reaches neither. That residue is tribal knowledge, and it creates the enterprise context silos that agents inherit.
Permissions. AWS is explicit that ACL-aware retrieval in Managed Knowledge Base is filtering, not authorization. Your application authenticates users and passes their identity, matched by email, and group memberships are only as fresh as the last sync. S3 and custom sources do not support real-time ACL verification. None of that is a flaw; it describes where the responsibility sits. It does mean the permission model for documents and the permission model for tables are two separate models, and an agent that reads both inherits the gap.
Freshness. A sync-based index and a crawler-based catalog both lag their sources. The freshness of agent context decides whether an answer reflects last quarter’s definition or this quarter’s, and the staleness of an LLM knowledge base is the failure mode when nobody owns it.
Shared meaning. Two systems with two vocabularies give agents conflicting answers to the same question. This is the problem that a semantic layer for AI agents and an ontology address, and our comparison of the context layer, data catalog, and semantic layer sets out how the three categories differ. The difference between active metadata and a context layer matters here too, because shared meaning only helps when it stays current.
How does Atlan work as the context layer beside both?
Atlan is the context layer for AI built on the Context Lakehouse, which stores context in an open, Iceberg-native form. The aim is a shared enterprise context layer that holds definitions, lineage, and ownership once, so a knowledge base and a catalog read the same meaning instead of each keeping its own. For the full definition, see what a context layer is.
The video shows two agents on the same LLM answering one customer refund question, one with Atlan’s context layer and one without. That is the retrieval-versus-context gap in a four-minute demo.
Four capabilities matter for the AWS stack described above.
- Connectors across the AWS estate. Atlan’s connector list includes Amazon S3, Amazon Redshift, Amazon DynamoDB, Amazon DocumentDB, AWS Glue, Amazon Athena, Amazon QuickSight, Amazon MSK, Amazon SageMaker, and Amazon MWAA through OpenLineage. Metadata from those services lands in one graph rather than in separate consoles.
- Enterprise Data Graph. One graph spans systems, mapping entities and their relationships. The Enterprise Data Graph is where an agent finds an asset, a metric, and the lineage behind it in one place, and it extends the idea of a knowledge graph to metadata. Lineage over MCP lets an agent trace where a number came from.
- Active ontology. Business terms, domains, and metrics live on the same graph, so a definition has one home. The ontology 101 explainer covers the concepts.
- MCP server. Per Atlan’s documentation, Atlan MCP is a hosted server that lets AI clients use Atlan as a context layer, with search and discovery, lineage traversal, governed definitions and glossaries, SQL execution, and metadata curation, and it respects existing Atlan permissions. Our guide to the Atlan MCP server shows how it builds context for AI tools.
Context also needs to be engineered and versioned, not only collected. Context Engineering Studio and Context Repos are where teams build, test, and ship that context, and Context Agents help draft and maintain it for human review. Because Atlan sits across systems rather than inside one cloud, the same context serves agents whatever framework or memory layer they use, and it stays consistent in multi-agent systems.
Which should you start with?
Start from the question your agent has to answer. If it answers from documents, such as policies, support articles, and contract text, a knowledge base is the right first move, and a managed one removes most pipeline work. If it answers about data, such as which table holds net revenue or who owns a pipeline, you need a catalog with real definitions before retrieval can help. If it does both, and most production agents do, you need the two to agree on meaning.
For a single-account AWS team with a small, stable document set, Knowledge Bases plus the Glue Data Catalog may cover the need for a long time, and adding another layer would be overhead. The case for a cross-system context layer grows with the number of agents, the number of source systems outside AWS, and the cost of two vocabularies disagreeing. The comparison of agent context layers and RAG and the way AI memory, RAG, and knowledge graphs differ help frame that decision. Teams that want to see the metadata gap on their own estate can start with the skill above, which sorts what they track into four layers and names the one that is thin.
FAQs about Amazon Bedrock Knowledge Bases vs. external data catalog
1. How does Amazon Bedrock Knowledge Bases work?
Amazon Bedrock Knowledge Bases connects to a data source, parses each document, splits it into chunks, converts the chunks into vector embeddings with an embedding model, and writes them to a vector index. At query time, it embeds the question, finds the most similar chunks, and returns them or passes them to a model to generate a cited answer. For structured stores, it can instead turn the question into SQL.
2. What is the difference between Bedrock Knowledge Bases and an external data catalog?
Knowledge Bases retrieves content from your documents so a model can ground its answers, while a data catalog stores metadata about data assets, such as schemas, glossary terms, lineage, and ownership. One answers what your content says. The other answers which data exists and what it means.
3. Can Bedrock Knowledge Bases work with an external data catalog?
Yes. Knowledge Bases can connect to structured data through Amazon Redshift and the AWS Glue Data Catalog, using schema metadata to turn natural language questions into SQL. An agent can also call a knowledge base and a catalog side by side as separate tools, which is how most teams combine them today.
4. Can an agent on Amazon Bedrock AgentCore use Atlan alongside a knowledge base?
Yes. AgentCore Gateway exposes a managed knowledge base as MCP tools and can aggregate other MCP servers as targets behind the same endpoint. Atlan provides a hosted MCP server for search, lineage, and governed definitions, so an agent can retrieve passages from the knowledge base and ask Atlan what a metric means in one session. Check that the authentication method on each target fits your gateway setup.
5. Do you still need a data catalog if you already use Bedrock Knowledge Bases?
It depends on the questions your agents answer. If they only answer from documents, a knowledge base may be enough. Once agents need to answer questions about tables, metrics, ownership, or data quality, they need the definitions and lineage a catalog holds, and the numbers are only as trustworthy as that metadata.
6. How does Atlan work with the AWS Glue Data Catalog?
Atlan lists AWS Glue among its supported connectors, so metadata held in Glue can sit in the same Enterprise Data Graph as metadata from warehouses, BI tools, and orchestrators. Atlan then adds definitions, ownership, and lineage across those systems, which a catalog scoped to one cloud does not span on its own.
Sources
- How Amazon Bedrock knowledge bases work, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-how-it-works.html
- Turning data into a knowledge base, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-how-data.html
- Retrieving information from data sources using Amazon Bedrock Knowledge Bases, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-how-retrieval.html
- Amazon Bedrock Managed Knowledge Base is now generally available, AWS What’s New, June 17, 2026. https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-bedrock-managed-knowledge-base/
- Build a managed knowledge base, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-build-managed.html
- Build a knowledge base with Amazon Neptune Analytics graphs, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/knowledge-base-build-graphs.html
- Access Control Lists awareness enablement, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-acl.html
- Connect to your knowledge base through AgentCore Gateway, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-gateway-target.html
- Data discovery and cataloging in AWS Glue, AWS Documentation, 2026. https://docs.aws.amazon.com/glue/latest/dg/catalog-and-crawler.html
- Adding business context, AWS Glue Documentation, 2026. https://docs.aws.amazon.com/glue/latest/dg/catalog-business-context.html
- Amazon Bedrock AgentCore Gateway, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway.html
- Supported targets for Amazon Bedrock AgentCore gateways, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-supported-targets.html
- MCP servers targets, AWS Documentation, 2026. https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-target-MCPservers.html
- Connectors and capabilities, Atlan Documentation, 2026. https://docs.atlan.com/product/connections/references/connectors-and-capabilities
- Atlan MCP server overview, Atlan Documentation, 2026. https://docs.atlan.com/product/capabilities/atlan-ai/how-tos/atlan-mcp-overview