Skip to main content

Data Classification and Tagging: What AI Agents Need

Emily Winks, Data Governance Expert, Atlan
Data Governance Expert
Updated:
|
Published:
15 min read

Key takeaways

  • Classification sorts assets by sensitivity and value; tagging turns that judgment into a label a system can read.
  • An agent checks four things before it acts: what is PII, what policy applies, who can access it, what is certified.
  • Snowflake, dbt, and BigQuery tag independently; without a shared definition, "Confidential" drifts across all three.
  • Mastercard enriched 30,000+ assets with one person and saved 6,000+ hours pairing AI-assisted tagging with human review.

What is data classification and tagging?

Data classification sorts assets by sensitivity, type, ownership, or business value; tagging turns that judgment into a metadata label a system can act on. Together they tell people and AI agents what an asset contains, who is allowed to touch it, and what happens if they get it wrong. Snowflake, dbt, and BigQuery each apply their own tags, which is useful inside one tool and a problem the moment an agent has to reason across all three: a governed context layer holds one definition of "Confidential" or "PII" that every system inherits, instead of three definitions quietly drifting apart.

What this page covers:

  • Classification vs. tagging vs. labeling: where each one starts and stops
  • Why an agent needs it: the four questions it has to answer before acting
  • Tagging across the stack: how Snowflake, dbt, and BigQuery each handle it
  • Getting it to scale: bulk workflows and AI-assisted tagging without losing accuracy

Is your tagging AI-ready?


Before an AI agent queries a table or writes back a value, it inherits who can see what. That inheritance only works if the underlying asset already carries a label: is this PII, what policy applies, who is allowed to touch it, is it certified. Classification is the judgment behind that label. Tagging is what makes the judgment machine-readable, and how AI agents get access to enterprise data at all.

Check Which Metadata Layers You Have


Give it what you track today or a sample of real field and table names. It sorts them into technical, business, operational, and social metadata and names the layer that’s missing. Read the skill.

Paste into a new chat

Use the skill at https://atlan.com/skills/metadata-layer-check.md to check which metadata layers are present before classifying an estate. Ask me for whatever it needs.

Run once in a terminal

curl -fsSL --create-dirs \
  -o ~/.agents/skills/metadata-layer-check/SKILL.md \
  https://atlan.com/skills/metadata-layer-check.md

For an agent

curl -fsSL https://atlan.com/skills/metadata-layer-check.md

Most data teams already have an inventory. Far fewer have a classification scheme consistent enough for an agent to trust. An inventory tells you an asset exists. Classification and tagging tell you, and anything reading the asset, what it is and how it should be handled.


Question What classification answers What tagging does about it
What is this data? Sensitivity, type, ownership, business value Applies a label like PII or Confidential
Who can touch it? The access tier the classification implies Attaches a policy the platform can enforce
Is it trustworthy? Freshness, quality, certification status Surfaces a tag like Verified or Stale
Can an agent act on it? Whether the first three are answered at all Makes the answer queryable at inference time

Data classification and tagging, defined

Data classification is the process of sorting data assets into categories based on defined criteria: sensitivity, ownership, quality, or business value. Tagging is the classification technique that turns that sorting into a metadata label attached to the asset itself. Labeling is a looser cousin of tagging: it applies the same idea without a fixed schema, which makes it faster to start and harder to govern at scale.

The distinction matters. A spreadsheet listing which tables are “sensitive” is classification without tagging: a judgment in someone’s head, not on the asset. A PII tag on a Snowflake column is classification made durable, and it is what types of metadata AI agents need to reason about the column. Only the tagged version survives a schema change or an agent querying the table six months later.

A data scientist choosing a training set relies on tags like accurate or latest instead of opening ten tables by hand. Governance runs on the same logic at higher stakes. Classifying a column as PII or Confidential is what lets access policies, masking, and retention attach automatically instead of by memory, and it is the foundation any data governance taxonomy is trying to formalize.

An inventory alone does not get you there. Listing every table an organization owns produces a catalog of names, not a system anyone can safely act on without first understanding what each entry is. Data catalogs that stop at inventory create exactly this gap, and it is one reason is your data catalog keeping up with what AI agents need is the question more teams are asking. Classification and tagging are the layer that closes it: a well-defined scheme turns an inventory into governed context that people and AI systems can use consistently, not just browse.


Why an agent needs this before it acts

An AI agent about to answer a question or write a value has to clear four checks, in order: what is classified as PII, what policy applies to it, who is allowed to access it, and whether it is certified. Classification supplies the judgment behind each check. Tagging is what makes it something the agent can read at inference time, not something a steward explains in a Slack thread. Skip the check and the ROI of AI agent governance inverts fast: an ungoverned agent answering confidently on a PII field it should never have touched costs more than asking a steward first would have.

This is a different posture than the compliance-era version of classification, where the goal was passing an audit. The agent-era version is offense, not defense: classification and tagging exist so an agent can move fast on the data it is cleared to touch, not just so an auditor can confirm nothing bad happened after the fact. Agentic stewardship puts a steward on the loop rather than in it: the agent does the volume tagging work, and a person approves the edge cases and exceptions, the same logic behind securing multi-agent systems in the enterprise rather than trusting each agent’s judgment blind.

General Motors classifies data governance by design, before production, rather than retrofitting it after an agent is already live: a shift-left posture that GM’s own team credits with 45 to 60% less effort on each subsequent agent it ships, because the classification and policy work for a data asset only has to happen once. Governed context becomes the thing every new agent reuses instead of rebuilding, which is also why most data governance implementations fail when classification is treated as a one-time project instead of infrastructure every new agent depends on.


Tagging and labeling: the techniques underneath classification

Tagging is the most common classification technique: it assigns a discrete label to an asset so it can be discovered, managed, and protected consistently. Labeling is the more flexible sibling. It does not require a fixed schema, which makes it useful for ad hoc annotation but harder to enforce a policy against at scale. Most mature programs start with labeling and converge on tagging once a classification scheme stabilizes.

Classification, tags, and labels get used interchangeably in casual conversation, but the distinction drives how a governance program actually matures. Classification is the decision. Tagging is the durable, structured version of that decision. Labeling is the informal version. A program that never moves past labeling ends up with inconsistent, undiscoverable metadata, closer to structured vs. unstructured data that nobody has actually structured; a program that jumps straight to rigid tagging schemas before anyone agrees on definitions ends up with tags nobody trusts enough to use.


Tagging across the modern data stack

Every major platform in the modern data stack ships its own tagging mechanism, and that is both the opportunity and the problem. Each one is genuinely useful inside its own tool. None of them, on their own, gives an agent a single answer when it has to reason across all three at once, which is the same limit that shows up in context layer for Snowflake coverage questions: native tags are real, and they stop at the platform’s edge.

Snowflake tags let organizations classify and protect objects, tables, views, columns, and schemas, by attaching a tag object directly to them. Tags can drive tag-based masking and row-access policies, so a PII tag does not just describe a column, it can trigger the masking policy that protects it.

dbt tags label resources, models, snapshots, seeds, so teams can select and run subsets of a project by tag instead of by name, applied at the table, column, or test level as a data contract for AI takes shape.

BigQuery and Google Cloud’s data catalog use aspects to attach business and technical metadata, including classification-relevant fields like PII and retention, to an asset.

The problem shows up the moment an agent has to answer a question that spans all three: is this customer table, this dbt model, and this BigQuery dataset governed by the same PII definition? Locally, yes, in each tool’s own vocabulary. Globally, only if something outside all three, a model-agnostic context layer, holds one shared definition each local tag maps back to.



Four benefits, once classification and tagging actually hold

Better data management. A clear view of what an asset contains and how it is structured lets teams organize, retrieve, and reuse it instead of rediscovering it every time. A sales team relying on a Confirmed Lead tag skips re-qualifying leads that are already dead or duplicate.

Stronger data security. Classifying sensitive data before it spreads reduces the odds it ends up somewhere it should not. Patient records tagged PII and PHI inherit the right protection policy automatically, instead of depending on whoever touches the table next remembering to apply one, the exact gap handling PII and sensitive data in AI pipelines has to close before an agent is allowed near the table.

Simpler compliance. Marking assets highly sensitive, confidential, or long-term retention gives an auditor something concrete to validate against, rather than a policy document and a promise. Data privacy governance programs run on exactly this kind of explicit, tagged evidence.

Improved governance. Classification-based policies enforce access by role instead of by memory: an Engineering-only tag on a schema does the job a hallway conversation used to do, and does it the same way every time, which is the whole premise behind context layer for data governance teams.


Classification and tagging as the input to governance

Data governance is one of the highest-value applications of classification, and one of the places where getting it wrong compounds fastest. Assigning tags to a predefined schema, sensitive, confidential, regulated, archival, is what lets governance policies actually attach to the right assets instead of applying uniformly, or not at all.

The modern version pushes classification upstream: enriching data with governance context closer to the source rather than retrofitting it downstream. Shift-left governance means embedding classification and policy earlier in the pipeline, before an asset reaches an analyst or an agent, so the labeling decision only happens once. Fivetran’s column-blocking, Snowflake’s tag-based masking, and dbt’s model-level tags are three examples of platforms moving classification earlier rather than treating it as downstream cleanup, the same instinct behind preparing enterprise data for AI agents instead of fixing it after the fact.


Real stories from real customers: classification and tagging at enterprise scale

"AI initiatives require more context than ever. Atlan's metadata lakehouse is configurable, intuitive, and able to scale to hundreds of millions of assets. As we're doing this, we're making life easier for data scientists and speeding up innovation."

— Andrew Reiskind, Chief Data Officer, Mastercard

"Context is the differentiator. Atlan gave our teams the shared vocabulary and lineage to move from reactive data management to proactive AI enablement."

— Kiran Panja, Managing Director, Cloud and Data Engineering, CME Group

One Mastercard steward classified and enriched more than 30,000 assets and saved 6,000-plus hours in the process, evidence that AI-assisted tagging changes what “one person” can cover, provided a human is still reviewing the output rather than rubber-stamping it.


Getting classification and tagging to scale

Manual tagging does not survive contact with a real data estate. A handful of tables, sure. A few thousand, added weekly across Snowflake, Databricks, and BigQuery, no. Bulk, rule-based workflows are the first lever: tag by pattern, source system, or naming convention, instead of one asset at a time. Data lineage for AI matters here too: once an asset is tagged at the source, lineage lets the tag propagate downstream.

AI-assisted classification is the second lever, and the one that changes the economics. Context Agents can read a table’s structure and documentation, then propose a classification for review, turning 9 to 12 months of manual stewardship into roughly 30 days. That is not full autonomy: current tools are not mature enough to replace human judgment on sensitive calls. The credible version is human on the loop: the agent proposes at volume, a steward corrects it, and the correction trains the next batch. Skipping this is exactly what data observability for AI pipelines exists to catch: a tag going stale silently is a quality problem before it is a governance one.

The gap this closes is quantified. Atlan Frontier Labs found that governed context, including PII, policy, and access classification, improved natural-language query accuracy 38% across 174 enterprise queries and 522 evaluations. The lift came from context layer evaluation criteria an agent has available, not a bigger model, and it holds whether the agent reaches that context directly or through an MCP server.


Bringing classification and tagging under one definition

Every platform in a modern stack will keep tagging its own way. That is not a problem to solve; it is how Snowflake, dbt, and BigQuery are built to work. The problem is when “Confidential” in Snowflake, confidential in dbt, and a custom aspect in BigQuery drift apart because nothing outside the three platforms holds one shared definition.

Atlan’s Enterprise Data Graph is where that shared definition lives. A steward defines PII or Confidential once, and it propagates to the local tag inside Snowflake, dbt, and BigQuery instead of being redefined three times. A data scientist finds a dbt model by its tag whether they started in dbt or in Atlan, and an MCP registry exposes that same definition to any agent asking for it.

As a data estate spans more platforms, this gets harder every quarter, not easier: keeping definitions synchronized is ongoing maintenance a shared layer solves. An AI-ready data checklist confirms a data estate has reached that point rather than assuming it has.


What classification and tagging owe an AI agent, specifically

Return to the four questions an agent has to answer before it acts: what is PII, what policy applies, who can access it, what is certified. Classification supplies the answer to each one. Tagging is what makes the answer available at inference time, in the format the agent’s tool calls actually expect, rather than buried in a wiki page a human would have to translate.

That is the real argument for treating classification and tagging as context infrastructure rather than a governance chore. An agent that cannot answer those four questions either stalls, waiting on a human, or answers anyway and gets it wrong, the same risk that shows up across context observability for AI agents once a team starts measuring how often agents act on stale or unlabeled data. Neither outcome is acceptable once agents are doing production work instead of demos. Classification and tagging, done consistently and read by every system an agent touches, are what make the fast, correct answer the default one.


FAQs about data classification and tagging

1. What is data classification and tagging?


Data classification sorts assets by sensitivity, type, or business value. Tagging applies that judgment as a metadata label so people and systems, including AI agents, can act on it consistently.

2. What is the difference between tagging and classification?


Classification is the judgment: deciding an asset is Confidential, PII, or low-value. Tagging is the mechanism: writing that judgment onto the asset as a label a query, policy, or agent can read.

3. What are the four common types of data classification?


Public, internal, confidential, and restricted are the four tiers most organizations start from, ordered by how much damage exposure would cause and how tightly access should be controlled.

4. Why does an AI agent need classification and tagging?


Before an agent queries or writes data, it has to know what is PII, what policy applies, who can access it, and whether it is certified. Classification and tagging are what make those four answers explicit instead of assumed.


Sources

  1. Snowflake, “Tag-based Data Protection Policies.” https://docs.snowflake.com/en/user-guide/tag-based-policies
  2. dbt Labs, “Tags Reference.” https://docs.getdbt.com/reference/resource-configs/tags
  3. Google Cloud, “Enrich Entries with Aspects.” https://docs.cloud.google.com/dataplex/docs/enrich-entries-metadata
  4. Fivetran, “Data Blocking and Column Hashing.” https://fivetran.com/docs/core-concepts/features/data-blocking-column-hashing

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI. It translates business knowledge, including data definitions, working procedures, and governance policies, into context AI can actually use. This knowledge lives in a single Enterprise Data Graph that every team and AI agent can reach.

In Atlan's AI Labs benchmark, adding this context improved AI's text-to-SQL accuracy by 38%.

Gartner recognizes Atlan as a sample vendor for AI Context Platforms in its Emerging Tech Impact Radar for Generative AI. Atlan is trusted by over 400 enterprises representing $10T+ in market cap, including Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, and Elastic.

Bridge the context gap.
Ship AI that works.