Most public build guides for an AI Center of Excellence solve the org chart and stop there, leaving out whether the CoE’s decisions are enforced anywhere. By the end of 2027, Gartner predicts more than 40% of agentic AI projects will be canceled due to escalating costs, unclear ROI, and inadequate risk controls (Gartner, June 2025), and a CoE with no enforcement teeth is a common contributor. Atlan treats building a CoE as a fixed sequence: charter, org model, staffing, intake workflow, cadence, then tooling. What this guide adds on top of that sequence: whether the CoE’s approvals live only in a charter, or in the same governed context layer the agents it approved already read from.
This guide walks the sequence step by step, reconciles the conflicting headcount advice in circulation, and ends with a diagnostic: how to tell if your CoE has quietly become the bottleneck it was built to prevent.
- Charter first. Roles and hiring plans are meaningless without a written scope and named decision rights.
- Hub-and-spoke wins by default. Centralized models bottleneck delivery; fully federated models produce shadow AI.
- Headcount bands aren’t contradictory. 3-5, 5-8, and 15-30 person CoEs describe maturity stages of the same org, not disagreement.
- A clean RACI can still fail if its approvals never touch the runtime path where agents retrieve data and act.
| Tutorial overview | |
|---|---|
| Difficulty | Intermediate (organizational design, not a technical build) |
| Time required | 90 days to a first production use case; ongoing to mature |
| Roles and functions used | Executive sponsor, CoE director, engineer, product manager, governance and risk role, business-unit AI leads |
| What you’ll build | A chartered, staffed, and running AI Center of Excellence with an intake workflow, a review cadence, and a tooling stack |
| Prerequisites summary | Executive sponsorship, at least one candidate AI use case, and a data platform capable of enforcing access policy |
The CIO's Guide to Context Graphs
A hub-and-spoke CoE still needs a shared, queryable definition layer underneath it. This guide covers how CIOs are architecting that layer for agentic AI.
Get the CIO Context GuideWhat an AI center of excellence actually does (and doesn’t do)
Permalink to “What an AI center of excellence actually does (and doesn’t do)”An AI Center of Excellence is the operational hub that coordinates AI use-case intake, cross-functional review, and the governance calendar, sitting between executive strategy and technical enforcement, not a replacement for either. It is Tier 2 of a three-tier AI governance operating model: Tier 1 is board-level strategy, Tier 3 is technical enforcement at the data platform, and the CoE runs the tier between them, its structure determining how accountability gets assigned for the AI agent primitives it approves, memory, tools, and orchestration alike.
Scope the CoE explicitly so it doesn’t expand past its lane: it makes advisory recommendations on most use cases and holds binding authority only on a defined subset, typically regulated data or customer-facing decisions. What it does not own is Tier 3 technical enforcement, which belongs to the AI Platform Team and the underlying AI agent access control infrastructure. Only 14% of enterprises enforce AI governance enterprise-wide (ModelOp, 2025), and that gap is what Tier 3 exists to close, not the CoE. A CoE that starts reviewing pull requests directly has stopped coordinating and become a second engineering team with worse tooling.
Keep one thing in mind as the build starts: a CoE’s decisions become real when they land in a governed, queryable context agents actually read from at runtime, not when they’re approved in a meeting. That’s the standard the finished CoE should be held to once Steps 1 through 7 are done. Start with Step 1.
Step 1: Secure an executive sponsor and write the charter
Permalink to “Step 1: Secure an executive sponsor and write the charter”The first artifact a CoE needs is a charter naming an executive sponsor, defining scope, and setting decision rights before any roles are hired. Without a written charter, a CoE’s authority is ambiguous the first time a business unit disputes a decision, and ambiguous authority loses that argument. Draft it as a decision-rights contract, not a mission statement.
| Charter section | What it defines | Example language |
|---|---|---|
| Mission | Why the CoE exists, in one sentence | “Coordinate AI use-case intake, review, and governance calendar across [org].” |
| Scope | Which AI initiatives fall under the CoE | “All production AI and agentic systems touching customer or regulated data.” |
| Decision rights | Advisory vs. binding, by use-case class | “Advisory on internal tooling; binding sign-off on customer-facing and regulated use cases.” |
| Budget | Funded headcount and tooling spend for the first year | “3 FTE plus a $[X] tooling allocation for year one.” |
| Success metrics | What the CoE proves in its first 12 months | “One production use case shipped per quarter; intake SLA held at 48 hours.” |
| Reporting line | Who the CoE answers to | “Reports to the CDO, with a dotted line to the CAIO on use-case approval.” |
Get the sponsor’s sign-off in writing before Step 2; an unsponsored CoE has no authority the first time its first hard call gets challenged. Data-driven initiatives typically report to the CDO, product-driven ones to the CAIO. Name which EU AI Act obligations apply and which use cases require entry into a formal AI registry, since Step 4’s intake design depends on that scope.
Step 2: Decide your org model: centralized, federated, or hub-and-spoke
Permalink to “Step 2: Decide your org model: centralized, federated, or hub-and-spoke”Most mid-to-large enterprises should default to a hub-and-spoke model: a central team sets standards and trains, while business units execute their own use cases inside those guardrails. Enterprise data leaders describe this as the practical middle ground, a central team that “trains people and dictates best practices” while business units “operate for their own departments.”
A fully centralized model enforces standards fast but becomes the delivery bottleneck: the team approving projects is also expected to deliver them (Agility at Scale). A fully federated model moves fast in the opposite direction: business units ship quickly but produce fragmented tooling and significant shadow AI adoption as teams route around any central process (AI Assembly Lines, 2026).
In practice, hub-and-spoke pairs centralized platform decisions (cloud, model provider, context management standards) with fully federated use-case delivery, each business unit running its own pilots inside those guardrails. One enterprise data leader called this the “sweet spot” between control and speed, echoing the centralized vs. federated debate most data organizations are already having. The org model choice is customer-evidence-led, not vendor-led.

Hub-and-spoke AI Center of Excellence org model: a central hub sets policy and tooling, business-unit teams execute on a shared governed context layer. Source: Atlan
Step 3: Staff it: roles and headcount by stage
Permalink to “Step 3: Staff it: roles and headcount by stage”CoE headcount ranges from 3 to 30 people depending on org size and use-case volume, not a single universal number. The competing bandings in circulation, 3-5 versus 5-8 versus 15-30, describe different maturity stages of the same organization, not disagreement.
| Stage | Headcount | Core roles | Trigger to next stage |
|---|---|---|---|
| Early stage | 3 to 5 | Director, 1 engineer, 1 product manager | First production use case shipped with baseline governance |
| Minimum viable CoE | 5 to 8 | Adds a governance or risk role | Repeatable intake process; several use cases in production |
| Scaled CoE | 15 to 30 | Embedded business-unit AI leads | Hub-and-spoke fully federated; enforcement embedded in the platform |
At the early stage, a director plus one engineer and one product manager ships a first production deployment and baseline governance (AI Assembly Lines, 2026). At the scaled stage, 15-30 people with embedded business-unit AI leads matches a fully federated hub-and-spoke model in motion (AppScale Blog, 2026). The two sources describe the same growth curve at different points, tied to org size and use-case volume.
Pro tip: Hiring a governance or risk specialist before the CoE ships its first production use case is a common overcorrection. Sequence staffing to what the intake queue actually needs, checked against where your org sits on an AI maturity framework rather than a headcount target borrowed from a different stage.
Context Maturity Assessment
Before you lock in headcount and staffing stage, check where your org's context maturity actually sits. Sequencing staffing to a stage you haven't reached yet is the overcorrection this step just warned about.
Assess Context MaturityStep 4: Design the intake and approval workflow
Permalink to “Step 4: Design the intake and approval workflow”Every working intake workflow answers four questions before a single use case moves: who submits it, who triages it and on what SLA, who scores it, and what triggers escalation. It is not an abstract RACI diagram consulted once and ignored under deadline pressure.
| Step | Owner | SLA | Escalation trigger |
|---|---|---|---|
| Submission | Business unit AI lead | N/A | N/A |
| Triage | CoE director | 48 hours | No triage owner available |
| Scoring (value, feasibility, risk) | Product manager + governance role | 5 business days | Score conflict between reviewers |
| Prioritization | CoE director | Weekly | Backlog exceeds capacity |
| Approval | Executive sponsor (binding use cases only) | 5 business days | Sponsor unavailable beyond SLA |
| Monitoring | Engineer | Continuous | Drift or policy violation detected |
Use a RACI-style skeleton, but name actual roles from Step 3 rather than generic titles like “governance team” (StackAI, 2026). Earlier-stage CoEs can run a lighter, three-function version: a minimal sign-off across AI, legal, and security. Every use case should also register in the AI registry for a running lifecycle record, not just a point-in-time approval.
This is also where AI agent access control, AI security review, and zero trust data governance principles enter, since the scoring criteria should reflect what the data platform can actually enforce, not just what the business case looks like on paper.
Step 5: Set the rituals and review cadence
Permalink to “Step 5: Set the rituals and review cadence”A CoE’s own operating rhythm runs on three cadences: weekly intake triage, monthly portfolio review, and quarterly charter review. This is distinct from, but consistent with, the operating model’s Tier cadence table, which runs quarterly strategic, monthly operational, and continuous technical review. The CoE’s cadence is its internal meeting rhythm; the Tier cadence is the org-wide structure it reports into.
| Ritual | Frequency | Attendees | Purpose |
|---|---|---|---|
| Intake triage | Weekly | CoE director, product manager | Score and prioritize new submissions |
| Portfolio review | Monthly | Full CoE, business-unit AI leads | Review use cases in flight, flag drift |
| Charter and mandate review | Quarterly | Executive sponsor, CoE director | Confirm scope, budget, and decision rights still hold |
The monthly portfolio review is also where memory and context governance issues surface for agentic use cases, since agent memory drifts in ways a one-time approval never anticipated. A CoE that only reviews at intake won’t catch this until a customer or auditor does.
Step 6: Pick the tooling stack
Permalink to “Step 6: Pick the tooling stack”Five categories make up a CoE’s tooling stack: registry and intake, risk scoring, policy and access enforcement, lineage and audit, and observability and monitoring. Evaluate by category before evaluating any specific vendor (StackAI, 2026).
- Registry and intake: tracks every AI use case and agent from submission through decommission, the lifecycle record customers most often ask for by name.
- Risk scoring: applies fixed criteria (value, feasibility, regulatory exposure) consistently instead of ad hoc judgment calls.
- Policy and access enforcement: blocks or allows an agent’s action based on what the CoE approved.
- Lineage and audit: proves, after the fact, what an agent touched and why.
- Observability and monitoring: catches drift, cost overruns, and policy violations in use cases the CoE already approved.
Atlan’s footprint sits at two categories: Context Governance & Observability handles policy and access enforcement; Policy Center and Data Lineage cover lineage and audit. A CoE that only tools up enforcement while leaving intake and risk scoring manual has solved half the problem. AI observability, AI agent observability, and an AI control plane round out a fully tooled stack; how to build a centralized AI platform covers the technical build underneath it, extended to context management for multi-agent systems once a use case involves more than one agent.
Step 7: Ship your first use case in 90 days
Permalink to “Step 7: Ship your first use case in 90 days”How does a new CoE prove itself inside its first quarter? By shipping one production use case, proving a measurable benefit, and establishing a repeatable path for the next ten initiatives, not a polished slide deck of future plans.
Days 1-30 are foundation: charter signed, sponsor confirmed in writing, first hire started. Days 31-60 build the team and infrastructure: remaining Step 3 roles hired, Step 4’s intake workflow live, Step 6’s tooling at least partially in place. Days 61-90 run the first pilot: one use case moves through the full intake-to-monitoring workflow, governance watching for drift from day one. A useful discipline: pick a hard checkpoint roughly a month before go-live and confirm the architecture is fully tested by then, rather than discovering integration problems the week of launch.
AI agent scaling in production and how to choose an agentic framework for the enterprise are worth reading before this step, since the framework decision greenlit here determines how much of Step 6’s tooling actually gets exercised on day one.
How do you tell if your CoE has already become a bottleneck?
Permalink to “How do you tell if your CoE has already become a bottleneck?”A CoE has become a bottleneck when its approval queue is growing, business units are adopting shadow AI to route around it, or it is doing delivery work itself instead of setting standards. The fix is connecting its decisions to something enforceable, not adding a fifth policy document.
| Symptom | Root cause | Fix |
|---|---|---|
| Approval queue growing | CoE centralizes review but not delivery capacity | Federate execution, keep policy central |
| Shadow AI appearing in business units | Intake process is too slow or too opaque | Publish the SLA from Step 4; make triage times visible |
| CoE doing delivery work itself | No clear line between advisory and binding scope | Revisit Step 1’s charter scope; it has drifted |
| Approved use cases stall before production | Decisions live in a charter or slide, not an enforced policy | Connect approvals to the access and policy layer from Step 6 |
This pattern shows up across enterprise conversations, not just public research: one financial services director described a newly launched CoE with “a lot of action” but little “results driven focus,” already restructuring toward a project-based model. At a large insurance enterprise, one stakeholder voiced a near-identical worry: that a promising initiative would become “a science experiment” that never reached production.
Neither example is a failure of the org design in Steps 1-5. Both are CoEs with a workable structure whose decisions were never enforced anywhere. A clean RACI is necessary, not sufficient; see the operating model this CoE plugs into for what has to be true underneath it. The same failure shows up when agent interoperability protocols get treated as a governance solution rather than plumbing, and when decision traces don’t exist to show what happened after a use case shipped. Both trace back to the root cause in why AI agents fail in production: approval without enforcement.
Next steps: from a working CoE to enforced governance
Permalink to “Next steps: from a working CoE to enforced governance”Once the CoE is running, the production-grade additions are security hardening on the intake tooling and monitoring that scales past the first handful of use cases. A CoE reviewing five use cases manually in year one needs automated risk scoring by year two, or the intake queue becomes the next bottleneck.
The advanced-configuration step is embedding enforcement directly into the platform as the CoE matures toward Tier 3, the point where policy stops living in an approved document and starts running as a technical control every agent hits automatically, per the operating model’s Tier 3 section.
That maturity path is what makes the agents a CoE approves genuinely enterprise-ready, not just approved on paper.
How Atlan connects CoE decisions to runtime enforcement
Permalink to “How Atlan connects CoE decisions to runtime enforcement”A CoE’s approvals only hold if enforced somewhere agents actually read from at runtime, not just documented in a charter. Everything in this guide so far, the charter, the intake workflow, the cadence, is manual, document-heavy work by design, and paperwork does not stop an agent from querying a restricted field.
Atlan connects the CoE’s decisions to a governed context layer agents already read from. An AI and agent registry gives the CoE the intake and lifecycle tracking function enterprise teams consistently ask for by name, rather than a spreadsheet nobody updates past month three. Context Governance & Observability is where an approved policy becomes a running access control instead of a sentence in a charter. Policy Center and Data Lineage supply the audit trail proving a CoE’s decisions actually held when a regulator asks. None of this replaces the org-design work in Steps 1-7; it is what makes that work stick past the first use case.
Across enterprise conversations spanning financial services, insurance, and retail, Atlan has heard the same organizational design requested independently: a hub-and-spoke or federated model paired with exactly this kind of runtime enforcement layer underneath it, evidence of a live operating need, not a hypothetical one. For the technical foundation, see how to build an AI agent harness, what context engineering means in this stack, and how to implement an enterprise context layer for AI. The AI governance framework covers the regulatory scaffolding a CoE’s charter should map to; context layer ROI is the business case for funding enforcement, not just org design.
Real stories from real customers: governance built to hold under agentic load
Permalink to “Real stories from real customers: governance built to hold under agentic load”"We're excited to build the future of AI governance with Atlan. All of the work that we did to get to a shared language at Workday can be leveraged by AI via Atlan's MCP server, and as part of Atlan's AI Labs, we're co-building the semantic layer that AI needs with new constructs, like context products."
— Joe DosSantos, VP of Enterprise Data & Analytics, Workday
"AI initiatives require more context than ever. Atlan's metadata lakehouse is configurable, intuitive, and able to scale to hundreds of millions of assets. As we're doing this, we're making life easier for data scientists and speeding up innovation."
— Andrew Reiskind, Chief Data Officer, Mastercard
WTF Is the Context Layer?
A CoE's approvals only hold if there's a governed layer underneath enforcing them. This live series unpacks what that layer actually looks like once it's running in production.
Watch the SeriesWhy a clean RACI is where AI governance starts, not where it ends
Permalink to “Why a clean RACI is where AI governance starts, not where it ends”Everything in this guide, the charter, the org model, the staffing plan, the intake workflow, the cadence, the tooling stack, is necessary. Skipping a step is exactly how a CoE ends up as the bottleneck the diagnostic above described. Building it in the right order matters more than building it fast.
But a well-built CoE is not sufficient on its own. Every failure mode here, shadow AI, stalled approvals, a science-experiment initiative that never ships, traces back to the same root cause: a decision made by capable people in a well-run process that never touched the runtime path where an agent retrieves data and acts, the same gap multi-agent system orchestration has to close at the technical layer. The governance operating model this CoE plugs into is what closes that remaining gap. Build the CoE first. Then make sure its decisions are enforced, not just documented.
FAQs about building an AI center of excellence
Permalink to “FAQs about building an AI center of excellence”1. How many people do you need to staff an AI Center of Excellence?
Permalink to “1. How many people do you need to staff an AI Center of Excellence?”Most CoEs start with 3 to 5 people: a director, one engineer, and one product manager, enough to ship a first production use case and baseline governance. A minimum viable CoE runs 5 to 8 people once a governance or risk role is added. A scaled CoE runs 15 to 30 people with embedded business-unit AI leads. The right number depends on org size and use-case volume.
2. What should an AI CoE charter include?
Permalink to “2. What should an AI CoE charter include?”A charter needs a named executive sponsor, a defined scope, decision rights that separate advisory guidance from binding approval, a budget, a 12-month set of success metrics, and a reporting line. Without a written charter, a CoE’s authority is ambiguous the first time a business unit disputes a decision.
3. Should an AI CoE be centralized, federated, or hub-and-spoke?
Permalink to “3. Should an AI CoE be centralized, federated, or hub-and-spoke?”Hub-and-spoke is the right default for most mid-to-large enterprises: a central team sets standards, trains staff, and makes platform decisions, while business units execute their own use cases inside those guardrails. Fully centralized models become delivery bottlenecks; fully federated models produce fragmented tooling and shadow AI risk.
4. What is the intake and approval process for new AI use cases?
Permalink to “4. What is the intake and approval process for new AI use cases?”A working intake workflow names who submits a use case, who triages it and on what SLA, who scores it against fixed criteria like business value, feasibility, and risk, who approves it, and who monitors it afterward. Each step needs a named owner, not a generic role title.
5. How do you avoid an AI CoE becoming a bottleneck?
Permalink to “5. How do you avoid an AI CoE becoming a bottleneck?”Watch for a growing approval queue, business units adopting shadow AI to route around the CoE, and the CoE doing delivery work itself instead of setting standards. The fix is federating execution while keeping policy central, and connecting decisions to an enforcement mechanism instead of a document nobody checks again.
6. How long does it take to stand up an AI Center of Excellence?
Permalink to “6. How long does it take to stand up an AI Center of Excellence?”A new CoE can reach its first production use case in roughly 90 days: charter and sponsor in the first 30 days, staffing and infrastructure in the next 30, and a first pilot with monitoring in the final 30. The bar for success is one shipped use case and a repeatable path for the next ten.
Sources
Permalink to “Sources”- Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- ModelOp, “AI Governance Benchmark Report,” 2025. https://www.modelop.com/ai-gov-benchmark-report
- Agility at Scale, “AI Center of Excellence: Why Most Become Bottlenecks and How to Build One That Scales.” https://agility-at-scale.com/ai/people-change/ai-center-of-excellence/
- AI Assembly Lines, “How Do You Staff an AI Center of Excellence? Roles and Org Design for Enterprise Leaders,” 2026. https://aiassemblylines.com/post/how-to-staff-ai-center-of-excellence-enterprise
- StackAI, “How to Build an Internal AI Center of Excellence (CoE): Roles, Processes, and Tooling,” 2026. https://www.stackai.com/insights/how-to-build-an-internal-ai-center-of-excellence-(coe)-roles-processes-and-tooling
- AppScale Blog, “How to Build an AI Center of Excellence: Structure & Roles,” 2026. https://appscale.blog/en/blog/how-to-build-ai-center-of-excellence-structure-roles-governance-2026
