AI Platform Team Playbook: Roles and Governed Context [2026]

Emily Winks, Data Governance Expert, Atlan
Data Governance Expert
Updated:07/24/2026
|
Published:07/24/2026
14 min read

Key takeaways

  • An AI platform team owns infrastructure such as the gateway, registry, and observability stack, not policy.
  • Hub-and-spoke is the dominant pattern, a central team owns shared tooling while product teams ship AI independently.
  • Team size scales from roughly 3 to 15 people early on to 15 to 50 or more at multi-business-unit scale.
  • Governance runs continuously only if the gateway and registry the platform team runs actually enforce it.

What does an AI platform team do?

An AI platform team is the group that builds and runs the shared infrastructure, including the gateway, model and agent registry, observability, and cost management, that every AI initiative in the business depends on. It is distinct from an AI Center of Excellence, which sets policy, approves use cases, and owns risk classification rather than infrastructure. Most organizations size this team at 15 to 50 people once it serves multiple business units, at roughly one platform engineer for every four to six people building models or agents, a widely used staffing benchmark.

What this team typically owns

  • Gateway and routing: the AI and LLM gateway every model call passes through
  • Model and agent registry: lifecycle tracking, ownership, and risk classification
  • Observability and cost management: monitoring, incident response, and FinOps for AI
  • The CoE handoff: the boundary where a policy decision becomes an infrastructure change

Is your data estate AI-agent ready?

Assess Your Readiness

Most organizations start structuring for AI the way they started structuring for cloud: one team builds a pipeline, then another rebuilds it six months later because nobody owns the shared parts. The fix is a dedicated AI platform team, typically 15 to 50 people once it serves multiple business units, at roughly one platform engineer per four to six people building models or agents, a widely used benchmark, per KORE1’s 2026 benchmarking. This playbook treats “run the platform reliably” and “govern AI responsibly” as the same job: governance only executes continuously if the platform team’s own gateway, registry, and observability stack enforce it.

What you get here: a 5-step build sequence, a maturity scorecard, twelve diagnostic questions, and the centralized-versus-federated decision framework.

Category AI platform team operating model
Key stakeholders Platform/AI Engineering Lead, CDO/CAIO, AI Center of Excellence
Team size range 3-15 people (early stage) to 15-50+ (multi-business-unit scale)

This guide picks up where the architecture question ends. If you have not yet built the underlying platform, start with how to build a centralized AI platform for enterprise; this playbook assumes that exists and focuses on who staffs it. It complements how to standardize AI tooling across business units, what is an AI control plane, and what is LLMOps.


Why enterprises need a dedicated AI platform team

Permalink to “Why enterprises need a dedicated AI platform team”

Enterprises with no dedicated platform function end up with AI scattered across business units, each rebuilding the same gateway, registry, and observability plumbing. Per OutSystems Research (2026), 94% of enterprises report AI sprawl as their top concern, and the pattern is not slowing.

The impact compounds quietly: a business unit without a platform function re-solves gateway routing and cost attribution independently. Per Gartner (2024), 74% of organizations already use governance tools for AI governance, adoption that assumes someone is running the tooling, not just buying it.

This guide serves the Platform or AI Engineering Lead who owns this team, the CDO or CAIO who funds it, and the AI Center of Excellence that needs this team’s enforcement to make policy real. The broader AI agent governance function depends on this team existing.

Buying governance tools and staffing the team that runs them are two different budget lines, the same gap AI risk management treats as a continuous practice, not a document.


What does an AI platform team own, and how is that different from an AI Center of Excellence?

Permalink to “What does an AI platform team own, and how is that different from an AI Center of Excellence?”

An AI platform team owns the infrastructure that makes AI systems run reliably in production: gateway, registry, observability, cost management. The AI Center of Excellence owns policy, use-case approval, and risk classification. Conflating the two is one of the most common org-design mistakes in this space.

Tooling category What the platform team owns
LLM/AI gateway Routing rules, rate limits, provider failover
Model/agent registry Registration, lifecycle status, risk classification
Observability Dashboards, alert thresholds, on-call escalation
Cost management / FinOps Cost attribution, budget alerts
Orchestration Runtime and workflow versioning
Access control Where the CoE’s policy executes

The platform team’s failure mode is invisibility: gateway routing breaks overnight with no on-call owner. The CoE’s failure mode is the opposite: a choke point that centralizes review without centralizing capability. For the charter and roles on the CoE side, see how to build an AI Center of Excellence.

AI platform team vs. AI Center of Excellence: who owns what

Clear swimlanes: the CoE defines the rules, the platform team enforces them at runtime, on one shared governed context layer. Source: Atlan

The AWS Well-Architected ML Lens frames clear ML roles as sound guidance, but written for a pre-agent era that never anticipated a registry tracking agents alongside models. When gateway and observability read from the same governed context, the CoE’s policies and the platform team’s runbooks become one enforcement mechanism, not two documents.

Inside Atlan AI Labs & The 5x Accuracy Factor

How governed context changes AI accuracy for the infrastructure your platform team runs.

Get the Ebook

Should your AI platform team be centralized or federated?

Permalink to “Should your AI platform team be centralized or federated?”

The centralized-versus-federated decision mirrors the classic build-versus-buy trade-off. Most enterprises land on a hub-and-spoke hybrid. Ayautomate’s 2026 analysis identifies five recurring org patterns: embedded AI engineers, a centralized platform team, a hybrid model, an AI council, and outsourced-build-with-internal-lead.

Centralized: consistent gateway and registry policy, a single audit trail, but can become a bottleneck if the team also tries to own delivery.

Federated (hub-and-spoke): faster embedded delivery. PostHog’s engineering handbook describes exactly this: its platform team owns architecture, performance, and UX while product teams ship their own AI capabilities without central approval. The trade-off: without a shared governed context source, federation produces shadow AI and configuration drift.

Decision framework: centralize with fewer than 5 product teams and tight cost control. Federate with 10 or more teams where delivery speed binds, but only if the hub still owns the shared gateway, registry, and context source, not just standards nobody’s tooling checks against.

General Motors’ AI program shows federated-but-governed at real scale: 70+ AI projects across cloud and on-premises environments, governance and engineering running in parallel, the centralized-versus-federated trade-off resolved in favor of a governed hub.


How do you build AI platform team maturity? A 5-step readiness framework

Permalink to “How do you build AI platform team maturity? A 5-step readiness framework”

Standing up the team follows the same fixed sequence: charter and CoE boundary, staff the roles, stand up the tooling, design the intake process, then on-call rituals. Skipping ahead is the most common reason teams stall.

Step 1: Charter and CoE boundary (1-2 weeks)

Document scope (infrastructure, not policy), hand-off points with the AI CoE, and success metrics (uptime, onboarding time, incident MTTR). A charter without this predicts the choke-point failure mode above.

Step 2: Staff the core roles and ratio (2-4 weeks to first hires)

A Platform or AI Engineering Lead, ML/AI Platform Engineers, a cost or FinOps owner (often fractional early on), and on-call members. Roughly one platform engineer per four to six model builders; team size scales from 3-15 early stage to 15-50 or more, per KORE1.

Step 3: Stand up the tooling stack (4-8 weeks)

Gateway, registry, observability, cost management, orchestration, and access enforcement, mapping to the table above. The LLM gateway middleware market is growing at a 49.6% CAGR through 2034, per Maxim AI; see LiteLLM vs. Portkey vs. Bedrock Gateway for the choice itself.

Step 4: Design the model-onboarding intake process (2-3 weeks)

Step Owner SLA
Risk and complexity triage AI Center of Excellence 3 business days
Registry entry and access provisioning Platform team 5 business days
Gateway routing configuration Platform team 3 business days

Answer who owns onboarding explicitly, or every new use case becomes a negotiation.

Step 5: On-call rituals and SLAs (ongoing, review quarterly)

Metric Target SLA
Gateway incident response 15 minutes
Registry approval time 5 business days
Observability alert acknowledgment 10 minutes
Cost anomaly escalation Same business day
Uptime for shared AI infrastructure 99.9% monthly

Opsio’s operational-readiness checklist treats a documented runbook, an on-call rotation, and an escalation procedure as the bar for “ready.”


AI platform team maturity scorecard: how ready is your team?

Permalink to “AI platform team maturity scorecard: how ready is your team?”

A weighted self-assessment scorecard turns “are we ready” from a gut feeling into a structured score your team can revisit every quarter.

Criterion Weight Your score (1-5)
Gateway ownership clarity 15%
Registry lifecycle coverage 15%
On-call coverage defined 15%
Model-onboarding SLA documented 15%
Cost attribution mechanism 10%
CoE boundary documented and agreed 15%
Incident postmortem process exists 10%
Context/MCP delivery consistent 5%

Scoring guide: 5 = fully operational and documented, 3 = exists but inconsistent, 1 = does not exist. Score your own team, not a vendor shortlist.

Signs your platform team has become a bottleneck: no one owns model onboarding end-to-end, the CoE approves use cases the platform team was never resourced to run, and business units stand up shadow gateways because the central one is too slow.

See debugging multi-agent systems and decision traces, the audit-trail infrastructure this team also maintains, and how to secure multi-agent systems in the enterprise.

Context Maturity Assessment

Score your own tooling and governance maturity before you fill out the scorecard above.

Assess Your Maturity

What questions should your AI platform team be able to answer?

Permalink to “What questions should your AI platform team be able to answer?”

Where the scorecard above produces a number, these twelve questions produce a conversation across tooling ownership, the CoE boundary, on-call readiness, and cost accountability. If your team cannot answer most of them today, that is the gap to close next.

Tooling ownership questions

Permalink to “Tooling ownership questions”
  1. Which team owns gateway routing and rate-limit rules when two business units conflict?
  2. How is a new model or agent registered, and who approves it?
  3. What percentage of incidents are caught by monitoring before a user reports them?

Governance and CoE boundary questions

Permalink to “Governance and CoE boundary questions”
  1. When the AI CoE approves a use case, who is responsible for provisioning it?
  2. What happens when a business unit deploys an AI tool the platform team did not know about?
  3. Is there a documented hand-off point between “policy decision” and “infrastructure provisioning”?

On-call and incident response questions

Permalink to “On-call and incident response questions”
  1. Who is on call specifically for AI/LLM infrastructure, separate from general on-call?
  2. What is the target time-to-acknowledge for a gateway or registry incident?
  3. Is there a postmortem process specific to AI incidents, such as prompt injection or model drift?

Cost and SLA questions

Permalink to “Cost and SLA questions”
  1. What SLA does the platform team offer for model-onboarding turnaround?
  2. Does the team have a documented uptime target for shared AI infrastructure?
  3. Is there a defined escalation path when AI cost exceeds budget mid-quarter?

Eric Paulsen, Field CTO, International at Coder puts it plainly: “Governance in regulated environments is simultaneously a security problem, a platform problem, a performance problem, and a cost problem.” Answering these in writing separates a team that can pass an audit from one that only feels ready.


How Atlan approaches the AI platform team’s tooling stack

Permalink to “How Atlan approaches the AI platform team’s tooling stack”

The tooling stack a platform team owns, gateway, registry, observability, cost management, only enforces governance continuously if every tool reads from the same governed context instead of siloed configs.

Most platform teams run their stack on separate configurations with no shared source of truth, so the AI CoE’s policies live in a document the infrastructure never checks against. Governance stays theoretical even when the team is fully staffed, a gap that shows up as an audit finding, not a system alert.

Atlan’s Context Governance & Observability and AI Governance Studio give the platform team a single model and agent registry with risk classification, owner, and deployment status in one place, the same context engineering discipline this playbook has been building toward. The MCP server and Context Repos give the gateway one place to route through, so every AI tool draws from the same governed context graph rather than rebuilding it per tool, the same shared-language problem Workday solved by drawing on context already built for the business; the customer story below covers that work firsthand.

See how to implement an enterprise context layer for AI and the broader AI governance framework this tooling enforces.


Real stories from real customers: governed context across the platform stack

Permalink to “Real stories from real customers: governed context across the platform stack”

"We're excited to build the future of AI governance with Atlan. All of the work that we did to get to a shared language at Workday can be leveraged by AI via Atlan's MCP server…as part of Atlan's AI Labs, we're co-building the semantic layer that AI needs with new constructs, like context products."

— Joe DosSantos, VP of Enterprise Data & Analytics, Workday

"Atlan is much more than a catalog of catalogs. It's more of a context operating system…Atlan enabled us to easily activate metadata for everything from discovery in the marketplace to AI governance to data quality to an MCP server delivering context to AI models."

— Sridher Arumugham, Chief Data & Analytics Officer, DigiKey

Context Layer ROI Calculator

Estimate what governed context is worth before your next tooling business case.

Calculate Your ROI

Why the platform team is where AI governance either executes or stalls

Permalink to “Why the platform team is where AI governance either executes or stalls”

The platform team and the AI Center of Excellence are not the same function: the CoE decides what is allowed, and the platform team’s infrastructure enforces it. Hub-and-spoke only works if the hub still owns the shared gateway, registry, and context source, not just written standards.

Treat the maturity scorecard above as a living quarterly document. The questions your team can and cannot answer today, more than any org chart, are the most honest read on where to invest next.

An AI platform team is where governance either becomes a runtime property of the systems already running, or stays a document nobody’s infrastructure checks against. Treating AI infrastructure and AI governance as one operating layer is what Atlan calls the enterprise AI operating layer; an agent harness without governed context fails for the same reason, quietly, at the worst possible time.


FAQs about the AI platform team’s operating model

Permalink to “FAQs about the AI platform team’s operating model”

1. What does an AI platform team do?

Permalink to “1. What does an AI platform team do?”

It runs the shared infrastructure every AI initiative depends on: the gateway, the model and agent registry, observability, and cost management. It does not set policy; that is the AI Center of Excellence’s job.

2. What is the difference between an AI platform team and an AI Center of Excellence?

Permalink to “2. What is the difference between an AI platform team and an AI Center of Excellence?”

The platform team owns infrastructure: gateway, registry, observability, cost. The CoE owns policy: use-case approval, risk classification, intake. Conflating the two is why governance stalls even when both exist.

3. How big should an AI platform team be?

Permalink to “3. How big should an AI platform team be?”

Early-stage teams run 3 to 15 people; teams serving multiple business units scale to 15 to 50 or more, at roughly one platform engineer per four to six model builders.

4. What roles are on an AI platform team?

Permalink to “4. What roles are on an AI platform team?”

A Platform or AI Engineering Lead, ML/AI Platform Engineers, a cost or FinOps owner (often fractional early on), and dedicated on-call members. Evaluation and incident-response specialists split off as the team scales.

5. What SLAs should an AI platform team offer the rest of the business?

Permalink to “5. What SLAs should an AI platform team offer the rest of the business?”

At minimum: gateway incident acknowledgment within 15 minutes, registry approval within 5 business days, observability alerts within 10 minutes, and a documented uptime target.

6. Who is on call when an AI agent or LLM pipeline breaks in production?

Permalink to “6. Who is on call when an AI agent or LLM pipeline breaks in production?”

The platform team’s dedicated on-call rotation, separate from general platform on-call, which typically lacks the context needed to diagnose an incident quickly.


Sources

Permalink to “Sources”
  1. Agentic AI Goes Mainstream in the Enterprise, but 94% Raise Concern About Sprawl, OutSystems Research, 2026. https://www.businesswire.com/news/home/20260407749542/en/Agentic-AI-Goes-Mainstream-in-the-Enterprise-but-94-Raise-Concern-About-Sprawl-OutSystems-Research-Finds
  2. Gartner Says Data and Analytics Governance Will Evolve to Ensure Organizational Alignment, Gartner, 2024. https://www.gartner.com/en/newsroom/press-releases/2024-10-22-gartner-says-data-and-analytics-governance-will-evolve-to-ensure-organizational-alignment
  3. MLOPS02-BP01: Establish ML Roles and Responsibilities, AWS Well-Architected Machine Learning Lens. https://docs.aws.amazon.com/wellarchitected/latest/machine-learning-lens/mlops02-bp01.html
  4. AI Engineering Org Structure (2026): 5 Patterns That Work, ayautomate. https://www.ayautomate.com/blog/ai-engineering-org-structure/
  5. PostHog Engineering Handbook: AI Platform Team Structure, PostHog. https://posthog.com/handbook/engineering/ai/team-structure.md
  6. AI Team Structure: Roles, Reporting & Headcount (2026), KORE1. https://www.kore1.com/building-ai-team-structure-2026/
  7. Top 5 Enterprise LLM Gateways in 2026, Maxim AI. https://www.getmaxim.ai/articles/top-5-enterprise-llm-gateways-in-2026/
  8. AI PoC to Production: Scaling from Pilot to Enterprise, Opsio. https://opsiocloud.com/blogs/ai-poc-to-production-scaling-guide/
  9. Platform Teams’ Playbook for Scaling AI Coding Agents in Regulated Industries, Eric Paulsen, Platform Engineering. https://platformengineering.org/blog/platform-teams-playbook-for-scaling-ai-coding-agents-in-regulated-industries

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI, a Leader in the Gartner Magic Quadrant for D&A Governance (2026) and the Forrester Wave for Data Governance (Q3 2025). Atlan unifies your data, business knowledge, and the meaning behind your terms into one Enterprise Data Graph that gives every team and every AI agent the trusted context they need. Trusted by Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, Elastic, and 400+ enterprises representing $10T+ in market cap.

Bridge the context gap.
Ship AI that works.

[Website env: production]