---
title: "How to Build an AI Platform Team From Scratch (And in What Order)"
url: "https://atlan.com/know/how-to-build-ai-platform-team-from-scratch/"
description: "Learn how to build an AI platform team from scratch: which role to hire first, headcount triggers for scaling, and the gateway-vs-context tooling order debate."
author: "Emily Winks"
author_role: "Data Governance Expert"
published: "2026-07-27"
updated: "2026-07-27T00:00:00.000Z"
---

---

Most guides on this topic answer the wrong question: they describe hiring the AI *development* team that builds models, not the AI *platform* team that runs the shared gateway, registry, and observability infrastructure, [per Credal AI's definition](https://www.credal.ai/blog/what-is-an-ai-platform-team). Only 19% of AI-building orgs have one, per a [2026 CNCF and SlashData survey](https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-platform-engineering-tools-maturing-as-organizations-prepare-for-ai-driven-infrastructure/).

This is a sequencing tool, not a maturity scorecard. The [sibling playbook](https://atlan.com/know/ai-agent/ai-platform-team-playbook/) covers org chart, ratios, and SLAs.

**Difficulty:** Intermediate. **Time required:** Months to over a year. **Frameworks used:** Headcount-band staging (Bessemer), five-friction-signal model (PlatformEngineeringCost.com), core-product-vs-bolt-on fork (KAIRI). **Prerequisites:** One validated AI use case in production, executive sponsorship, and an honest read on core-vs-supporting.

---

## How does the AI platform team build-order decision tree work?

This section replaces the "Month 1 / Month 2" calendar with a two-axis decision tree: a scenario fork crossed with headcount-band triggers. [Evidently AI's review of ten ML platform builds](https://www.evidentlyai.com/blog/how-to-build-ml-platform) found no universal best order; [KAIRI's 2026 hiring-strategy analysis](https://medium.com/kairi-ai/two-paths-to-building-an-ai-ml-team-the-modern-hiring-strategy-45066e99b5c4) shows it flips entirely: ML engineer first if core product, data engineer first if bolt-on.


  Is AI your organization's core
  product, or a bolt-on?





  Core product
  First hire: ML engineer


  Bolt-on
  First hire: data engineer





  Seed-10: ML engineer


  10-20: + data engineer


  20-50+: + MLOps engineer


  Seed-10: data engineer


  10-20: + MLOps engineer


  20-50+: + ML engineer










  Both forks converge by
  20-50+ engineers

*A scenario fork crossed with three headcount bands. Source: Atlan, synthesized from Bessemer and KAIRI research.*

---

## What has to be true before you build a dedicated AI platform team?

Clear a readiness bar first: a headcount threshold, five friction signals, and one validated AI use case already in production, ruling out a dedicated *team* or *governance tooling* ahead of the evidence, this research's most common mistake.

The ROI threshold clusters around 30 product engineers, per [PlatformEngineeringCost.com's 2026 analysis](https://platformengineeringcost.com/when-to-invest): below it, a gateway costs more than it saves; above it, engineers rebuild the same [data infrastructure for AI](https://atlan.com/know/data-infrastructure-for-ai/) independently.

| Signal | Threshold | What it means |
|---|---|---|
| Deploy cycle for AI changes | Over 30 minutes | Manual review is the bottleneck |
| Onboarding to first AI contribution | Over two weeks | No shared scaffolding exists |
| AI-related infra incidents | 2+ per week | Ad hoc infra is now a risk |
| Senior time on AI infra help | 30%+ of their week | Expensive people doing platform work informally |
| New AI service spin-up time | Over a day | No standard pattern to fork from |

[OutSystems Research (2026)](https://www.businesswire.com/news/home/20260407749542/en/Agentic-AI-Goes-Mainstream-in-the-Enterprise-but-94-Raise-Concern-About-Sprawl-OutSystems-Research-Finds) found 94% of enterprises cite AI sprawl as a concern.

---

  Inside Atlan AI Labs & The 5x Accuracy Factor
  See what changes when the context layer feeds your models, before you finalize your platform team's tooling stack.
  Get the Ebook

---

## What order do you hire and build in?

A scenario fork at the root, three headcount bands, and a trigger signal for the next.

| Band | Core-product fork | Bolt-on fork | Next-band signal |
|---|---|---|---|
| Seed to 10 engineers | ML → data → MLOps engineer | Data → MLOps → ML engineer | Tooling decisions become expensive |
| 10 to 20 engineers | + technical leader (architecture calls) | + technical leader (architecture calls) | Shifts from "tool" to "people/architecture" |
| 20 to 50+ engineers | Diagnose people vs. architecture first | Diagnose people vs. architecture first | Ratios and CoE boundary (sibling playbook) |

### The fork: is AI your organization's core product, or are you retrofitting it onto an existing one?

If AI is your core product, [KAIRI's hiring-sequence research](https://medium.com/kairi-ai/two-paths-to-building-an-ai-ml-team-the-modern-hiring-strategy-45066e99b5c4) supports ML engineer first; if it's a bolt-on, data engineer first. Can't decide? Check whether a pipeline already moves production data reliably.

### Seed to 10 engineers: hire for velocity, not architecture

Make your first one to two hires per the fork above, hands-on coders who ship solo, not architects, per [Jessica Popp](https://www.bvp.com/atlas/inside-ai-pilled-engineering-teams-five-lessons-for-scaling-without-losing-the-plot) (Bessemer) and [KORE1](https://www.kore1.com/hire-first-ai-engineer/). DoorDash started with one prediction service; Uber's [Michelangelo platform](https://www.uber.com/us/en/blog/michelangelo-machine-learning-platform/) prioritized training and serving instead. Shopify didn't standardize its tool layer but still centralized cost control via a shared gateway.

### 10 to 20 engineers: the architecture-lock inflection point

Architectural decisions become permanent here, requiring experienced leadership over more hands-on coders.

### 20 to 50+ engineers: diagnose the constraint before you hire more

Figure out whether the bottleneck is people or architecture; the sibling playbook's roles, ratios, and [AI Center of Excellence](https://atlan.com/know/ai-agent-governance/) boundary are the reference.

> **Pro tip:** Still resolving gateway-versus-registry past the 20-50 band? That's a signal you under-invested in the [context and governance foundation](https://atlan.com/know/what-is-context-engineering/) early.

Skip a band and you'll rebuild your [AI control plane](https://atlan.com/know/ai-control-plane/) later instead of now.

---

## Should you stand up the LLM gateway or the model and agent registry first?

External sources argue gateway-first; Atlan's own published build-order research argues context-and-governance-first. This is a genuine, unresolved debate, and this section adjudicates it honestly rather than asserting one side.

Gateway-first is the more commonly documented sequence among practitioners. [Portkey's 2026 buyers' guide](https://portkey.ai/buyers-guide/leading-llm-gateway-platforms) treats the gateway as the natural first stop for routing, rate limits, and cost control (compare [LiteLLM versus Portkey versus Bedrock Gateway](https://atlan.com/know/litellm-vs-portkey-vs-bedrock-gateway/) and [manage multiple LLM providers at scale](https://atlan.com/know/manage-multiple-llm-providers-scale/)). Atlan's own build-order research sequences differently: scope and governance, then the data and context foundation, then model choice, then the [AI gateway](https://atlan.com/know/what-is-ai-gateway-llm-gateway/), then orchestration, then observability, flagging context as the most commonly skipped layer.

Here's the honest complication: Atlan's customer conversations show a recurring pattern, a directional signal from sales calls, not a published statistic, of teams standing up a gateway first and only later confronting fragmented context and registry sprawl. That's the risk of gateway-first, not proof context-first is uncontested. Teams leading with the gateway get [LLM cost management](https://atlan.com/know/llm-cost-management-enterprise/) fast and a working [model and agent registry](https://atlan.com/know/what-is-ai-registry/) later, once sprawl forces the issue.

| Consideration | Gateway-first argument | Context-first argument (Atlan POV) | What actually happens in practice |
|---|---|---|---|
| Time to first value | Fastest way to get cost control and routing live | Slower to first tool, avoids rework later | Most teams observed choose gateway-first |
| Risk if skipped | Context/registry sprawl discovered later, harder to retrofit | Gateway choices made without governance context to constrain them | Gateway-first is common, sprawl surfaces later |
| Who should own it first | Whoever ships cost/routing control fastest | Whoever owns shared context and policy | Genuinely contested, this guide states both sides |

Both sides are defensible, and the evidence shows most teams pick speed over sequencing discipline. Naming who owns shared context and policy costs nothing, it's a decision, not a build, and it's cheap insurance against the harder retrofit: undoing governance gaps under a gateway already in production. The honest answer isn't which comes first in theory, it's which retrofit you're willing to accept later.

---

## What's the difference between an AI platform team and an AI center of excellence, and which do you build first?

A CoE without a platform team sets policy nobody can enforce; a platform team without a CoE builds tooling with no policy to encode. Neither has to form first, but a minimal CoE makes the platform team's tooling worth standardizing. See [how to build an AI Center of Excellence](https://atlan.com/know/ai-agent/ai-agent-governance/how-to-build-ai-center-of-excellence/) and [AI agent governance](https://atlan.com/know/ai-agent-governance/) for the full charter.

Atlan's angle: policy context surfaced where tools enforce it, [access control](https://atlan.com/know/ai-agent-access-control/), PII propagation along lineage, makes a CoE's decisions executable; without it, [securing multi-agent systems](https://atlan.com/know/ai-agent/ai-agent-governance/how-to-secure-multi-agent-systems-enterprise/) becomes a manual, skippable step.

---

## What are the most common mistakes in AI platform team build order?

Ordering errors: doing the right thing, wrong sequence.

| Mistake | Why it happens | What to do instead |
|---|---|---|
| Researchers before data infrastructure | Feels like the "AI" part | Hire for the data foundation |
| Governance tooling before a validated use case | Feels like the responsible move | Validate one use case in production |
| Premature tool standardization | Pressure to pick a stack early | Shopify's non-standardization + shared gateway is valid |

["Six Sins of Platform Teams"](https://serce.me/posts/2025-01-07-six-sins-of-platform-teams) warns against structuring a team around a tool, not the problem it solves; build on evidence, a real, running [AI agent harness](https://atlan.com/know/how-to-build-ai-agent-harness/), not a hypothetical [AI risk management](https://atlan.com/know/ai-readiness/ai-risk-management/) program nobody checks.

---

  Context Maturity Assessment
  Score your own tooling and governance maturity before you commit to a hiring order or a tooling stack.
  Assess Your Maturity

---

## What do you build next once you're past the 20-50 engineer band?

Past this band: federate tooling ownership across business units, mature the CoE-platform-team boundary, and revisit the gateway-registry order with real usage data, not projections. [Multi-agent system orchestration](https://atlan.com/know/multi-agent-system-orchestration/) and [AI agent observability](https://atlan.com/know/ai-agent-observability/) become live concerns your incident reviews will need.

See [how to standardize AI tooling across business units](https://atlan.com/know/how-to-standardize-ai-tooling-across-business-units/) for the federated model, the sibling playbook for steady-state ratios and SLAs.

---

## Real stories from real customers: context as the shared language behind the build order

Two enterprises that scaled past ad hoc describe the same need this guide keeps returning to: a shared, governed source of context underneath whatever gateway, registry, or team structure they chose.



      "We're excited to build the future of AI governance with Atlan. All of the work that we did to get to a shared language at Workday can be leveraged by AI via Atlan's MCP server…as part of Atlan's AI Labs, we're co-building the semantic layer that AI needs with new constructs, like context products."


      — Joe DosSantos, VP of Enterprise Data & Analytics, Workday




    Watch Now




      "Atlan is much more than a catalog of catalogs. It's more of a context operating system…Atlan enabled us to easily activate metadata for everything from discovery in the marketplace to AI governance to data quality to an MCP server delivering context to AI models."


      — Sridher Arumugham, Chief Data & Analytics Officer, DigiKey




    Watch Now


  AI Agent Context Readiness Checklist
  Run through the readiness checklist before you lock a hiring order or a tooling stack.
  Check Your Readiness

---

## Why the context foundation matters more than which fork you choose

Whichever fork and band describe your organization, every path here eventually needs a shared, governed context layer underneath the gateway, registry, and observability tooling: Atlan's customer conversations show gateways going live first, sprawl discovered later.

Atlan's [Enterprise Data Graph](https://atlan.com/know/what-is-a-context-graph/) gives lineage, ownership, and policy in one governed graph agents can query, so every tool the platform team runs draws from the same source instead of rebuilding it per tool. See [what is the enterprise context layer](https://atlan.com/know/what-is-the-enterprise-context-layer/) if you're earlier in the process. It governs the payload beneath the gateway, not the gateway itself.

There is no single build order: the fork and band determine the sequence, and the tooling order remains genuinely unresolved.

  Book a Demo

---

## Sources

1. [CNCF and SlashData Report Finds Platform Engineering Tools Maturing as Organizations Prepare for AI-Driven Infrastructure, CNCF, 2026](https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-platform-engineering-tools-maturing-as-organizations-prepare-for-ai-driven-infrastructure/)
2. [How to Build an ML Platform: Lessons from 10 Companies, Evidently AI, 2026](https://www.evidentlyai.com/blog/how-to-build-ml-platform)
3. [Two Paths to Building an AI/ML Team: The Modern Hiring Strategy, KAIRI (Medium), 2026](https://medium.com/kairi-ai/two-paths-to-building-an-ai-ml-team-the-modern-hiring-strategy-45066e99b5c4)
4. [When to Invest in a Platform Team, PlatformEngineeringCost.com, 2026](https://platformengineeringcost.com/when-to-invest)
5. [Agentic AI Goes Mainstream in the Enterprise, but 94% Raise Concern About Sprawl, OutSystems Research, 2026](https://www.businesswire.com/news/home/20260407749542/en/Agentic-AI-Goes-Mainstream-in-the-Enterprise-but-94-Raise-Concern-About-Sprawl-OutSystems-Research-Finds)
6. [Inside AI-Pilled Engineering Teams: Five Lessons for Scaling Without Losing the Plot, Bessemer Venture Partners, 2026](https://www.bvp.com/atlas/inside-ai-pilled-engineering-teams-five-lessons-for-scaling-without-losing-the-plot)
7. [Meet Michelangelo: Uber's Machine Learning Platform, Uber Engineering, 2017](https://www.uber.com/us/en/blog/michelangelo-machine-learning-platform/)
8. [Best AI Gateway Solutions, Portkey, 2026](https://portkey.ai/buyers-guide/leading-llm-gateway-platforms)
9. [Ask HN: How Do You Organize a Platform Team?, Hacker News, 2022](https://news.ycombinator.com/item?id=32823551)
10. [Six Sins of Platform Teams, serce.me, 2025](https://serce.me/posts/2025-01-07-six-sins-of-platform-teams)
11. [How to Hire Your First AI Engineer, 2026, KORE1](https://www.kore1.com/hire-first-ai-engineer/)
12. [What Is an AI Platform Team?, Credal AI, 2026](https://www.credal.ai/blog/what-is-an-ai-platform-team)

---

## FAQs about building an AI platform team from scratch

### 1. What is an AI platform team, and how is it different from an AI development team?

An AI platform team builds and runs the shared infrastructure other teams depend on: the gateway, the model and agent registry, and observability. An AI development team builds the AI products and models themselves, data scientists and ML engineers. Most search results answer the development-team question, not this one.

### 2. How much does it cost to build an AI platform team from scratch?

Cost is not a single figure. It scales with your headcount band, roughly 30 product engineers is the point where a dedicated function typically starts paying back, and with which tools you choose first. The gateway-versus-context tooling order changes total cost sequencing more than any single tool's price tag.

### 3. How big should an AI platform team be at each stage?

Seed to 10 engineers typically needs one to two platform-adjacent hires, scaling in three stages: 10 to 20, then 20 to 50 or more. For specific staffing ratios at multi-business-unit scale, see the AI platform team playbook this guide defers to.

### 4. Who should you hire first if AI is your core product vs. if it's supporting an existing one?

If AI is your core product, hire an ML engineer first to validate the core capability actually works, then a data engineer, then an MLOps engineer. If AI supports an existing product, the order reverses: data engineer first, then MLOps engineer, then ML engineer, since that product's data reliability comes first.

### 5. When should you build a dedicated AI platform team instead of embedding AI engineers in product teams?

Around 30 product engineers, once at least two of five friction signals appear: slow deploys, long onboarding, frequent AI-related incidents, seniors spending significant time on infrastructure help, or slow new-service spin-up. Below that threshold, embedding AI engineers in product teams is usually more efficient.

### 6. What if our organization doesn't fit either the "AI-as-core-product" or "AI-as-bolt-on" pattern cleanly?

Default to whichever fork matches your existing data infrastructure maturity: if a pipeline already moves production data reliably, treat yourself as the bolt-on fork; if not, treat yourself as the core-product fork. No source in this research supports ignoring that check.