How to Build an AI Platform Team From Scratch (And in What Order)

Emily Winks, Data Governance Expert, Atlan
Data Governance Expert
Updated:07/27/2026
|
Published:07/27/2026
13 min read

Key takeaways

  • There is no universal build order: hire order reverses depending on whether AI is your core product or a bolt-on.
  • Headcount-band triggers, not calendar months, signal when to add the next hire or tool (seed-10, 10-20, 20-50+).
  • Gateway-first is the common practice, but Atlan customer data shows it often causes context sprawl later.
  • Validate one AI use case in production before hiring platform staff or standing up governance tooling.

What order should you build an AI platform team in?

There is no single build order for an AI platform team. The correct first hire and first tool depend on whether AI is your organization's core product or an add-on to an existing one, and the order reverses between the two. Most teams follow a scenario fork crossed with three headcount bands (seed-10, 10-20, 20-50+ engineers) rather than a fixed calendar.

What determines your build order

  • The fork: whether AI is your core product or a bolt-on decides your first hire
  • Headcount bands: seed-10, 10-20, and 20-50+ each carry a different priority
  • The tooling debate: gateway-first versus context-first, argued honestly, not resolved with one answer
  • The readiness bar: a validated use case and roughly 30 product engineers before you dedicate a team at all

Want the tooling order mapped out?

Get the AI Context Stack

Most guides on this topic answer the wrong question: they describe hiring the AI development team that builds models, not the AI platform team that runs the shared gateway, registry, and observability infrastructure, per Credal AI’s definition. Only 19% of AI-building orgs have one, per a 2026 CNCF and SlashData survey.

This is a sequencing tool, not a maturity scorecard. The sibling playbook covers org chart, ratios, and SLAs.

Difficulty: Intermediate. Time required: Months to over a year. Frameworks used: Headcount-band staging (Bessemer), five-friction-signal model (PlatformEngineeringCost.com), core-product-vs-bolt-on fork (KAIRI). Prerequisites: One validated AI use case in production, executive sponsorship, and an honest read on core-vs-supporting.


How does the AI platform team build-order decision tree work?

Permalink to “How does the AI platform team build-order decision tree work?”

This section replaces the “Month 1 / Month 2” calendar with a two-axis decision tree: a scenario fork crossed with headcount-band triggers. Evidently AI’s review of ten ML platform builds found no universal best order; KAIRI’s 2026 hiring-strategy analysis shows it flips entirely: ML engineer first if core product, data engineer first if bolt-on.

Is AI your organization's core product, or a bolt-on? Core product First hire: ML engineer Bolt-on First hire: data engineer Seed-10: ML engineer 10-20: + data engineer 20-50+: + MLOps engineer Seed-10: data engineer 10-20: + MLOps engineer 20-50+: + ML engineer Both forks converge by 20-50+ engineers

A scenario fork crossed with three headcount bands. Source: Atlan, synthesized from Bessemer and KAIRI research.


What has to be true before you build a dedicated AI platform team?

Permalink to “What has to be true before you build a dedicated AI platform team?”

Clear a readiness bar first: a headcount threshold, five friction signals, and one validated AI use case already in production, ruling out a dedicated team or governance tooling ahead of the evidence, this research’s most common mistake.

The ROI threshold clusters around 30 product engineers, per PlatformEngineeringCost.com’s 2026 analysis: below it, a gateway costs more than it saves; above it, engineers rebuild the same data infrastructure for AI independently.

Signal Threshold What it means
Deploy cycle for AI changes Over 30 minutes Manual review is the bottleneck
Onboarding to first AI contribution Over two weeks No shared scaffolding exists
AI-related infra incidents 2+ per week Ad hoc infra is now a risk
Senior time on AI infra help 30%+ of their week Expensive people doing platform work informally
New AI service spin-up time Over a day No standard pattern to fork from

OutSystems Research (2026) found 94% of enterprises cite AI sprawl as a concern.


Inside Atlan AI Labs & The 5x Accuracy Factor

See what changes when the context layer feeds your models, before you finalize your platform team's tooling stack.

Get the Ebook

What order do you hire and build in?

Permalink to “What order do you hire and build in?”

A scenario fork at the root, three headcount bands, and a trigger signal for the next.

Band Core-product fork Bolt-on fork Next-band signal
Seed to 10 engineers ML → data → MLOps engineer Data → MLOps → ML engineer Tooling decisions become expensive
10 to 20 engineers + technical leader (architecture calls) + technical leader (architecture calls) Shifts from “tool” to “people/architecture”
20 to 50+ engineers Diagnose people vs. architecture first Diagnose people vs. architecture first Ratios and CoE boundary (sibling playbook)

The fork: is AI your organization’s core product, or are you retrofitting it onto an existing one?

Permalink to “The fork: is AI your organization’s core product, or are you retrofitting it onto an existing one?”

If AI is your core product, KAIRI’s hiring-sequence research supports ML engineer first; if it’s a bolt-on, data engineer first. Can’t decide? Check whether a pipeline already moves production data reliably.

Seed to 10 engineers: hire for velocity, not architecture

Permalink to “Seed to 10 engineers: hire for velocity, not architecture”

Make your first one to two hires per the fork above, hands-on coders who ship solo, not architects, per Jessica Popp (Bessemer) and KORE1. DoorDash started with one prediction service; Uber’s Michelangelo platform prioritized training and serving instead. Shopify didn’t standardize its tool layer but still centralized cost control via a shared gateway.

10 to 20 engineers: the architecture-lock inflection point

Permalink to “10 to 20 engineers: the architecture-lock inflection point”

Architectural decisions become permanent here, requiring experienced leadership over more hands-on coders.

20 to 50+ engineers: diagnose the constraint before you hire more

Permalink to “20 to 50+ engineers: diagnose the constraint before you hire more”

Figure out whether the bottleneck is people or architecture; the sibling playbook’s roles, ratios, and AI Center of Excellence boundary are the reference.

Pro tip: Still resolving gateway-versus-registry past the 20-50 band? That’s a signal you under-invested in the context and governance foundation early.

Skip a band and you’ll rebuild your AI control plane later instead of now.


Should you stand up the LLM gateway or the model and agent registry first?

Permalink to “Should you stand up the LLM gateway or the model and agent registry first?”

External sources argue gateway-first; Atlan’s own published build-order research argues context-and-governance-first. This is a genuine, unresolved debate, and this section adjudicates it honestly rather than asserting one side.

Gateway-first is the more commonly documented sequence among practitioners. Portkey’s 2026 buyers’ guide treats the gateway as the natural first stop for routing, rate limits, and cost control (compare LiteLLM versus Portkey versus Bedrock Gateway and manage multiple LLM providers at scale). Atlan’s own build-order research sequences differently: scope and governance, then the data and context foundation, then model choice, then the AI gateway, then orchestration, then observability, flagging context as the most commonly skipped layer.

Here’s the honest complication: Atlan’s customer conversations show a recurring pattern, a directional signal from sales calls, not a published statistic, of teams standing up a gateway first and only later confronting fragmented context and registry sprawl. That’s the risk of gateway-first, not proof context-first is uncontested. Teams leading with the gateway get LLM cost management fast and a working model and agent registry later, once sprawl forces the issue.

Consideration Gateway-first argument Context-first argument (Atlan POV) What actually happens in practice
Time to first value Fastest way to get cost control and routing live Slower to first tool, avoids rework later Most teams observed choose gateway-first
Risk if skipped Context/registry sprawl discovered later, harder to retrofit Gateway choices made without governance context to constrain them Gateway-first is common, sprawl surfaces later
Who should own it first Whoever ships cost/routing control fastest Whoever owns shared context and policy Genuinely contested, this guide states both sides

Both sides are defensible, and the evidence shows most teams pick speed over sequencing discipline. Naming who owns shared context and policy costs nothing, it’s a decision, not a build, and it’s cheap insurance against the harder retrofit: undoing governance gaps under a gateway already in production. The honest answer isn’t which comes first in theory, it’s which retrofit you’re willing to accept later.


What’s the difference between an AI platform team and an AI center of excellence, and which do you build first?

Permalink to “What’s the difference between an AI platform team and an AI center of excellence, and which do you build first?”

A CoE without a platform team sets policy nobody can enforce; a platform team without a CoE builds tooling with no policy to encode. Neither has to form first, but a minimal CoE makes the platform team’s tooling worth standardizing. See how to build an AI Center of Excellence and AI agent governance for the full charter.

Atlan’s angle: policy context surfaced where tools enforce it, access control, PII propagation along lineage, makes a CoE’s decisions executable; without it, securing multi-agent systems becomes a manual, skippable step.


What are the most common mistakes in AI platform team build order?

Permalink to “What are the most common mistakes in AI platform team build order?”

Ordering errors: doing the right thing, wrong sequence.

Mistake Why it happens What to do instead
Researchers before data infrastructure Feels like the “AI” part Hire for the data foundation
Governance tooling before a validated use case Feels like the responsible move Validate one use case in production
Premature tool standardization Pressure to pick a stack early Shopify’s non-standardization + shared gateway is valid

“Six Sins of Platform Teams” warns against structuring a team around a tool, not the problem it solves; build on evidence, a real, running AI agent harness, not a hypothetical AI risk management program nobody checks.


Context Maturity Assessment

Score your own tooling and governance maturity before you commit to a hiring order or a tooling stack.

Assess Your Maturity

What do you build next once you’re past the 20-50 engineer band?

Permalink to “What do you build next once you’re past the 20-50 engineer band?”

Past this band: federate tooling ownership across business units, mature the CoE-platform-team boundary, and revisit the gateway-registry order with real usage data, not projections. Multi-agent system orchestration and AI agent observability become live concerns your incident reviews will need.

See how to standardize AI tooling across business units for the federated model, the sibling playbook for steady-state ratios and SLAs.


Real stories from real customers: context as the shared language behind the build order

Permalink to “Real stories from real customers: context as the shared language behind the build order”

Two enterprises that scaled past ad hoc describe the same need this guide keeps returning to: a shared, governed source of context underneath whatever gateway, registry, or team structure they chose.

"We're excited to build the future of AI governance with Atlan. All of the work that we did to get to a shared language at Workday can be leveraged by AI via Atlan's MCP server…as part of Atlan's AI Labs, we're co-building the semantic layer that AI needs with new constructs, like context products."

— Joe DosSantos, VP of Enterprise Data & Analytics, Workday

"Atlan is much more than a catalog of catalogs. It's more of a context operating system…Atlan enabled us to easily activate metadata for everything from discovery in the marketplace to AI governance to data quality to an MCP server delivering context to AI models."

— Sridher Arumugham, Chief Data & Analytics Officer, DigiKey

AI Agent Context Readiness Checklist

Run through the readiness checklist before you lock a hiring order or a tooling stack.

Check Your Readiness

Why the context foundation matters more than which fork you choose

Permalink to “Why the context foundation matters more than which fork you choose”

Whichever fork and band describe your organization, every path here eventually needs a shared, governed context layer underneath the gateway, registry, and observability tooling: Atlan’s customer conversations show gateways going live first, sprawl discovered later.

Atlan’s Enterprise Data Graph gives lineage, ownership, and policy in one governed graph agents can query, so every tool the platform team runs draws from the same source instead of rebuilding it per tool. See what is the enterprise context layer if you’re earlier in the process. It governs the payload beneath the gateway, not the gateway itself.

There is no single build order: the fork and band determine the sequence, and the tooling order remains genuinely unresolved.


Sources

Permalink to “Sources”
  1. CNCF and SlashData Report Finds Platform Engineering Tools Maturing as Organizations Prepare for AI-Driven Infrastructure, CNCF, 2026
  2. How to Build an ML Platform: Lessons from 10 Companies, Evidently AI, 2026
  3. Two Paths to Building an AI/ML Team: The Modern Hiring Strategy, KAIRI (Medium), 2026
  4. When to Invest in a Platform Team, PlatformEngineeringCost.com, 2026
  5. Agentic AI Goes Mainstream in the Enterprise, but 94% Raise Concern About Sprawl, OutSystems Research, 2026
  6. Inside AI-Pilled Engineering Teams: Five Lessons for Scaling Without Losing the Plot, Bessemer Venture Partners, 2026
  7. Meet Michelangelo: Uber’s Machine Learning Platform, Uber Engineering, 2017
  8. Best AI Gateway Solutions, Portkey, 2026
  9. Ask HN: How Do You Organize a Platform Team?, Hacker News, 2022
  10. Six Sins of Platform Teams, serce.me, 2025
  11. How to Hire Your First AI Engineer, 2026, KORE1
  12. What Is an AI Platform Team?, Credal AI, 2026

FAQs about building an AI platform team from scratch

Permalink to “FAQs about building an AI platform team from scratch”

1. What is an AI platform team, and how is it different from an AI development team?

Permalink to “1. What is an AI platform team, and how is it different from an AI development team?”

An AI platform team builds and runs the shared infrastructure other teams depend on: the gateway, the model and agent registry, and observability. An AI development team builds the AI products and models themselves, data scientists and ML engineers. Most search results answer the development-team question, not this one.

2. How much does it cost to build an AI platform team from scratch?

Permalink to “2. How much does it cost to build an AI platform team from scratch?”

Cost is not a single figure. It scales with your headcount band, roughly 30 product engineers is the point where a dedicated function typically starts paying back, and with which tools you choose first. The gateway-versus-context tooling order changes total cost sequencing more than any single tool’s price tag.

3. How big should an AI platform team be at each stage?

Permalink to “3. How big should an AI platform team be at each stage?”

Seed to 10 engineers typically needs one to two platform-adjacent hires, scaling in three stages: 10 to 20, then 20 to 50 or more. For specific staffing ratios at multi-business-unit scale, see the AI platform team playbook this guide defers to.

4. Who should you hire first if AI is your core product vs. if it’s supporting an existing one?

Permalink to “4. Who should you hire first if AI is your core product vs. if it’s supporting an existing one?”

If AI is your core product, hire an ML engineer first to validate the core capability actually works, then a data engineer, then an MLOps engineer. If AI supports an existing product, the order reverses: data engineer first, then MLOps engineer, then ML engineer, since that product’s data reliability comes first.

5. When should you build a dedicated AI platform team instead of embedding AI engineers in product teams?

Permalink to “5. When should you build a dedicated AI platform team instead of embedding AI engineers in product teams?”

Around 30 product engineers, once at least two of five friction signals appear: slow deploys, long onboarding, frequent AI-related incidents, seniors spending significant time on infrastructure help, or slow new-service spin-up. Below that threshold, embedding AI engineers in product teams is usually more efficient.

6. What if our organization doesn’t fit either the “AI-as-core-product” or “AI-as-bolt-on” pattern cleanly?

Permalink to “6. What if our organization doesn’t fit either the “AI-as-core-product” or “AI-as-bolt-on” pattern cleanly?”

Default to whichever fork matches your existing data infrastructure maturity: if a pipeline already moves production data reliably, treat yourself as the bolt-on fork; if not, treat yourself as the core-product fork. No source in this research supports ignoring that check.

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI, a Leader in the Gartner Magic Quadrant for D&A Governance (2026) and the Forrester Wave for Data Governance (Q3 2025). Atlan unifies your data, business knowledge, and the meaning behind your terms into one Enterprise Data Graph that gives every team and every AI agent the trusted context they need. Trusted by Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, Elastic, and 400+ enterprises representing $10T+ in market cap.

Bridge the context gap.
Ship AI that works.

[Website env: production]