Skip to main content

Four Architectures That Make AI Work

Decisions, meaning, proof, people. The four systems that decide whether it delivers.

Colin Hardie

Colin Hardie

Chief Data Architect

September 21, 2026·20 min read

Gartner’s 2026 CIO and Technology Executive Survey tells a familiar story: just 17% of organisations have deployed AI agents, while more than 60% expect to within two years. That is a lot of flight plans filed and not a lot of planes in the air. Agentic AI has become what Dan Ariely once said of Big Data:

  • Everyone talks about it.
  • Nobody really knows how to do it.
  • Everyone thinks everyone else is doing it.
  • So everyone claims they are doing it.
Intent is not the same as a journey to production. Many of the organisations in that gap will end up where the [42% who shelved their initiatives](https://www.ciodive.com/news/AI-project-fail-data-SPGlobal/742590/) ended up, hoping nobody would ask about them at the next board meeting.

But when you really look at the reasons why things haven’t worked out, for GenAI or its Agentic progeny, you begin to see the same story recurring: the model isn’t the problem, everything else is.

An AI model is a component, but components do not produce value. Systems produce value, and a component’s contribution is determined almost entirely by the system it sits inside. The same model that transforms one organisation is an expensive ornament in another.

Moreover, AI does not automatically create new organisational capability. It exposes and magnifies what is already there. Every previous technology left room for poor design to hide:

  • The analyst who compensated for ambiguous definitions
  • The workaround that papered over a broken process
  • The human patience that absorbed what the system could not
AI removes those hiding places. Point an agent at a business that does not agree with itself about what its own words mean, and it will not fail politely; it will scale the disagreement at machine speed.

Boards keep asking whether the organisation is ready for AI. The better question is whether the organisation itself works, because AI is about to demonstrate it either way.

The discipline of designing systems so that the parts add up to something useful has a name: architecture. Not architecture in the narrow sense of boxes and arrows on a reference diagram, but in the older, broader sense: the deliberate design and arrangement of parts, and the relationships between them, in service of a purpose.

Everything that makes AI work is architecture. The architecture of decisions. The architecture of meaning. The architecture of proof. The architecture of people. Get these four right and a competent model has a real chance of delivering value. Get them wrong and no model can save you, no matter how mythical or fabulous it may be.


The Architecture of Decisions

The first architecture is the deliberate design of the path from strategic objective to operational outcome: a closed loop that connects what the organisation wants to achieve with the decisions and actions where that achievement is won or lost, and with the capabilities that support them.

Somewhere in the organisation, somebody has to decide something, and then act on it, and the entire apparatus of data, analytics and AI exists to make that decision better and that action faster. When IBM’s Hans Peter Luhn coined the modern usage of business intelligence in 1958, his stated objective was information that could guide action towards a desired goal. The industry then spent 60+ years losing the plot.

Generations of warehouses, marts, lakes, dashboards and self-service platforms promised insight, and stopped there: an artefact that a human might look at, which might inform a decision, which in turn might change something, rather than the closed loop that delivered the outcome that was actually needed.

If decisions and actions are the unit of value, the first architectural act is choosing which ones matter. Too many data and AI strategies are technology selections masquerading as strategy. “We need a lakehouse.” “Let’s implement governance.” “AI is the future.” These are shopping lists, usually delivered by an executive freshly returned from a conference, glowing with conviction and carrying a tote bag.

Real strategy asks a tougher question: where in the business are your objectives won or lost, and what are the decisions at that point?

Below is the framework I use for this, with the decision at its centre.

Strategy framework diagram. Strategic objectives decompose into decisions and actions, then a prioritised portfolio of initiatives, then foundational capabilities, with measurement running as a vertical spine through all four layers and the outcome feeding back to the objective.

Measurement runs as a vertical spine, and the outcome feeds back to the objective. | Source: Colin Hardie

Strategic objectives decompose directly into the decisions and actions where those objectives are won or lost: which customers do we intervene with, which trades do we approve, which suppliers do we renew. Each decision implies a process, and each process must be redesigned around a closed loop so that intelligence terminates in action and the outcome feeds back into the next cycle. Those decisions translate into a prioritised portfolio of initiatives, each gated by feasibility, impact and risk, each required to name the decision it serves. Underneath are the foundational capabilities those initiatives require, built incrementally rather than as a precondition. Some problems need AI. Some need a well-configured report, a cleaner process, or just a conversation. Choosing the boring option that works is not a failure of ambition; it is what having a strategy is for.

Measurement runs as a vertical spine through all four layers, connecting the outcome of each decision back to the objective it was supposed to serve. This is what turns strategy from a document into a feedback loop.

The process redesign is not a separate step. It is inherent in naming the decision. McKinsey’s 2025 State of AI survey tested 25 organisational attributes and found that fundamental workflow redesign had the single largest effect on EBIT impact from AI. In the years since, high performers have been roughly three times as likely as everyone else to have fundamentally redesigned workflows as part of their AI efforts. An agent can observe, orient, decide, and act within a workflow. That makes process design more consequential, not less.

An agent dropped into an undesigned process is not a better dashboard. It is a dashboard that can act, which is worse.

Consider retail banking. The bolt-on version buys a churn model and emails the retention team a ranked list every Monday, where it competes with everything else in the inbox and mostly goes unread. The redesigned version starts from the decision: which customers do we intervene with this week, and with what offer? It surfaces the customers, pushes them into the queue the agent already works from, holds back a control group, and feeds the result into next week’s model. Same data, same team, same technology, different process. One produces an artefact. The other produces retained customers.

The first architecture, then, is a single thread from objective to outcome: identify the decisions where your objectives are won or lost, build the initiatives that support them, redesign the process so the loop closes, and measure the result against what you promised. None of this involves the model. All of it determines whether the model matters.


The Architecture of Meaning

The second architecture is the deliberate design of shared understanding: a governed path that ensures the system and the organisation agree on what their words mean, how they are computed, and where to go when the answer is ambiguous.

The perennial failure of self-service analytics was never access. Every generation of tooling made data easier to reach, and every generation foundered on the same rock: the consumer did not carry the context needed to interpret it. Which table is trusted? Which revenue definition applies? Is the fraud filter included? A human analyst compensates for that ambiguity; they know the terrain, and when they do not, they pick up the phone and ask. An agent does not compensate. It selects the most likely interpretation and proceeds, returning an answer that is plausible, well-formatted, and wrong.

Anthropic found that without structured context, their own analytics agent never exceeded 21% accuracy. After it encoded analytical workflows and business context as governed, reusable skills, aggregate accuracy exceeded 95%. Same model, same data. The difference was the context provided. If a single number explains why AI makes data foundations more important rather than less, it is that one.

The context an agent needs is not mystical, and it is narrower than the full enterprise context layer: Prukalpa has drawn that broader map in What an Enterprise Context Layer Actually Is. The smallest version that stops an agent guessing is three layers of shared understanding, represented by three questions:

  • What do we mean when we say this word?
  • How do we compute it?
  • Where do I go to answer this question, and what do I do when the question is ambiguous?
Three questions, three layers of context. Definition, Computation, and Procedural Guidance. The governed path. A question, what was EMEA revenue, passes down through Layer 1 definitions, Layer 2 computation and Layer 3 procedural guidance before reaching the data. A direct route from the question to the raw data is crossed out and marked blocked.

Every query traverses the governed path. The direct route to raw data is blocked. | Source: Colin Hardie

Together they form a governed path that every query must traverse, end to end, so that a dashboard, a notebook and an agent asking for the same metric all receive the same number. The alternative is the ritual in which the CFO and the Head of Sales argue over whose revenue figure is right while the data team updates their CVs under the table.

The structural enforcement matters as much as the content. If agents can bypass one layer and query raw tables directly, they will, because raw tables are faster and require no configuration. The architectural decision that all queries must pass through the governed layers before touching the data does more for meaning than any amount of documentation, because documentation relies on someone reading it and architecture does not. This is not governance in the compliance sense. It is plumbing: the system is built so that the wrong path is harder to take than the right one.

Fixing this does not require a two-year metadata programme. It begins with a basic version of all three layers and the discipline to grow them together. The definitions can start as a glossary of thirty to fifty terms for one domain, maintained in markdown, each term linked to the physical table or column it describes. The computation layer can start just as simply: a shared set of metric definitions, each specifying the exact logic that produces the number, maintained alongside the data transformations so there is one canonical calculation for “EMEA revenue” rather than seven spreadsheets arriving at seven answers. Procedural guidance can begin as a reference document per domain that tells the agent which definitions to use for which questions, which tables are trusted, and what gotchas to watch for: the kind of knowledge that a senior analyst carries in their head and that nobody has ever written down.

At that basic stage the three layers are small enough to live in the same repository as your data transformations, which means a developer changing a revenue calculation cannot avoid seeing the glossary term and the metric definition in the same review. That coupling is the architectural principle: things that must change together must live together. At the markdown stage the coupling is literal: same directory, same pull request. As the layers grow, the coupling becomes engineered: version-controlled schemas deployed through CI pipelines so that changing the data model still triggers a review of the definition, even though the systems are now separate.

What you add next is a capability question, not a rung on a ladder. Controlled vocabularies, taxonomies, ontologies and knowledge graphs are not sequential stages that each mature into the next; they solve different representation problems. You add formal vocabulary, relationship modelling, inference or graph structure when a use case demands that specific capability, and not before. Which is which, and when each earns its place, is a discipline in its own right, and one Jessica Talisman mapped for C&C in Ontologies, Context Graphs, and Semantic Layers: What AI Actually Needs in 2026. The architectural question is never whether you need a knowledge graph. It is when your current level of semantic precision becomes the bottleneck. Many organisations may never need the full stack. All of them need more than they have.

But a system that can answer questions correctly is not yet a system anyone should rely on. That requires something else.


The Architecture of Proof

The first two architectures design a system capable of producing value. The third establishes whether it actually does, not as a milestone you pass on the way to production, but as a continuous property that requires its own architecture: a structure for evidence, a structure for governance, and a structure for integrity.

Evidence

Without a structure for evidence, nothing else can be evaluated. The measurement spine that runs through the strategy framework does not stop at deployment; it extends into production, where the question shifts from “what outcome will this deliver?” to “did it?”

IBM’s 2026 study of 2,000 technology executives found that 85% still lack full visibility into real-time AI spend and 84% have not operationalised AI financial management. Without a baseline measured before implementation, there is no way to attribute impact and no rational basis for the next investment decision. The organisations compounding AI value define the metric, capture the baseline, and agree the target before a line of code is written. That is the difference between a leader who can say “this initiative delivered a 14% reduction in customer churn” and one who says “we deployed a model and things seem better, probably.” One of them gets the next round of funding.

Governance

A lot of what passes for governance in modern enterprises is theatre. Committees, gates and documents that slow everything while protecting nothing, which is how shadow systems end up flourishing in every department while the official governance forum meets quarterly to review a slide deck that nobody updates. The teams doing the most dangerous, ungoverned work are the ones who found the official process so painful they stopped using it.

The route to production has to be the easiest route available, or people will route around it. That is not a concession to speed, it is the only design that actually governs anything. What is needed are risk-differentiated tracks: pre-approved patterns that mean the majority of use cases never reach a committee and most stage gates become thirty-minute conversations rather than six-week reviews. Lightweight controls for the dashboard showing last month’s sales figures. Comprehensive review for the model that autonomously approves insurance claims. Match the cost of oversight to the cost of being wrong.

Integrity

Production integrity is where most organisations stop paying attention. The lifecycle of AI does not stop at launch. It is a living system, and everything it depends on is perishable. The model that tested well before launch will drift as the world it was trained on changes beneath it. Without ongoing evaluation, nobody notices until the decisions built on the model’s outputs have already gone wrong. Provenance is the minimum defence: every answer should carry a record of where the data came from, how current it is, and which definition was applied. In practice, few systems provide it. How that decay unfolds after go-live, and why nobody ends up accountable for catching it, is the subject of Ankit Anand’s Production AI & The False Finish Line.

The third architecture, then, is proof as a continuous system property: evidence that the investment is justified, governance that makes the governed route the path of least resistance, and integrity that keeps a production system honest after the launch fanfare has faded and the implementation team has moved on to the next shiny thing.


The Architecture of People

The fourth architecture is the deliberate design of the organisational forces that determine whether the other three survive contact with reality.

Everything so far could, in principle, be drawn on a whiteboard. The fourth architecture cannot, because the components have opinions, ambitions, professional identities, and the capacity to resist.

Four forces govern organisational architecture, three borrowed from system design and one that exists only because the components are human. Coupling describes how tightly bound your teams are to each other: too tight and nothing moves independently, too loose and three business units build three different customer models using three different definitions. Cohesion describes whether each team has a clear, internally consistent purpose. Interfaces describe how teams interact: defined contracts with agreed accountability, or informal relationships that collapse when someone leaves.

Four forces of organisational architecture. Coupling, cohesion and interfaces transfer directly from system design and sit above a dividing line. Incentives sits below, exists only because the components are human, and determines whether the other three function as designed. A footer reads Conway's Law, your systems mirror your org chart, and vice versa.

Three forces transfer from system design. The fourth exists only because the components are human. | Source: Colin Hardie

These three follow the same rules whether you are designing systems or organisations. Conway’s Law reads as a diagnostic as well as a prediction: if your data estate is fragmented across incompatible platforms, look at your org chart, because the fragmentation is not a technology problem, it is an organisational one that expressed itself in technology.

There is also a fourth force, and it is the one that makes organisational architecture harder than system design: incentives. Software components do what they are designed to do. They do not form allegiances, compete for budget, or sit in a restructuring workshop thinking “this threatens my influence” while nodding along and saying “sounds great.” People do all of these things, and incentives determine whether they do what the operating model says or what their performance review rewards. Organisations can launch the platform, hold the kickoff, print the branded mugs, and then watch the whole thing gather dust because not a single KPI in anyone’s review mentions collaboration. In my experience, the mugs outlast the initiative. Incentives are not a change management afterthought. They are the architectural force that determines whether the other three function as designed. Why incentives decide whether AI foundations hold, and what organisational self-awareness is needed to see them clearly, is the argument Gaurav Ramesh makes in The Human Elements of the AI Foundations.

How you resolve these four forces is the actual work of operating model design. Every operating model choice is a power allocation decision wearing a structural disguise. In practice, most organisations end up somewhere between fully centralised and fully distributed, but “somewhere in the middle” is not a design until you specify exactly what is centralised, what is distributed, and what the interface contract looks like. What matters more than the reporting line is the operating relationship: a data function without a seat when strategy is discussed is an order-taker, regardless of how talented its people are. The question is not who do you report to. The question is are you in the room when it matters.


The Work

None of the four architectures is new. Every one of them predates the current AI hype cycle by years or decades. That is not a coincidence. The organisations compounding value from AI are not doing anything exotic. They are doing the unglamorous work that was always available. None of it will demo well, and no vendor can sell the whole of it to you. Platforms can provide the infrastructure; they cannot choose your decisions, settle what your words mean, redesign your workflows, or realign your incentives. And that is the answer to the obvious objection. If none of this is new, why has nobody done it? Because the moment a true idea becomes something you can buy, organisations buy it and skip the work. The work is the point. It is slower than the hype cycle, cheaper than the failures, and it compounds, which the hype never does.

None of which means the model never matters. Sometimes it is the binding constraint: a task sitting at the edge of what current models do reliably, where no amount of context design closes the gap and the honest answer is to wait for the next generation or not to build at all. A narrow, low-risk use case does not need all four architectures at equal maturity either. The four architectures decide the outcome when the model is good enough for the job. Judging whether it is good enough is the first architectural decision you make.

The models will keep improving whether you do this work or not, and they will keep improving for your competitors at the same rate, often from the same vendor, because a commodity is available to everyone by definition. Prukalpa made that argument in If Intelligence Is Abundant, What Is the Moat?. The abstract version is settled; what is left is the work itself. What will differ is everything around them: which decisions you chose, what your data means, whether you proved it works well enough to act on, and whether you enabled your people to work in a way that gets the best out of them. When the examination comes, and it always comes, the model was never the thing being tested.

You were.



The Cats of Context & Chaos

Three-panel cartoon titled AI doesn't fix the process, it hits copy. An orange tabby holding a folder marked current process tells a robot labelled AI agent to make it autonomous. The robot feeds the folder into a photocopier marked automate, which spits out sheets marked automated decision. In the final panel both stand cheering on a mountain of identical automated decision sheets while a grey professor cat in glasses holds up two of them and says you made the broken process faster.

You made the broken process faster. | Source: Context and Chaos




Note: Contributions reflect their authors’ views, not ours. We curate them for wider access and discussion, and vet every submission for quality and relevance. Information-first always, with no promotions, paid or otherwise.




About Context & Chaos

Context & Chaos isn’t just a newsletter. It’s shared community space where practitioners, builders, and thinkers come together to share stories, lessons, and ideas about what truly matters in the world of data and AI: context engineering, governance, architecture, discovery, and the human side of doing meaningful work.

Our goal is simple, to create a space that cuts through the noise and celebrates the people behind the amazing things that are happening in the data & AI domain.

Whether you’re solving messy problems, experimenting with AI, or figuring out how to make data more human, Context & Chaos is your place to learn, reflect, and connect.


Got something on your mind? We’d love to hear from you.

Share this article