Every technology and vendor selection process is subjective. The choice of Vendor A over Vendor B is more often a personality decision than a technical one, shaped by who has the relationship, who carries the institutional memory from the last procurement cycle, and who has the standing to say “we tried that before” in a way that anyone will actually listen to.
The scoring matrix, RFP response and technical evaluation matter. But they record only what the process was designed to capture. By the time a vendor presents, prior experience, relationships and constraints have already narrowed what the organization is willing to consider.
Informal authority has always shaped those decisions and always will. The trouble starts when an agent deployed into procurement, negotiation or contract management cannot recover that judgment from the formal artifacts alone. It was configured against the scoring matrix, not against the room.
When it defers only to the documented approver, leaves out someone whose judgment the decision relies on, or ignores a constraint everyone in the process already knew, the model has not malfunctioned. It was never given the full decision-making spectrum. That is an authority failure.
I have watched it happen during a vendor selection in an earlier engagement. The documented path required formal sign-off from architecture, legal, procurement and security. Engineering was exerting significant influence too, and none of it was obvious or recorded. The selection was completed cleanly on paper. What followed was considerable technical debt, rework of the implementation, several months of delay and real disruption to business as usual.
That is the shape of the bill, and most of it is not paid in tokens. The team corrects the finding offline, the next decision of the same shape gets escalated rather than trusted, and a rollout, contract or launch waits while someone reconstructs who should have been involved. The agent also retrieves more context to compensate for the missing judgment. The largest cost is the human time it was bought to release.
In my experience a single authority misread costs one to three additional escalation cycles, each pulling four to six people for two to six hours, and at enterprise scale across hundreds of decisions that is the order of magnitude. Any risk or liability consequence of a missed authorization makes it considerably worse. I should be equally clear that this is experience rather than measurement. Putting an empirical number on it is the next test for the framework.
The governance problem and the inference bill are the same problem: resolving authority is how the agentic investment returns anything at all.
Authority is not one thing
That procedure, decision history and authority have no home in any of our systems is not a new observation here. Part of the reason is that authority is not one thing. The word covers five different objects.
There is the formal right: what policy says a role may decide. There is the system permission: what the software will allow. There is the recorded approval: whose name lands in the audit trail. There is practical influence: whose judgment the decision turns on. And there is the agentic mandate: what the agent itself is permitted to see and do.
Those are the objects the Authority Resolution Framework, or ARF, has to resolve, and it resolves them across five domains. The objects tell you what kind of authority is in question; the domains tell you where its evidence lives. The social domain identifies who holds and influences a role. The business domain defines the terms and policies involved. The process domain encodes the approval path. The machine domain determines what the system permits. The real-world domain records the state and consequence the decision acts on.
The first three leave records and can be audited separately even when they contradict one another. Practical influence leaves fragments in meeting invitations, message threads, document histories and people’s memories, and no governed record of its own. It is also the one item on the list that is not really authority at all. It is the evidence that tells you whether the other four are being exercised as intended.
The fifth object is the newest and least examined. Someone approves the agent’s objective, sets the boundary of its actions and chooses which systems it may treat as authoritative. Those decisions happen across engineering, procurement, security and vendor configuration before the agent encounters any specific decision, which makes this the least visible form of all and the most likely to go unaudited: it looks like a configuration choice rather than a governance one. Authority did not disappear when the agent arrived. It moved to the people shaping what the agent can see and do. Governing agents as compositions rather than documents asks what an agent is made of. This asks who is allowed to choose the parts.
This also extends the stale-metadata problem Amanda Darcangelo described here. Stale metadata once had a recorded state that can, in principle, be refreshed. Practical influence never entered a governed record at all. There is nothing to restore, and remediation programs never find it, because they look for things that went wrong and this never went wrong.
The new hire learns what the agent cannot see
If it was never written down, the obvious question is how anyone in the organization knows it. As Vivek Dubey argued in this newsletter, a human hire spends weeks absorbing definitions, trust hierarchies and unwritten rules. Authority is the sharpest case: a new hire notices who gets called, whose objection stops the room and whose approval is ceremonial. An agent sees policies, role records, workflow definitions, permissions and whatever examples someone placed in its context.
If those artifacts name the formal approver, the agent will use the formal approver. It will be right about the organization that was documented.
That fails in two shapes, and the second one is worse.
Only one of the two looks wrong on paper. | Source: Context and Chaos
In the first, someone exercises more influence than their formal mandate grants, and the agent excludes a person the organization informally depends on. This is the more findable of the two, because action and mandate visibly diverge and auditors are built to notice that. Findable is not the same as found: in the vendor selection above, nobody noticed until the rework did.
In the second, a formally authorized person approves the action while the substantive judgment came from somewhere else. The record is valid, the approver holds the role, and the safeguard has quietly become ceremonial. An agent can obtain the correct signature without ever obtaining the independent judgment that signature was meant to represent, and nothing in the trail registers a problem.
Score the gap, not the legitimacy
My earlier formulation treated practical influence as itself a form of authority. On reflection, the more useful reading is diagnostic: influence tells you whether the documented structure is being exercised as intended or has become ceremonial. It is also what makes divergence measurable.
The cleanest representation separates two questions. Magnitude asks how far documented and practised authority diverge, and zero means the two are aligned. Direction carries the sign. A positive reading means practice exceeds the mandate: someone outside the documented chain is substantively driving the decision. A negative reading means the mandate exceeds practice: the documented holder retains the formal authority while the judgment the control was meant to supply comes from somewhere else. Both failure shapes above can be read off that sign.
Neither field measures legitimacy. A large divergence does not tell the agent that informal practice is right, and a small one does not prove the formal path is safe. The result tells the system whether the question needs asking and which discrepancy must be investigated.
Magnitude measures the gap. Direction names the failure. | Source: Context and Chaos
Zero settles the authority question and nothing more; the action still runs under its own risk tier. A non-zero result is interpreted through a risk policy defined in advance: low-risk cases may proceed within approved bounds, while material divergence or any case involving independent review escalates or stops. Not every discrepancy carries the same consequence.
What the ontology has to hold
ARF is not an exercise in discovering who is really in charge. It asks the organization to encode, openly, how its formal structure relates to the judgment it relies on in practice. The ontology must therefore hold both sides: roles, policies, permissions and approval paths; and whose judgment a decision relies on, and why. The definition of a context layer published here by Prukalpa names the same gap: a semantic layer can tell an agent what gross margin is, but not which approval path matters in practice.
An agent using only the structural half will produce outputs that are formally correct and operationally wrong, and it will be wrong in a way an audit built around formal authorization does not detect. Without an ontology that maps both halves accurately enough for the agent to act, nobody trusts it with a decision that matters, and an agent nobody trusts never scales.
ARF makes that requirement checkable through an Authority Relation: an actor performing an action on an object under a stated mandate. The relation links the role and its holder to the governing policy, process step, system permission and relevant real-world state. A resolvesTo relationship lets a validator check whether those references exist, agree and remain current. An affirmedBy field makes business, technical and governance review part of the record rather than a separate procedural hope.
At runtime, an agent does not need to retrieve a policy PDF, CRM export and email thread and reconcile them from scratch. It can traverse from the permission or object it is about to rely on to the governing Authority Relation, then read the divergence magnitude, direction, evidence and affirmation status. The ontology does not decide whom the agent should obey. It tells the agent whether the authority it is about to rely on is sufficiently aligned and governed for it to act.
The ontology does not decide whom the agent should obey. It tells the agent whether authority is governed enough to act. | Source: Context and Chaos
Three instruments get you there. Compare documented approvers against decision logs. Work with the organization to map where influence concentrates, with the people on the map knowing they are on it and why. Interview people inside the flow to surface judgment and reliance that leave no clean trail.
I should be plain that none of the three has been empirically tested. I regard the measurement methodology as the most original thing in the framework and, at the same time, its largest open empirical question. I would rather see it replicated, criticized and extended than treated as settled.
Missing capability or missing independence?
Suppose ARF identifies someone exercising authority they were never granted. The same finding supports two opposite readings. Ask whether the informal path routes around missing capability or around independent review.
The same divergence can signal adaptation or control failure. | Source: Context and Chaos
Usually it is the first, and the vendor selection above is the clearest case I have seen. Engineering had never been deemed part of that decision workflow at all. The influence it exerted was routing around an absent credential, a compensatory mechanism to keep the outcome aligned rather than a way to evade a control. Formalizing that path is the right response, with a second benefit: the agent learns from the behaviour as well as the paperwork.
The other case is rarer and more serious. Where the informal path routes around segregation of duties, an independent regulatory sign-off or a required second pair of eyes, the slowness was the mechanism. Encoding that path into an agent would convert a control failure into an executable rule. Where independence was the purpose of the control, divergence is evidence of risk, never evidence that practice should prevail.
Then what? Four options, one I would pick today
An agent facing material authority divergence has four options. It can defer to the informal decision-maker, which encodes the workaround. It can escalate to a human, which works but does not scale. It can refuse to act above an agreed risk threshold. Or the organization can formally delegate or revoke the authority before the agent proceeds.
I would choose human escalation today, with refusal above an agreed threshold where escalation is unavailable. I would not encode the informal path: the divergence record alone cannot distinguish a useful workaround from a control failure.
Refusal has its own cost. If it happens frequently, people learn to route around the system, which is the behaviour the framework exists to surface in the first place, and they begin treating the threshold as the problem. Calibration therefore requires a governance judgment: compare the cost of refusing a legitimate decision with the cost of proceeding on an illegitimate one. The answer moves with industry, domain and jurisdiction. Without that baseline, any threshold is guesswork.
One decision is enough to start
The maturity model runs from authority held only as tacit knowledge, through authority written down but never checked against practice, to divergence measured with a direction. I expect most enterprises deploying agents today to sit in the first two levels for their high-stakes decisions.
Start with one consequential decision in one business unit, and apply the three instruments to that alone. Ask the people inside the flow separately, and record where their accounts diverge from each other as well as from the documentation. Two practitioners disagreeing about who decides is itself evidence.
No divergence finding should be used until business, technical and governance have each affirmed it independently. Each catches something the others cannot: whether the documented structure is current policy or already superseded, whether the evidence is complete and temporally sound, and whether the intended use is proportionate to the stakes.
The technical details matter. A verbal approval logged after the fact may carry the timestamp of the logging rather than the decision, materially changing the divergence calculation. The disagreement being measured is organizational before it is technical, and a score one team can produce alone is exactly the kind that goes unchallenged and eventually wrong.
The framework cannot legitimize what it finds
Writing down who shapes a decision is only the diagnostic step. The organization must still decide whether that influence should be recognized, formally delegated, constrained or stopped. No coefficient can make that call, and no ontology can turn repeated behaviour into legitimate authority by itself.
The same constraint applies to the people implementing ARF. The teams collecting, modelling and operationalizing authority should not be able to validate their interpretation alone. That is why business, technical and governance affirmation belongs inside the record.
I have argued that agent-design authority should sit under independent oversight rather than inside the function building the agents. It limits the discretionary authority of the very role I have a professional interest in occupying, and I take that to be the correct consequence of applying the framework’s logic consistently rather than an inconvenience to be elided. If ARF examines undocumented authority everywhere except inside the team operating it, it has reproduced the problem it was built to measure.
An agent that cannot resolve authority has two bad options. It can carry more uncertainty into every call and hand the decision back to people, or it can act confidently on an incomplete model of the organization. One erodes the value the agent was bought to create. The other creates risk at machine speed. Making that choice visible before the agent makes it for you is what this framework is for.
The views expressed in this article are his own, based on his experience and the independent research he conducted for ARF. They do not represent the views of any organisation.
The Cats of Context & Chaos
You recorded who said yes. Who decided? | Source: Context and Chaos
About Context & Chaos
Context & Chaos isn’t just a newsletter. It’s shared community space where practitioners, builders, and thinkers come together to share stories, lessons, and ideas about what truly matters in the world of data and AI: context engineering, governance, architecture, discovery, and the human side of doing meaningful work.
Our goal is simple, to create a space that cuts through the noise and celebrates the people behind the amazing things that are happening in the data & AI domain.
Whether you’re solving messy problems, experimenting with AI, or figuring out how to make data more human, Context & Chaos is your place to learn, reflect, and connect.
Got something on your mind? We’d love to hear from you.
Share this article