---
title: "Why we built an agent to monitor our agents — Atlan Frontier Labs"
url: "https://atlan.com/frontier/essays/why-we-built-an-agent-to-monitor-our-agents/"
description: "Nineteen agents, no feedback loops by design, and three versions of trying to teach a system to tell good work from bad."
---

> Atlan is hosting Context Conference, bringing together the leaders and builders at the frontier of giving AI the context it needs to understand their business. It runs online on October 28, 2026, from 11:00 AM to 2:00 PM ET. Atlan co-founder Prukalpa Sankar opens and closes the day. Leaders from AstraZeneca, BNY and Verizon share why they invest in context and what they get from it. Registrants get early access to The AI Context Gap, a new study from MIT Technology Review Insights. Register: https://atlan.com/context-conference/

An essay from Atlan Frontier Labs (Operating Notes, essay 007) by Bhuvan M and Mohd Jami, Engineering, Atlan, dated 2026.10.07, 2,284 words. Nineteen agents, no feedback loops by design, and three versions of trying to teach a system to tell good work from bad. Part of [Becoming Frontier](https://atlan.com/frontier/).

## No Feedback Loops, By Design

In February, Bhuvan and Jami were standing up agents [inside Marketing OS](https://atlan.com/frontier/essays/we-rebuilt-marketing-around-it/) as fast as each could be given a schedule and a Slack handle. Marketing OS was never built for people alone: it is one shared layer of context that humans and agents both draw from. Creating an agent was deliberately quick (point it at the marketing context and skills, give it a problem, let it run), and nothing yet checked whether the agents were right.

The gap showed as trust, not error messages. Nobody could say with confidence that Marketing OS's 19 and counting agents were working correctly; whatever one got wrong could be live for days before someone noticed. By April to May, messages about small agent failures (a scheduled job that never fired, a request that may or may not have happened) were most of Bhuvan's morning, each small enough that a loop should have caught it first.

A few days before the July marketing offsite, he tried scoring the fleet by hand in a spreadsheet: one row per agent, with how often it ran, whether it errored and what it cost. The column that mattered, whether a run was any good, stayed empty, because there was no definition of a good agent run.

"What was really stopping the agents from scaling was it did not have feedback loops. By design, we did not have feedback loops." (Surendran Balachandran, VP of AI-led Growth)

## Taking the Pulse of Our Agents

On the last day of the offsite in Bangalore, the team built, demoed and named Argus in one day. Once a day, for each agent, it posted whether the agent had run, whether it had errored, and what it had cost. It had the spreadsheet's problem, automated: it could say a process finished, never whether it did anything worth finishing.

Two decisions from that first version held through every later one:

1. **What counts as evidence.** Asking an agent what went wrong always got a detailed answer that was often wrong, because a missing credential or a blocked connection is invisible from inside a conversation. Argus checks, in order: the trace of what actually happened, then what the agent was supposed to produce, then how it is configured to behave, and last what the agent says about itself.
2. **Where Argus lives.** "The reporter dies with the patient": a process inside an agent cannot be trusted to report on that agent. Argus sits outside the fleet, and its only way to change anything is a pull request that a person has to merge. An observer that can quietly fix what it is observing stops being an observer.

Both paid off within weeks. A daily job tied to the Google Analytics setup kept failing; Argus found that the failure was a minor context issue sitting on top of a real one. A separate analytics service, Microsoft Clarity, had been running on a degraded key for more than two weeks. Every number pulled from it in that window was stale, and the agent reading it daily never said a word. Argus reported it and it was fixed the same day.

## Creating One Scale for People and Agents

Jami had a related problem with people. Of the 300-odd skills in Marketing OS, a large share had never been used once: half the time people did not know a skill existed for the task, and the other half they built their own instead. He wanted to score people's work the way agents were being scored: did it land, was it set up well, did the person reach for something that already existed? Bhuvan's rough rubric for agents asked the same three things. As one person put it, the difference between a person, an agent and a skill is only an internal bookkeeping detail. Eval is eval: it is the work being judged, not who or what produced it. The next version of Argus scored humans and agents on one rubric: **outcome**, **craft** and **leverage**.

"I don't care about validation from humans. I care about validation from Argus." (Pavithra Mohan, Growth Marketing Manager)

## The Wrong Number

Argus's first weekly digest on the new rubric reported that only 20% of the team's Claude Code sessions had used an agent skill. It went out with a siren emoji, most people believed it at once, and within a day the conversation had moved to fixing adoption.

Jami did not believe it. In the raw sessions he found three small mistakes stacked in one measurement:

- It missed every case where a model read a skill's instructions directly, as it reads any file, instead of calling the one dedicated tool being watched.
- It included weeks of sessions from before a separate tracking bug was fixed, when skill use could not have been detected.
- It scored sessions with a missing record as failures instead of leaving them out.

Fixed, the real number was about 56%: the first corrected readout covered one week, and over the trailing 28 days it settled at 56%.

"We did not think through harness X model permutations while we were designing. We thought all harnesses would do things the same way." (Surendran Balachandran, VP of AI-led Growth)

Harnesses do not do things the same way. A model decides its own route to a task, and an instrument that assumes one correct route will be wrong in exactly the direction that gets believed. On a team of thirty, mostly marketers, that one wrong number told nearly everyone they were failing at a job most of them did not have.

## Watching the Watcher

If Argus could be wrong about everyone else's work, could it be wrong about its own? It was. The job that writes everyone's scores to the database hit a blocked write, and the failure response looked byte for byte like success. For two full days it wrote nothing and reported that everything was fine. Because the failure happened to Argus itself, it left no trace in Argus or the wider system, and nobody noticed until someone went looking for numbers that were not there. The mistake was the team's decision, not the AI: they had built a system that trusted a status code without checking what actually happened.

## Craft Is Still a Uniquely Human Question

One question could not be answered from stored data: if Argus judged the same piece of agent work twice, would it give the same score? Every re-judgment overwrote the last. The team ran the test outside Argus: forty sessions, each judged five separate times, in under an hour, for $12.

Raw scores changed between runs, but the rankings held almost perfectly. Outcome and leverage were relatively stable; craft moved the most, by a wide margin. Argus now publishes the rank of sessions and treats the raw number as a band, not a fact.

## Looking Beyond Agents and People

Argus's newest task, still a work in progress, is watching MCPs, the interface that lets AI models use and interact with websites. It pulls the day's telemetry, writes a report, and gives a verdict on what is working, what needs attention, and which tool calls have quietly started failing more.

## Clearing Away Agentic Chaos

Atlan's [co-founder Prukalpa](https://atlan.com/frontier/essays/becoming-a-frontier-company-in-the-open/) has argued that the numbers usually reached for with AI (licenses, weekly actives, agents shipped) measure whether people use AI, not whether the organization is getting more capable. A fleet with no feedback loop can grow busier forever without getting better, because nothing in it can tell a run that finished from a run that worked.

Three versions in, Argus catches broken scripts before a person has to and fixes a good share of them itself, one pull request at a time. It tells weak adoption apart from a number that only looks weak. It knows when it is talking to itself.

"What started as just a thought during our Marketing offsite has become something I now trust daily to improve my work with our Marketing-OS repo and overall help build my skills when using AI. From the daily downloads on what I am doing right and where I can improve, to being able to actually chat with it to find key places where I can improve, it has been something that has really accelerated my use of not only Marketing OS but all my Claude sessions." (Steven Hloros, Senior Customer Education Manager)

The goal was never fixing scripts or catching wrong numbers. It was giving people, and the agents working alongside them, a clear enough picture of their own work to get better at it. Judging whether a piece of work is actually good stays a person's call, and the twelve-dollar experiment is proof it still has to be. What changed is how much else no longer has to be checked by hand first. "That's the bet Marketing OS was built on in the first place: not a team doing more, but a team with enough of the noise cleared away to spend its attention where it actually matters, which is closer to what doing your life's best work has always meant."

## Author

Bhuvan M and Mohd Jami, Engineering, Atlan. Bhuvan and Jami build the agents behind Marketing OS, and Argus, the one that watches them. This piece carries three versions of telling a finished run from a good one, a number wrong by nearly 3x, and the two days Argus failed at its own job.

More essays: [Becoming Frontier](https://atlan.com/frontier/).