Nineteen agents, no feedback loops by design, and three versions of trying to teach a system to tell good work from bad. By Bhuvan M and Mohd Jami.
No Feedback Loops, By Design
In February, my colleague Jami and I were standing up agents inside Marketing OS as fast as we could give one a schedule and a Slack handle. Marketing OS was never built for people alone. It started as one shared layer of context that both humans and agents drew from. We then added agents to it, built around all the problems that people wanted to solve. Creating an agent was intentionally quick and easy at the start: point an agent at our marketing context and skills, give it a problem, and let it run. However, we hadn’t realized the impact of building that many agents that fast, and we hadn’t yet built anything to check whether they were actually right.
The gap didn’t take long to show, and it showed as trust, not error messages. None of us could say, with any real confidence, that Marketing OS’s 19 and counting agents were working correctly. We just assumed they were, until someone told us otherwise, and by the time someone did, whatever it got wrong could already have been live for days.
The first message was about a scheduled job that hadn’t been fired. The second was a colleague asking whether something she’d requested had actually happened. By April to May, messages like that weren’t occasional interruptions anymore. They were most of my morning, and every single one of them was small enough that a loop should have caught it before it ever reached a person.
So, a few days before a marketing offsite that July, I tried to see whether the fleet could be scored at all. Nothing automated, just one afternoon with a spreadsheet. One row for each agent, answering a few basic questions: How often did it run? Did it error? What did it cost? I could fill in all of that by hand. Then I got to the column that actually mattered, whether a run was any good. It stayed empty because we didn’t have any definition for what a good agent run looked like. The gap was clear, but how to solve it was a big question.
“What was really stopping the agents from scaling was it did not have feedback loops. By design, we did not have feedback loops.” Surendran Balachandran, VP of AI-led Growth
Taking the Pulse of Our Agents
On the last day of our marketing offsite in Bangalore, we shipped the first stab at a solution, something small but deliberately so. Once a day, for each agent, it posted whether an agent had run, whether it had errored, and what it had cost. We built, demoed, and named it Argus all in one day. However, I knew even while building it that this suffered from the same problem as my spreadsheet, just automated: a monitor that could only tell you a process had finished, never whether it had done anything worth finishing.
A weekly pulse for one agent: whether it ran, whether it errored, and what it cost.
Even though this was a first solution, it already carried two decisions that persisted through every version that came after.
The first is what counts as evidence. For months, the fastest way to find out what had gone wrong had been to ask the agent directly. It always answered in detail, but the answer was often wrong because agents are inherently blind. A missing credential or a blocked connection isn’t visible from inside a conversation, even if they lead to that agent breaking. Argus was built to first check the trace of what actually happened, then what the agent was supposed to produce, then how it’s configured to behave, and only last, underneath all of that, what the agent says about itself.
The order Argus checks in: what actually happened first, what the agent says about itself last.
The second is where Argus lives. If an agent can’t be fully trusted to report on itself, a process running inside that same agent can’t be either. As one person said, the reporter dies with the patient. Argus sits outside our fleet of agents, and it has exactly one way to change anything: it can raise a pull request, which a person has to merge. After all, an observer that can quietly fix what it’s observing stops being an observer.
A fix Argus proposed, as a pull request. A person merged it.
Both decisions paid off within Argus’ opening weeks. A colleague had a job related to our Google Analytics setup that ran daily and kept failing. I looked at it, told him I didn’t see an infrastructure problem, and suggested we wait a day and see if Argus flagged the problem. Argus came back with an answer to a question nobody had asked. The failing job was a minor context issue, sitting on top of a real one. A second, unrelated analytics service, Microsoft Clarity, had been quietly running on a degraded key for more than two weeks. Every number we’d pulled from it in that window was stale, and the agent reading that data every day had never said a word, because nothing in it was built to say anything else. Argus found it and reported it, and we fixed it that day.
Creating One Scale for People and Agents
At first, Argus only watched agents. Shortly after it shipped, the same gap I’d hit with my spreadsheet, how to tell if agent work was good and not just whether it happened, turned up again somewhere I hadn’t thought to look for it — with Atlan’s people.
Jami, who works alongside me in the agent fleet, had been fighting a problem that sounded unrelated. Out of the 300-odd skills built into Marketing OS, a large share had never been used once, and it wasn’t because people were lazy. Half the time they didn’t know a skill existed for the task in front of them. The other half, they built their own version instead of looking for one. Jami wanted to score people’s work the way I’d been trying to score agents: did it land, was it set up well, did the person reach for something that already existed instead of reinventing it?
As it turns out, I’d sketched a rough version of that for agents already: did the run achieve what it was supposed to, was it done cleanly, did the agent use the tools already built for it rather than working around them? We compared notes and realized these were the same rubric wearing different names.
This wasn’t a coincidence. As one person said, the difference between a person, an agent, and a skill is only an internal bookkeeping detail. Eval is eval. It’s the work being judged, not who or what produced it. And so our next version of Argus scored both humans and agents on the same rubric: outcome, craft, and leverage.
“I don’t care about validation from humans. I care about validation from Argus.” Pavithra Mohan, Growth Marketing Manager
The same rubric, sent to a person: one day’s sessions scored on outcome, craft and leverage, with tips.
The Wrong Number
Argus’ first weekly digest after adding this rubric reported that only 20% of the team’s Claude code sessions had used an agent skill. It went out with a siren emoji, signaling a big problem with adoption.
Most people believed it instantly. We were already worried about adoption of our AI agents, so this number intuitively felt right. This kind of stat automatically gets acted on, not double-checked, so within a day the conversation had moved to fixing adoption.
Jami didn’t believe it. He went back into the raw sessions and found the gap wasn’t one mistake. It was three small ones, stacked in the same measurement. The count missed every case where a model read a skill’s instructions directly, the way it reads any other file, instead of calling the one dedicated tool we’d been watching for. It included weeks of sessions from before a separate tracking bug had even been fixed, when skill use couldn’t have been detected no matter what happened in them. And it scored sessions with a missing record as failures instead of leaving them out.
When we fixed these mistakes, we learned the real number was about 56%.
The first corrected readout, one week of sessions. Over the trailing 28 days it settled at 56%.
“We did not think through harness X model permutations while we were designing. We thought all harnesses would do things the same way.” Surendran Balachandran, VP of AI-led Growth
We had learned that harnesses don’t do things the same way. A model decides its own route to a task, and an instrument built on the assumption that there’s one correct route will be wrong in exactly the direction that gets believed. On a team of thirty, mostly marketers, that one wrong number had told nearly everyone they were failing at a job most of them didn’t actually have.
Watching the Watcher
Getting that number right raised an uncomfortable follow-up. If Argus could be wrong about everyone else’s work, could it be wrong about its own?
It was. The job that writes everyone’s scores to our database hit a write that got blocked, and the failure response it got back looked, byte for byte, exactly like success. So the job that scores the whole team’s work ran for two full days, wrote nothing, and reported that everything was fine the entire time.
A blocked write that looked like success. Argus reported all fine for two days while no scores were written.
It was the exact kind of failure Argus exists to catch. But because it happened to Argus itself, leaving behind no trace in Argus or the larger system, no agent or person noticed until someone went looking for numbers that weren’t there.
Someone summed it up by saying that it was our decisions, not AI, that was wrong. We hadn’t built a system that lies. We’d built one that trusted a status code without ever checking what actually happened, which is a much more human mistake to make.
Craft Is Still a Uniquely Human Question
Even after these improvements and learnings, one question was still open, and we couldn’t answer it from the stored data: if Argus judged the same piece of work from an agent twice, would it get the same score? Every re-judgment overwrote the last one, so we had no idea.
We ran the test outside of Argus. It was a quick forty sessions, judged five separate times each — under an hour with a cost of just $12.
The results were unexpected. While raw scores changed across each run, the rankings held almost perfectly across every run. The axis that moved the most, by a wide margin, was craft. We had already guessed a model would struggle to judge craft reliably, but that was just a guess. All it took was a twelve-dollar experiment to answer this definitively. Now we publish the rank of sessions and treat the raw number as a band, not a fact.
Forty sessions, judged five times each, for $12. The rankings held. Craft moved the most.
Looking Beyond Agents and People
Once Argus was good at pulling telemetry out of our own agents, the obvious next question was, why stop there?
Its newest task, which is still a work in progress, has nothing to do with agents or people. Now it’s also watching MCPs, the interface that lets AI models use and interact with websites. Argus has now started auditing our MCP usage too: pulling the day’s telemetry, writing a report, and giving a verdict on what’s working, what needs attention, and which tool calls have quietly started failing more. These are the same habits Argus learned watching agents and humans, now pointed at something new.
Clearing Away Agentic Chaos
Our co-founder Prukalpa has talked about how the numbers everyone reaches for with AI (licenses, weekly actives, agents shipped, and so on) just measure whether people are using AI, not whether the organization is actually getting more capable. Argus was a great demonstration of this. A fleet of agents with no feedback loop can grow busier forever without ever getting better, because nothing inside it can distinguish between a run that finished and a run that actually worked. For a while, neither could we, from the outside.
Three versions in, that’s no longer true. Argus catches a broken script before a person has to and fixes a good share of them itself now, one pull request at a time. It tells the difference between adoption that’s actually weak and a number that only looks that way. It knows when it’s talking to itself.
“What started as just a thought during our Marketing offsite has become something I now trust daily to improve my work with our Marketing-OS repo and overall help build my skills when using AI. From the daily downloads on what I am doing right and where I can improve, to being able to actually chat with it to find key places where I can improve, it has been something that has really accelerated my use of not only Marketing OS but all my Claude sessions.” Steven Hloros, Senior Customer Education Manager
But fixing scripts and catching wrong numbers was never really the goal. It was always about giving people, and the agents working alongside them, a clear enough picture of their own work to get better at it: to build better agents, and to work better with the ones already here.
None of this effort was ever meant to replace judging whether a piece of work is actually good. That was always going to be a person’s call, and the twelve-dollar experiment is proof it still has to be. What changed is how much of everything else Jami, I, and the rest of the team no longer have to check by hand before we get to make that call. That’s the bet Marketing OS was built on in the first place: not a team doing more, but a team with enough of the noise cleared away to spend its attention where it actually matters, which is closer to what doing your life’s best work has always meant.