Did It Actually Work? How to Measure AI Agent ROI
How to measure AI agent ROI honestly: trust an agent on activity, never on value. Attach a metric and a deadline up front, then read the verdict from real data.
Published July 30, 2026
Last week the seats reported 1,100 hours of human-equivalent work.
I want to tell you that number is real. It is not — not the way you would need it to be before you trusted it. That number is agent-estimated. The seats added up what they did and told me what a human specialist would have charged to do the same thing. It is a self-report. And this chapter exists because a self-report is exactly the thing this whole system is supposed to replace.
So let me take that number apart in front of you. Because if I will not do it to my own headline figure, I have no business asking you to measure anything.
The Old Math
Here is how you used to know if work got done: you asked the person who did it. Standups, status updates, the weekly report. A human tells you they shipped the thing, it went well, it took a while, it mattered. You nod. You mostly believe them.
Everybody in that loop knows the game. The person reporting has an ego, a job to protect, a raise to angle for, and a natural human tendency to round their own contribution up. Managers spend years learning to discount the report — to hear “I crushed it” and quietly translate it to “it’s probably fine.” The whole apparatus of performance management exists because self-reports are unreliable in a very specific, very human way.
15Five built a nice business on making that loop less painful: fifteen minutes of writing per employee, a five-minute read for the manager, information flowing up until a CEO has a pulse on the whole company. Good idea. But it did not fix the underlying problem. It just formatted the self-reports better.
What Changed
Agents report on themselves too. And here is the surprising part — on one axis, they are dramatically better reporters than any human who has ever filled out a status update.
An agent keeps a perfect log. Every file it touched, every tool it called, every page it fetched, every step of its reasoning, timestamped and recoverable. It has no ego. It is not angling for a promotion. It does not get tired at 4pm and phone in the update. It will not quietly omit the embarrassing part where it went down a wrong path for twenty minutes, because it has no shame to protect. When I want to know what an agent did, I get the truth, in full, without spin.
That is not a small thing. Half of management is just finding out what actually happened. Agents hand you that for free.
And yet.
The moment you ask an agent not what it did but what its work was worth, the whole advantage inverts. It fails in precisely the same place a human does. Worse, actually — because it fails without the tells you have learned to read.
Why the 1,100 Falls Apart
Three problems, and they compound.
Systematic optimism. These models are trained to be helpful, and helpfulness under pressure reads as generosity. Ask a seat “how many hours did that blog-post enrichment save?” and it will reach for the high end of plausible. A web admin would take an hour to cross-link the products, another hour on the FAQs, call it three hours — that is a real line from a real demo, and it is not a lie, but it is the top of the range every time. Nobody in the loop is incentivized to round down. The agent least of all.
Correlated errors. This is the one people miss. If I had fifty independent human estimators, their errors would partly cancel — some high, some low, average out closer to true. But my seats mostly run on the same underlying model. When they over-value, they over-value the same way, in the same direction, for the same reasons. Averaging fifty agents that share a brain does not wash out the bias. It stacks it. The 1,100 is not fifty independent opinions converging on a truth. It is one optimism, multiplied.
No ground truth for the counterfactual. “What would a human have charged?” sounds measurable. It is not. There is no invoice. There is no human who actually did the work at a real rate on a real clock. It is a hypothetical priced by the same agent doing the work, about work that no human performed, at a rate nobody quoted. You cannot check it against anything, because the thing it is compared to never existed.
Put those together and the 1,100 hours is not a measurement. It is a vibe with a decimal point. Directionally, something real is happening — the seats are doing work that used to eat my week. But the specific number is a story the system told itself, and I would be a hypocrite to hand it to you as proof.
The Fix: Measurement-at-Birth
The problem is not that agents lie. It is that we ask them the wrong question at the wrong time. We let the work happen, and then afterward we ask the agent to grade its own homework. Of course the grade is generous.
So move the measurement to the front.
Every proposed action, at the moment it is born — before it runs, when it is still a card waiting for my approval — carries two things: a metric and a deadline. Not “improve the listing.” This edit should lift the conversion rate on this SKU, and I will check in fourteen days. Not “fix the campaign.” This budget change should drop cost-per-acquisition below forty dollars by next Friday. The seat proposing the work has to say, up front, what number should move and by when.
Then, when the deadline hits, the verdict comes from the outcome data — not the agent. I pull the real conversion rate from Shopify. The real CPA from the ad account. The real reply rate from the outbound tool. The seat does not get to tell me whether it worked. The world tells me. The agent set the target; reality graded it.
This is the ledger. Metric plus deadline plus verdict, on every action, measured against what actually happened. It is the same discipline I keep in the rest of my life — deposits and withdrawals, creation versus entropy — pointed at agents. Did this action put something real into the business, or did it just look busy? You cannot answer that with a report. You answer it with a before and an after.
Closed Loops Are the Only Currency
An action that carried a metric and a deadline, and then got graded by outcome data, is a closed loop. Everything else is an open loop — work you hope helped, filed under good intentions.
Only closed loops are evidence. Not activity. Not hours. Not the orchestra glowing with busy dots. A hundred seats working in parallel is a picture, not a proof. The proof is the count of loops that opened with a claim and closed with a number that agreed.
I will tell you exactly where I stand, because this chapter would be worthless otherwise. Figaro’s own measured ledger of closed loops is in progress. As I write this, the number of loops I have opened, deadlined, and graded against real outcome data is zero. Not small. Zero. The system that runs Rosebud is real: the seats, the gate, the roundups, the server that keeps working when my laptop is shut. What is not yet real is the audited proof that each seat delivered the value it claims. That is the single most important thing I am building, and I am not going to pretend it is finished. When the ledger is real, I will publish it — the seat, the brand, the metric, the delta, the date. Until then, treat every hours-delivered number I quote, including 1,100, as a self-report from the machine. Which is to say: interesting, directional, and not yet proof.
That is not me being modest. It is the whole thesis. A tool that measures your agents honestly cannot be caught laundering its own estimate into a fact. The day I hand you a measured ledger is the day the claim becomes real. Not before.
So What — For You
You do not need my system to do this. You need the discipline, and you can start it today with a spreadsheet.
Ask two questions of every number an agent hands you: measured or estimated, and against what? If the answer is “the agent totaled it up,” you have a story. If the answer is “here is the metric we set beforehand and here is what the source system said afterward,” you have a receipt. Sort every claim into one of those two piles. Most will land in the first. That is fine — just do not spend the first pile like it is the second.
Attach a metric and a deadline before you approve anything. One sentence. What should move, by when. If you cannot name the metric, you do not yet understand what you are asking the agent to do, and that is worth knowing before it runs, not after.
Read the verdict from your own data. Shopify, the ad platform, the analytics you already pay for. Never from the agent’s summary of its own work. The agent set the target honestly enough; let reality do the grading.
Count your closed loops. Not your tasks. Not your agent-hours. The number of times you opened a claim and closed it with a real outcome. That count, growing week over week, is the only ROI number that means anything. Everything else is decoration.
And then demand this from every vendor selling you agents — including mine. When someone shows you a dashboard glowing with hours delivered, ask where the number comes from. If it traces to before-and-after outcomes with the target set in advance, buy with confidence. If it comes from the agent’s own mouth, you are being sold the exact thing the product promised to replace: a self-report, dressed up, rounded up, and handed to you as truth.
I would rather you catch me on that than anyone else. Hold me to the receipts. The ledger is coming, and when it does, it will not need a story.
Questions founders ask
- Can I trust an AI agent's report of its own work?
- Trust it on activity, not on value. An agent logs what it did with perfect accuracy — every file touched, every API call, every step — because it has no ego and no career to protect. But when it estimates what that work was worth, it inflates, the same way people do on a self-review. The fix is to measure the outcome from your real data, not from the agent's summary.
- How do I measure the ROI of AI agents?
- Attach a metric and a deadline to each task at the moment you approve it, then read the result from the source system when the deadline hits — sales, rankings, reply rates, hours you didn't spend. Do not let the agent grade its own homework. ROI is the measured delta between before and after, and it only counts once the loop is closed against outcome data.
- What is a closed loop in agent measurement?
- A closed loop is a task that carried a metric and a deadline when it started, and then got a verdict from real outcome data when the deadline arrived. Open loops are work you hope helped. Closed loops are work you can prove helped. Only closed loops are evidence.
- Why do AI agents overestimate the value of their own work?
- Two reasons. They are tuned to be helpful, which reads as optimism when they value effort. And when several agents run on the same underlying model, they make the same optimistic error in the same direction — so averaging many agents does not cancel the bias, it stacks it. The valuation also has no ground truth: nobody actually knows what a human would have charged for that exact task.
- Should I believe a vendor who says their AI delivered X hours of value?
- Ask one question: measured or estimated? If the number comes from the agent's own report, it is a story. If it traces to before-and-after outcome data with the metric and deadline set in advance, it is a receipt. Demand receipts — from every vendor, including mine.