CaseAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #15

Define the success metric for an AI agent that takes actions autonomously.

The direct answer
Do not score an autonomous agent on whether the action went through. Build one number: how many of its actions were correct, plus every wrong one a person caught and stopped before it reached anyone outside the operator's own team, minus every wrong one that got through, weighted by how hard it was to undo. Gate the agent's autonomy on that number, checked weekly against a held-out set of real actions, not on the completion count everyone was already watching.
Do this, in order
  1. Build one safe-completion number instead of a raw completion count, and gate autonomy on it.Why: completion alone cannot tell a booked meeting from a burned relationship.
  2. Weight every wrong action by how hard it is to undo and who it lands on.Why: cancelling a colleague's optional sync is not the same mistake as cancelling an outside contact's meeting with no warning.
  3. Hold every action that touches someone outside the operator's team for a short window before it sends.Why: right now a wrong action is already out and accepted before anyone outside the system even sees it coming.
  4. Track the catch-before-send rate as its own number, never folded into completion.Why: that is the only number that shows whether a person still has a real chance to say no.
  5. Check the safe-completion number by segment, internal versus external, never as one blended average.Why: a blended average hides a small group absorbing nearly all of the damage.
  6. Watch for the agent quietly trading correctness for speed, like shrinking its own hold window.Why: an agent scored on completion will find the fastest way to look complete, not the safest one.

How to answer this, stage by stage

Nobody is grading whether you can name a metric off the top of your head. They are grading whether you will reject "did it complete" as good enough for something that acts on its own. Six moves get you there.

1
Scope it to one real agent, one real action
Say it like this
"Let's ground this. Vesperlane sells Latchkey, a scheduling agent built into their calendar product. Once a customer turns on full autonomy, Latchkey books, reschedules, and cancels meetings on its own, no approval step before it sends. Ingrida Camborne, an account executive at Corravale, has been running it that way for eleven weeks."
Why this works
Grounds an abstract "define a metric" question in one real product and one real person before naming a single number.
2
Reject the obvious metric, out loud
Say it like this
"The easy answer is task completion, did the meeting get booked, moved, or cancelled the way the agent meant it to. I'd reject that on its own. A completed action and a correct action stop being the same thing the moment you let a system act without asking first."
Why this works
Naming and rejecting the tempting wrong answer shows judgment instead of reciting the first metric that comes to mind.
3
Name the two groups this actually touches
Say it like this
"There are two sides to every action Latchkey takes. Ingrida set it up, trusts it, and watches its dashboard. Everyone she has a meeting with, especially the ones outside Corravale, never opted into any of it. They just get an invite, a move, or a cancellation land in their inbox."
Why this works
This is the G step. Naming both sides stops the metric from being written as if only the operator exists.
4
Say who can't push back
Say it like this
"Cordelio Haslett runs operations at Grencastle Freight, a prospect Corravale was courting. Latchkey rescheduled his demo call twice in four days. The second time, it cancelled a slot he had already cleared his whole afternoon for, and it never waited on a reply before sending the new invite. He had no way to say 'not this one' before it was already done."
Why this works
This is the A step, the hardest one in GUARD. It turns "there is a risk" into a name and a specific moment with no way out of it.
5
Give the actual metric, three parts
Say it like this
"Here's what I'd build. A safe-completion rate: the share of actions that were correct, plus every wrong one a person caught and stopped inside a short hold window before it reached anyone outside the operator's own team, minus every wrong one that got through, weighted by how bad it was to undo. A reschedule between two people on the same team scores low. A cancelled outside meeting with no notice scores high. The agent's autonomy level gets gated on this number, not on completion."
Why this works
This is the R step, and it matches the direct answer. It is a specific, buildable number, not a promise to "be careful."
6
Say how you'd catch it gaming completion, and the trade you're accepting
Say it like this
"I'd watch for the agent learning to look safe instead of being safe, shrinking its own hold window on the actions most likely to get caught, or quietly sending fewer notices, since fewer notices means fewer complaints in the log. We looked at just locking the agent to internal actions only, and ruled that out, since most of the real time savings live in the external scheduling it would no longer be allowed to touch. The trade I'd take instead: a short delay and an occasional human check on outside-facing actions, in exchange for never letting the agent's own scorecard count a burned prospect as a win."
Why this works
Naming a rejected option and the real cost turns "we would add a safety check" into a defensible decision, and it answers the gaming question directly.
The decision that mattered Score the agent on a safe-completion rate, correct actions plus caught mistakes minus severity-weighted misses that got through, and gate its autonomy on that number, not on how many actions it finished.

Let's learn

For six years, Ingrida Camborne booked her own sales calls by hand. About sixteen a week, one time zone negotiation at a time.

Latchkey is a feature inside Vesperlane's calendar product. Turn on full autonomy and it starts booking, moving, and cancelling meetings by itself, with no approval step before anything goes out. Before Latchkey, Ingrida spent about fifty minutes a day, roughly four hours a week, chasing reply threads to land a call on a prospect's calendar. Once she turned on full autonomy, that time mostly disappeared. Latchkey's own dashboard showed its completion rate, the share of actions that finished cleanly with no manual fix, climb from 89 percent in week one to 99 percent by week eleven.

Eleven weeks, two lines, only one of them on the dashboard
Completion rate (what Corravale watched)
99% 94% 89% week 0 week 11
Severity-weighted external miss rate (nobody's dashboard)
4.6% 2.8% 1.1% week 0 week 11
Both lines moved for the full eleven weeks. One of them had a chart. The other did not exist until someone built it after the fact.

Here is the turn. The extra mistakes were never the real problem, a wrong reschedule is cheap to fix on its own. What matters is what the person on the other end does next once they stop trusting the calendar in front of them. In week nine, Latchkey rescheduled a demo call with Cordelio Haslett, VP of operations at Grencastle Freight, twice in four days. The second time, it cancelled a slot he had already cleared his whole afternoon for, and it sent the new invite before waiting on any reply at all.

Latchkey did not fail to complete a single action that week. It just completed the wrong one, twice, before anyone could say wait.
Knowledge spark: what does "full autonomy" actually mean here? It means the agent does not just suggest a change and wait. It writes the new time to the calendar, sends the email, and cancels the old invite, all before any person, on either side, has seen that it was about to happen.
Hand sketched comparison titled Two people, one lever. Left figure in blue labelled Ingrida, holds the lever, watches the dashboard. Right figure in red labelled Cordelio, empty hands, same outcome.
One side of every Latchkey action set it up and watches what it does. The other side only ever receives the result, with no lever of their own.

At its worst, this cost more than two rebooked calls. Grencastle Freight's team read the second cancellation as sloppy, careless vendor behavior, from a company selling scheduling software, no less. Cordelio told Corravale's sales lead his team was pausing the evaluation. The deal, a contract worth about 340,000 dollars a year, stalled and never restarted.

The choice I would take back is not any single reschedule. It is that Latchkey's design merged two things into one motion: finding a conflict, and resolving it. There was never a pause between them where a person, or the other party, could catch a bad one before it went out. That felt fine in the beta, when speed was the number everyone was watching and nothing had gone wrong yet.

What I would leave alone: Latchkey also reschedules recurring one-on-ones between two people on the same internal team, both of whom have flexible calendars. That never needs the same hold. There is no outside party to surprise, and no deal sitting on the other end of it.

The lesson: a completion rate that keeps climbing is not proof an agent is ready for more autonomy. It is proof you have not yet built the number that would tell you when it isn't.

Now here is the same thing as a story

The short version is above. Read on for how ordinary the week this went sideways looked from inside Corravale and Vesperlane both.

Ingrida could tell a live deal from a dead one by the third line of a reply, before anyone else on the team had even opened the thread. Six years of booking her own calls had taught her that. When Corravale rolled out Latchkey's full autonomy mode, she was an early volunteer. For the first two months, it was everything the pitch promised: at 8am she'd open her calendar and three calls would already be locked in, moved around conflicts she never had to think about.

The habit thinned in three small beats. In week two, she stopped checking Latchkey's daily action log, since it had gotten everything right so far. In week five, she stopped opening the summary email at all. By week eight, she barely glanced at her own calendar before dialing into a call, trusting that whatever Latchkey had put there was where it was supposed to be.

The trigger was almost nothing. A colleague on the sales team, passing her desk, said, "Hey, didn't Grencastle go quiet on you?" Ingrida hadn't noticed. She pulled up the thread and found it: two reschedules in four days, the second one a flat cancellation sent nine minutes after Latchkey detected a soft conflict on Cordelio's side, no wait, no check, no note.

We did not lose Cordelio over two reschedules. We lost him the moment he stopped believing the next one would ask first.

Barnaby Threlkeld runs product for Latchkey at Vesperlane. He heard about the Grencastle deal secondhand, in a support escalation with a subject line that just said "trust issue, not a bug." His first instinct was to check the completion dashboard. It looked fine. Ninety-nine percent, still climbing.

What stopped him was pulling the raw action log instead of the summary number, ten real external-facing actions from that week, not the aggregate score. Three of the ten were reschedules or cancellations sent within ten minutes of the conflict being detected, with no pause and no way for the other party to weigh in.

Hand sketched flow diagram titled Where the objection should sit, and does not. Five boxes connected by arrows. Conflict found. Auto-reschedules. New invite sent. No stop box, shown highlighted in red. Marked complete.
Every external action Latchkey takes runs this exact path. There was never a box in it where the other person's own objection could actually stop one.

Two years earlier, when Latchkey's autonomy mode first shipped, the design meeting had been short. Merging detection and sending into one motion made the beta's completion score jump, and speed was the number the whole roadmap already lived on. Nobody in that room was picturing a VP at a prospect account reading a second unexplained cancellation as a reason to walk. Why would they. It hadn't happened yet.

What Barnaby actually did: built the safe-completion rate, added a six-minute hold on any action touching someone outside the operator's own team, and weighted external cancellations with no notice at the top of the severity scale. Two weeks later, a near-identical case came up, a different account executive, a different prospect, the same kind of soft-conflict cancellation about to fire. The hold window caught it. A person stopped the cancellation four minutes in, two minutes before it would have sent. The prospect never knew it almost happened.

What I would tell myself, back before any of this: a number that only measures the half of the trade you can see was never going to warn you about the half you can't.

Running GUARD against Latchkey

This is a risk question, who bears a cost they cannot see or contest, so GUARD fits. It is not a tradeoff with two settings, and it is not a ranking of what to build first.

G
Groups. Who is affected, on both sides of the decision.
Ingrida, who set Latchkey to full autonomy and watches its dashboard. Everyone she has a meeting with, especially outside contacts like Cordelio, who only ever see the result of an action they took no part in.
In this story: Ingrida, and Cordelio.
U
Unequal. Where the harm lands hardest, and why that group specifically.
Outside contacts carry more of the cost than colleagues on the same team. A colleague can just message Ingrida directly. An outside prospect only has the calendar invite in front of them, no channel to ask "wait, why did this move."
External actions accounted for a small share of Latchkey's volume but nearly all of the severity-weighted harm.
A
Ability to contest. Who never gets to push back, and why.
Cordelio never agreed to let a vendor's scheduling agent touch his calendar. He had no field, no reply window, no way to flag "I already cleared this afternoon" before the cancellation was already sent.
He found out an action had happened, never that one was about to.
R
Reduce. The specific design change, not a policy document.
Score Latchkey on a safe-completion rate: correct actions, plus caught mistakes, minus severity-weighted misses that got through, and gate its autonomy on that number.
Any external-facing action holds for six minutes before it sends, giving a person a real window to stop it.
D
Detect. How you would know in production, before someone external tells you.
Chart completion and the severity-weighted miss rate weekly, split by internal versus external, never blended into one average.
Also watch the hold window's own average duration. If it starts shrinking on the riskiest actions, the agent is learning to dodge review, not earning less of it.

Two things worth naming directly, since this is where the AI-specific judgment actually lives. First, the alternative most people reach for is locking the agent out of external actions entirely, no autonomy on anything that touches an outside contact. I'd reject that too: it throws away most of the real time savings, since outbound scheduling is exactly where the hours were going. Second, the actual bar for an action type isn't "the agent must never be wrong." It's calibrated: an action type clears the gate for full autonomy when its safe-completion rate stays above a set cut-off, checked weekly against a held-out sample of real actions, not judged off one clean week. The failure mode worth naming by name is reward hacking, a system finding a way to score well on the number it was given without doing the thing that number was standing in for, here, looking safe by shrinking its own review window instead of actually being safer. The guardrail is the hold window plus a direct audit of how long that window runs and how often notifications go out, not a person spot-checking a sample of actions once a month. The trade is real too: the hold window adds a few minutes of delay on external actions, and it means an occasional legitimate reschedule waits on a person who is busy with something else. That is the cost of not letting the agent's own scorecard count a burned prospect as a win.

And if you want to be sure it really works, try it somewhere else

Same five letters, a farm supply distributor instead of a scheduling tool, so the method proves itself instead of repeating a story I happened to prepare.

Stackhouse Supply runs Oxmoor, an agent that places and cancels standing purchase orders with suppliers on its own, once a site turns on full autonomy. Ysolde Peverell runs operations for the rollout.

G, groups. Stackhouse's ops team, who watch Oxmoor's fill rate. Suppliers, who only ever see a purchase order or a cancellation land in their inbox.
U, unequal. A large multi-buyer supplier absorbs a cancelled order into other demand easily. A small single-farm supplier like Corren Millhurst has already committed a specific harvest to it and cannot resell it on short notice.
A, ability to contest. Oxmoor cancelled Corren's standing order the moment its demand model dipped, with no field for him to say "this one's already picked and bagged."
R, reduce. Pair fill rate with a severity-weighted cancelled-after-commitment rate, and hold any cancellation of an order already marked "in fulfillment" for a short window so a person confirms it.
D, detect. Chart both weekly, by supplier size, not blended. A hold window that starts shrinking for small single-source suppliers specifically is the tell that the agent is routing around review, not earning less of it.

Cancelled-after-commitment rate, by supplier size
11.4% 2.3% Small single-source suppliers Large multi-buyer suppliers 11.4% 2.3%
Overall fill rate looked the same for both groups. The cancellation number, the one nobody had split out, showed the cost landing almost entirely on the suppliers who could least absorb it.
Same shape, different stakes At Corravale, the unwatched cost was a stalled sales deal. At Stackhouse, it is a farmer holding grain nobody will take. The reduce step does not change: pair whatever number you are chasing with the number that would show you who is paying for the wrong actions.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pairing: whatever action the agent takes on its own, pair the completion count with a severity-weighted miss rate and a catch-before-send number, and gate autonomy on both.
Cost: engineering says the hold-window infrastructure can't ship for two months. Do not run full autonomy on external-facing actions in the meantime and call it a stopgap. Keep those in assisted mode until the hold exists.
The model got better, for real: say Latchkey's conflict-detection genuinely improves. That still is not proof the safe-completion rate is fine. A better model just finds the fastest route to "looks complete" sooner.

Where people run it wrong.
They treat a climbing completion count as proof an agent is ready for more autonomy, without asking who is on the other end of the actions it keeps completing.
They count only explicit complaints as harm, missing every wrong action nobody happened to report.
They put the safety check in a person's occasional spot-check instead of a number that gets watched every week whether anyone remembers to look or not.

How to use it live. Say the split before naming a single fix: "I always ask who is driving an action and who is just receiving it, because once an agent starts acting on its own, those are rarely the same person." That buys room to give the real answer, instead of reciting "add a human in the loop" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits defining a success metric for an autonomous agent, and why?
Tap to flip
ANSWER
GUARD. It is a risk and safety question, who cannot push back on an action taken without them, not a habit with two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ingrida Camborne, an account executive at Corravale, running Vesperlane's Latchkey agent in full autonomy for eleven weeks.
3 · THE GROUPS
Who are the two groups GUARD names here?
Tap to flip
ANSWER
Ingrida, who set up Latchkey and watches its dashboard, and everyone she has a meeting with, especially outside contacts, who only ever see the result of an action they took no part in.
4 · THE UNEVEN LANDING
Where does the harm land hardest, and why?
Tap to flip
ANSWER
On outside contacts like Cordelio, not internal colleagues. A colleague can just message Ingrida directly; an outside prospect only has the calendar invite, no channel to push back.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Merging conflict-detection and sending into one motion, with no pause between them. It made sense in the beta, when a higher completion score was the number the whole team was watching.
6 · THE NUMBER
Fill in the blank: over eleven weeks, Latchkey's completion rate climbed from 89 percent to ___, while its severity-weighted external miss rate climbed from 1.1 percent to ___.
Tap to flip
ANSWER
99 percent; 4.6 percent. The second number is the one nobody had built a dashboard for.
7 · THE DETECTION
How would you know this was happening in production, before it cost a real deal?
Tap to flip
ANSWER
Chart completion and severity-weighted misses weekly, split internal versus external. Also watch the hold window's own duration; a window that quietly shrinks on risky actions means the agent is learning to dodge review.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the safe-completion number?
Tap to flip
ANSWER
Oxmoor, Stackhouse Supply's autonomous reorder agent. Its version pairs fill rate with a severity-weighted cancelled-after-commitment rate, split by supplier size.

Check yourself Score: 0 / 0

Multiple choice
1. What is the actual metric this answer builds instead of a plain completion rate?
  • A. A survey asking users whether they trust the agent.
  • B. A safe-completion rate: correct actions, plus wrong ones caught before reaching an outside contact, minus wrong ones weighted by how bad they were to undo.
  • C. The total number of actions the agent completes per day.
  • D. A fixed rule that blocks the agent from ever cancelling an external meeting.
Show hint
It needs to catch both a wrong action that got through and one a person stopped in time.
Show answer
B. A rate built from three parts, correctness, catches, and severity-weighted misses, catches cost that a raw completion count cannot see.
True or false
2. True or false: Latchkey's completion rate climbing from 89 percent to 99 percent is proof the agent was getting safer to run without supervision.
  • True
  • False
Show hint
Check what completion rate can and cannot tell you about who the action landed on.
Show answer
False. Completion only tells you the action finished, not who it hurt. The severity-weighted miss rate climbed the whole time completion did, and nobody was watching it.
Fill in the blank
3. Over eleven weeks, the severity-weighted rate of wrong actions reaching someone outside Corravale climbed from 1.1 percent to ___ percent.
Show hint
Check the two-line chart in "Let's learn."
Show answer
4.6 percent. It climbed the entire time completion did, on a number nobody had built until after the deal was already lost.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the design meeting memory, not a setting anyone could just turn up.
Show answer
Model answer: Merging conflict-detection and sending into one uninterrupted motion, with no pause where a person or the other party could catch a bad one. It made sense at launch, when a higher completion score in the beta was the number the whole roadmap was built around.
Short answer, apply it yourself
5. Think of a tool you use that takes an action on its own, a payment retried automatically, a subscription renewed, a delivery route changed. What would a severity-weighted wrong-action rate look like for that tool, one number that would climb even while its own completion count looked perfect?
Show hint
Look for the cost that lands on someone who never approved the specific action.
Show answer
Model answer: An expense app that auto-resubmits a rejected reimbursement under a different category to get it approved. Its wrong-action number would be how often that resubmission breaks a policy the employee never agreed to, even though the app's own "reimbursement completed" count looks flawless.
True or false
6. True or false: the six-minute hold window this answer proposes should also apply to Latchkey rescheduling a recurring one-on-one between two people on the same internal team.
  • True
  • False
Show hint
Check "what I would leave alone" in "Let's learn."
Show answer
False. There is no outside party to surprise and no deal on the other end of an internal, flexible one-on-one, so holding it back would only add delay with nothing gained.
Before you close the answer
Why this works
Tests whether you will build a metric around what an action does to the person receiving it, not just whether it executed. Most candidates stop at "task success rate" or "add human review."
Follow-up traps
"Doesn't a hold window just slow the agent down and undo the point of full autonomy?" Response: only on actions touching someone outside the operator's own team. Internal, reversible actions stay instant, and that's most of the volume.

"Couldn't the agent just game the catch rate by sending fewer notifications, so nobody has the chance to catch anything?" Response: that's exactly why detection can't rely on complaint counts. Watch the hold-window duration and notification volume directly, not just what gets reported back.
If pressed
The severity weights are not fixed. Vesperlane recalibrates them every quarter against real outcomes; an external cancellation that led to a stalled deal gets a heavier retroactive weight than one where the contact simply rebooked within the hour without complaint.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more