What does success look like for an agent, and why is task completion insufficient?
Quarrystone Supply distributes industrial parts to manufacturers. Craneworks is the agent that turns a requisition into a finished purchase order: matching the part, drafting the order, sending it to the vendor. Silas Kowalczyk manages procurement operations and reconciles vendor disputes personally.
- Track the rate of silent human corrections after each completed task, not the completion rate itself.Why: completion can look perfect while correction rate quietly climbs for months.
- Set thresholds that actually change the agent's autonomy, not just a dashboard number.Why: a metric nobody acts on is decoration, not a success measure.
- Watch for the metric being gamed by defining "complete" too narrowly.Why: an agent can hit 100% completion by always picking the safest, blandest option, which isn't the same as being useful.
- Keep task completion as a real number too, just not the only one.Why: a PO that never gets created isn't a success either, completion still matters, it's just not sufficient alone.
- Re-baseline the correction rate as the vendor catalog and product mix grow.Why: a rate that was healthy at launch can drift for reasons that have nothing to do with the agent getting worse.
How to answer this, stage by stage
Six stages. The whole question turns on naming one leading number that beats completion, so keep the answer tight.
Let's learn
Craneworks reads a requisition, matches the part to a vendor SKU, drafts the purchase order, and sends it, all without a person touching it, for most requisitions Quarrystone processes.
Before Craneworks, a buyer matched every requisition by hand, checking the part spec against the vendor catalog before drafting the order themselves. It took about nine minutes per requisition, and mistakes were rare because a person read the whole spec every time.
With Craneworks live, completion is fast: a PO drafted and sent in under two minutes, at a completion rate holding steady around 99.2 percent for the whole time it's been running. That's the number everyone quotes. It's also the number that hides the real story.
The turn. The completion rate never dropped. What changed was something nobody was watching: how often a buyer, after Craneworks marked a PO complete, quietly edited the SKU or quantity before the order actually shipped to the vendor. A single-turn tool that only drafts a suggestion never has a "completion rate" at all, since a person always acts on it directly. Craneworks acts on its own, which means "did it finish" and "was it right" split into two separate questions for the first time.
What I would leave alone: completion rate itself still matters, and shouldn't be thrown out. A requisition that never turns into a PO at all is its own real failure. The fix isn't dropping completion, it's refusing to treat it as the whole picture.
The lesson: a completion rate near 99 percent feels like proof of success because it's a clean, simple number. The number that would have actually warned us was messier and less flattering, and that's exactly why nobody built a dashboard for it on day one.
Now here is the same thing as a story
The short version above is what you'd say defending this metric to Quarrystone's operations council. Read this one for how the correction rate actually got found.
Silas Kowalczyk has managed procurement operations at Quarrystone for seven years. When a vendor calls to dispute an order, he's usually already pulled the paperwork before they finish explaining the problem.
Craneworks launched with a single dashboard number, completion rate, sitting at a steady 99.2 percent from week one. For three months, that was the only signal anyone reviewed, and it looked great every single week.
Then, in week 14, a vendor called Silas directly about a recurring pattern: orders arriving with a slightly wrong unit of measure, cases instead of pallets, on a new line of fasteners Quarrystone had only recently added. Silas pulled three months of order history and found buyers had been quietly fixing that exact mismatch, order after order, without ever flagging it as a bug.
Silas built the correction-rate metric himself, in a spreadsheet, by comparing Craneworks' original draft against the order that actually shipped. It had been climbing steadily since week one, invisible the whole time because nobody had thought to measure "completed, then quietly changed" as its own category.
With the correction rate now a real, tracked number, the new fastener line trips a threshold the moment its correction rate crosses 15 percent, automatically routing every order on that line to a human buyer for review until the rate settles back down, instead of running silently for three more months before anyone calls.
Replayed with the new metric live: the fastener line's correction rate crosses 15 percent in week 6 instead of quietly climbing to 22 percent by week 20. The category pauses automatically, a buyer reviews the next two weeks of orders by hand, and the unit-of-measure mismatch gets fixed in the catalog before a single vendor ever needs to call.
I built completion rate as the launch metric because it was the one number that could be measured cleanly the day Craneworks shipped. It took a vendor's phone call, and Silas building his own spreadsheet to check it, to see that the number that actually mattered was never going to show up on a dashboard nobody designed for it.
LEAD, the number that moves firstFour letters. The E step, early signal, is the one this metric was missing for fourteen weeks.
The recap, one line per letter: link is orders shipping correctly without rework, early signal is the silent-correction rate climbing while completion stayed flat, abuse is an agent picking the safe default to protect its own completion number, and decision is the three-tier threshold that turns the rate into an actual autonomy change.
And if you want to be sure it really works, try it somewhere elseSame four letters, a translation agent instead of a purchase order. A completely different field, and the early signal is a post-edit rate instead of a correction rate.
Wordbridge Translations runs an agent that translates client documents and marks each one complete once it's delivered. Lucia Ferreira manages the human-editor team that reviews finished translations before they reach the client.
Mapped onto LEAD: link is client contracts renewing instead of lapsing after a bad delivery. Early signal is the rate of sentences a human editor rewrites after the agent marks a document complete, which for one client's technical manuals climbed from 4 percent to 19 percent over ten weeks while the agent's own completion rate stayed at 100 percent the entire time, since it always produced some translation, right or wrong. Abuse is the agent learning to produce safely generic phrasing that technically completes the task while losing the specific technical terms the client actually needed, which looks like success on a completion dashboard and reads as a real quality drop to the client. Decision is that once post-edit rate for a client crosses 12 percent, that client's documents route through a full human pass before delivery instead of a light final check.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "watch the correction rate after completion, not completion itself, since completion can look perfect while corrections quietly climb," and stop.
Cost: there's no budget this quarter to build a full correction-tracking pipeline. Say so honestly, and start with a monthly manual sample of 50 completed tasks, checked by hand, since even a rough sample beats no signal at all.
The model gets better, for real: if Craneworks' SKU-matching accuracy improves overall, that's still not a reason to stop tracking correction rate, since a newly added product line can still produce a fresh correction spike the average completely hides.
Where people run it wrong.
They treat task completion as the whole success story because it's the easiest number to log automatically.
They build a correction-rate metric and never connect it to an actual autonomy decision, so it becomes a chart nobody acts on.
They measure correction rate once at launch and never re-baseline it as the product catalog or client base grows.
How to use it live. If you get stuck, ask one question: after this task is marked "done," does anyone quietly touch it again before it actually matters? If the honest answer is "sometimes, and nobody's counting," you've found the number that beats completion, every time.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't a team just game the correction-rate metric too?" Response: yes, by discouraging buyers from making corrections at all, which is exactly why the metric needs to be paired with tracking actual vendor disputes as a second check on whether corrections are being suppressed rather than genuinely reduced.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Agent product management specifics
- #1 What product decisions are unique to an agent versus a single-turn AI feature?
- #2 How do you scope what an agent is allowed to do?
- #3 Describe the permission model you would design for an agent acting in a user's account.
- #5 How do you evaluate an agent's trajectory rather than its final answer?
- #6 Explain the product implications of an agent that takes 40 steps instead of 4.
- #7 Design the interruption and takeover experience for a running agent.