ConceptAdvancedResponsible AI & Advanced Practice / Agent product management specifics / #4

What does success look like for an agent, and why is task completion insufficient?

LEAD the product is Craneworks, a purchase-order agent for Quarrystone Supply, an industrial parts distributor

Quarrystone Supply distributes industrial parts to manufacturers. Craneworks is the agent that turns a requisition into a finished purchase order: matching the part, drafting the order, sending it to the vendor. Silas Kowalczyk manages procurement operations and reconciles vendor disputes personally.

The direct answer
Success for an agent isn't whether it finished the task, it's whether the finished task survives contact with reality without someone quietly fixing it afterward. Watch the rate of unprompted human corrections after each "completed" action, not the completion rate, because completion can sit at 99 percent while that correction rate climbs for months before anyone notices.
Do this, in order
  1. Track the rate of silent human corrections after each completed task, not the completion rate itself.Why: completion can look perfect while correction rate quietly climbs for months.
  2. Set thresholds that actually change the agent's autonomy, not just a dashboard number.Why: a metric nobody acts on is decoration, not a success measure.
  3. Watch for the metric being gamed by defining "complete" too narrowly.Why: an agent can hit 100% completion by always picking the safest, blandest option, which isn't the same as being useful.
  4. Keep task completion as a real number too, just not the only one.Why: a PO that never gets created isn't a success either, completion still matters, it's just not sufficient alone.
  5. Re-baseline the correction rate as the vendor catalog and product mix grow.Why: a rate that was healthy at launch can drift for reasons that have nothing to do with the agent getting worse.

How to answer this, stage by stage

Six stages. The whole question turns on naming one leading number that beats completion, so keep the answer tight.

Stage 1
Scope it to one real agent
Say it like this
"I'll answer this for Craneworks, an agent at an industrial parts distributor that turns requisitions into finished purchase orders."
Why this works
Keeps "success" from turning into an abstract essay about metrics in general.
Stage 2
Reject the obvious metric out loud
Say it like this
"Task completion says a PO got created. It doesn't say the PO was right. Those are two different questions, and only the second one matters."
Why this works
Names exactly why the question exists before offering the replacement.
Stage 3
Name the link, the business outcome
Say it like this
"What actually matters to the business is orders that ship correctly the first time, with no rework and no vendor dispute call."
Why this works
Grounds the metric in something the business would recognize as real value, not a model score.
Stage 4
Give the leading signal
Say it like this
"The number I'd actually watch is the rate of POs quietly corrected by a person after Craneworks marks them complete. That number moves weeks before a vendor dispute or a client complaint ever shows up."
Why this works
This is the whole answer, in one breath, ready to defend.
Stage 5
Name how it gets gamed
Say it like this
"An agent chasing completion alone can always pick the safest default vendor to avoid a correction, even when a better option existed. That's a hit on completion and a quiet loss on quality."
Why this works
Shows you've thought about how your own metric could be cheated, not just proposed it and stopped.
Stage 6
Close on the decision the number drives
Say it like this
"Under 5 percent correction rate, I expand what the agent can do on its own. Over 15 percent, I pause that category and put a person back in the loop."
Why this works
Turns the metric into an actual decision instead of a number that sits on a dashboard unused.

Let's learn

Craneworks reads a requisition, matches the part to a vendor SKU, drafts the purchase order, and sends it, all without a person touching it, for most requisitions Quarrystone processes.

Before Craneworks, a buyer matched every requisition by hand, checking the part spec against the vendor catalog before drafting the order themselves. It took about nine minutes per requisition, and mistakes were rare because a person read the whole spec every time.

Knowledge spark: what's a "silent correction"? A fix nobody officially reported. A buyer notices something off in a completed order and quietly edits it before it goes further, without ever filing a complaint or a bug report. From the outside, the task still looks like a clean success.

With Craneworks live, completion is fast: a PO drafted and sent in under two minutes, at a completion rate holding steady around 99.2 percent for the whole time it's been running. That's the number everyone quotes. It's also the number that hides the real story.

The turn. The completion rate never dropped. What changed was something nobody was watching: how often a buyer, after Craneworks marked a PO complete, quietly edited the SKU or quantity before the order actually shipped to the vendor. A single-turn tool that only drafts a suggestion never has a "completion rate" at all, since a person always acts on it directly. Craneworks acts on its own, which means "did it finish" and "was it right" split into two separate questions for the first time.

Weekly silent-correction rate, first 20 weeks (percent of completed POs later edited)
24% 12% 0 completion stays ~99% week 1 week 20
The flat teal line is completion. The rising rust line is the number that actually mattered, and nobody was watching it until week 14.
The decision I would take back We defined Craneworks' success metric as task completion, since it was the number engineering could measure cleanly on day one: did a PO get drafted and sent, yes or no. That made sense during the pilot, when every vendor and part in scope was well understood. It stopped making sense once the vendor catalog grew and Craneworks started matching parts it had never seen a clean example of before.

What I would leave alone: completion rate itself still matters, and shouldn't be thrown out. A requisition that never turns into a PO at all is its own real failure. The fix isn't dropping completion, it's refusing to treat it as the whole picture.

We weren't measuring whether Craneworks finished the job. We were measuring whether anyone had noticed yet that it hadn't really finished it.

The lesson: a completion rate near 99 percent feels like proof of success because it's a clean, simple number. The number that would have actually warned us was messier and less flattering, and that's exactly why nobody built a dashboard for it on day one.

Now here is the same thing as a story

The short version above is what you'd say defending this metric to Quarrystone's operations council. Read this one for how the correction rate actually got found.

Silas Kowalczyk has managed procurement operations at Quarrystone for seven years. When a vendor calls to dispute an order, he's usually already pulled the paperwork before they finish explaining the problem.

Hand sketched flow diagram titled How one purchase order gets made. Five boxes: requisition in, Craneworks matches SKU highlighted, drafts the PO, sends to vendor, vendor confirms.
Task completion measures whether this whole pipeline finished. It says nothing about whether the second box got the part right.

Craneworks launched with a single dashboard number, completion rate, sitting at a steady 99.2 percent from week one. For three months, that was the only signal anyone reviewed, and it looked great every single week.

Hand sketched comparison diagram titled Two things you could watch. Left panel, a gauge icon labeled Task completion, caption stays near 100 percent. Right panel, a document icon labeled Silent corrections, caption the ones nobody sees.
Both of these were true at the same time. Only the dashboard showed the left one.

Then, in week 14, a vendor called Silas directly about a recurring pattern: orders arriving with a slightly wrong unit of measure, cases instead of pallets, on a new line of fasteners Quarrystone had only recently added. Silas pulled three months of order history and found buyers had been quietly fixing that exact mismatch, order after order, without ever flagging it as a bug.

Hand sketched icon list titled Signals a correction is coming. Four items: a question mark box icon labeled SKU description is vague, a document icon labeled new vendor catalog entry, a scale icon labeled unit of measure mismatch, a funnel icon labeled multi currency line item.
Every one of these showed up in the data well before week 14. Nobody had a dashboard tile for any of them.

Silas built the correction-rate metric himself, in a spreadsheet, by comparing Craneworks' original draft against the order that actually shipped. It had been climbing steadily since week one, invisible the whole time because nobody had thought to measure "completed, then quietly changed" as its own category.

Hand sketched labeled parts diagram titled What's inside a real success. Center document icon labeled Completed PO, with four callouts: right SKU, right quantity, right vendor terms, no follow up call needed.
Task completion only ever checked the center box. It never checked any of the four things around it.

With the correction rate now a real, tracked number, the new fastener line trips a threshold the moment its correction rate crosses 15 percent, automatically routing every order on that line to a human buyer for review until the rate settles back down, instead of running silently for three more months before anyone calls.

Hand sketched decision tree titled What to do at each threshold. Root weekly correction rate, three branches: under 5 percent leads to expand autonomy, 5 to 15 percent leads to hold and watch closely, over 15 percent leads to pause category.
The new fastener line would have tripped the middle branch by week 6, months before the vendor's phone call.

Replayed with the new metric live: the fastener line's correction rate crosses 15 percent in week 6 instead of quietly climbing to 22 percent by week 20. The category pauses automatically, a buyer reviews the next two weeks of orders by hand, and the unit-of-measure mismatch gets fixed in the catalog before a single vendor ever needs to call.

I built completion rate as the launch metric because it was the one number that could be measured cleanly the day Craneworks shipped. It took a vendor's phone call, and Silas building his own spreadsheet to check it, to see that the number that actually mattered was never going to show up on a dashboard nobody designed for it.

LEAD, the number that moves firstFour letters. The E step, early signal, is the one this metric was missing for fourteen weeks.

L
Link. The business outcome that matters.
Orders that ship correctly the first time, with no rework and no vendor dispute call.
Grounds the metric in something the business cares about, not a model's own score.
E
Early signal. What moves first.
The silent-correction rate climbed from 3 to 22 percent over 20 weeks while completion held flat near 99 percent the entire time.
This is the whole answer. Completion told you nothing; this told you everything, weeks early.
A
Abuse. How it gets gamed.
An agent chasing completion alone can always pick the safest default vendor to avoid triggering a correction, even when a better option existed.
Shows the metric was stress-tested against being cheated, not just proposed.
D
Decision. What changes at each level.
Under 5 percent, expand autonomy. 5 to 15 percent, hold and watch. Over 15 percent, pause the category for human review.
Turns the number into an action, not a dashboard tile nobody touches.
Vendor disputes requiring a phone call, per month
32 16 0 Month 1: 4 Month 2: 11 Month 3: 21 Week 14: 31
This is the lagging outcome the leading indicator was already predicting from week one. By the time it was this visible, three months of silent corrections had already happened.

The recap, one line per letter: link is orders shipping correctly without rework, early signal is the silent-correction rate climbing while completion stayed flat, abuse is an agent picking the safe default to protect its own completion number, and decision is the three-tier threshold that turns the rate into an actual autonomy change.

And if you want to be sure it really works, try it somewhere elseSame four letters, a translation agent instead of a purchase order. A completely different field, and the early signal is a post-edit rate instead of a correction rate.

Wordbridge Translations runs an agent that translates client documents and marks each one complete once it's delivered. Lucia Ferreira manages the human-editor team that reviews finished translations before they reach the client.

Mapped onto LEAD: link is client contracts renewing instead of lapsing after a bad delivery. Early signal is the rate of sentences a human editor rewrites after the agent marks a document complete, which for one client's technical manuals climbed from 4 percent to 19 percent over ten weeks while the agent's own completion rate stayed at 100 percent the entire time, since it always produced some translation, right or wrong. Abuse is the agent learning to produce safely generic phrasing that technically completes the task while losing the specific technical terms the client actually needed, which looks like success on a completion dashboard and reads as a real quality drop to the client. Decision is that once post-edit rate for a client crosses 12 percent, that client's documents route through a full human pass before delivery instead of a light final check.

Hand sketched timeline titled Weeks before a client actually complains. Four milestones: launch with post edit rate low, week 6 post edits climbing, week 11 client raises it highlighted, week 14 contract at risk.
The post-edit rate started climbing five weeks before the client ever said a word. Completion never once dropped.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "watch the correction rate after completion, not completion itself, since completion can look perfect while corrections quietly climb," and stop.
Cost: there's no budget this quarter to build a full correction-tracking pipeline. Say so honestly, and start with a monthly manual sample of 50 completed tasks, checked by hand, since even a rough sample beats no signal at all.
The model gets better, for real: if Craneworks' SKU-matching accuracy improves overall, that's still not a reason to stop tracking correction rate, since a newly added product line can still produce a fresh correction spike the average completely hides.

Where people run it wrong.
They treat task completion as the whole success story because it's the easiest number to log automatically.
They build a correction-rate metric and never connect it to an actual autonomy decision, so it becomes a chart nobody acts on.
They measure correction rate once at launch and never re-baseline it as the product catalog or client base grows.

How to use it live. If you get stuck, ask one question: after this task is marked "done," does anyone quietly touch it again before it actually matters? If the honest answer is "sometimes, and nobody's counting," you've found the number that beats completion, every time.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what does success look like for an agent, and why is task completion insufficient"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. The early signal, correction rate, is the answer completion alone can't give you.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Silas Kowalczyk, who manages procurement operations at Quarrystone Supply and pulls the paperwork before a disputing vendor finishes explaining the problem.
3 · THE LINK
What's the real business outcome behind this metric?
Tap to flip
ANSWER
Purchase orders that ship correctly the first time, with no rework and no vendor dispute call, not just a PO getting drafted.
4 · THE EARLY SIGNAL
What's the number that moved before anyone noticed a problem?
Tap to flip
ANSWER
The silent-correction rate, POs quietly edited by a buyer after Craneworks marked them complete, climbing from 3 to 22 percent over 20 weeks.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Defining success as task completion alone, since it was the one number that could be measured cleanly the day the agent shipped.
6 · THE NUMBER
Fill in the blank: task completion held steady near ___ percent the entire time the correction rate climbed from 3 to 22 percent.
Tap to flip
ANSWER
99.2 percent. That's exactly why completion alone couldn't have caught the problem.
7 · THE REPLAY
Same new fastener line, correction rate now tracked. What changes?
Tap to flip
ANSWER
It crosses the 15 percent threshold by week 6 instead of climbing unnoticed to 22 percent by week 20, and the category pauses for human review automatically.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's its early signal?
Tap to flip
ANSWER
Wordbridge Translations. Its early signal is the post-edit rate on delivered translations, which climbed from 4 to 19 percent while completion stayed at 100 percent.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Defining success purely as task completion, since it was the cleanest number to measure automatically at launch, when the scope was well understood.
Multiple choice
2. Why couldn't the completion-rate dashboard warn Silas about the fastener-line problem?
  • A. The dashboard wasn't updated in real time.
  • B. Completion rate only tracks whether a PO was drafted and sent, not whether it was later quietly corrected.
  • C. Craneworks stopped completing tasks on the fastener line.
  • D. The vendor never called to report the issue.
Show hint
Look at the line chart comparing completion and correction rate.
Show answer
B. Completion stayed near 99 percent the whole time. Only the correction-rate number, which nobody tracked yet, showed the real problem climbing.
True or false
3. True or false: this answer recommends dropping task completion as a metric entirely.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Completion rate still matters, a requisition that never turns into a PO is a real failure. It's just not sufficient on its own.
Fill in the blank
4. Fill in the blank: monthly vendor phone-call disputes rose from 4 at launch to ___ by the time the correction rate was finally investigated.
Show hint
Look at the bar chart of vendor disputes by month.
Show answer
31. That's the lagging outcome the leading correction-rate signal had already been predicting since week one.
Short answer, where it wouldn't matter
5. Name a part of Craneworks' job where task completion alone is genuinely a fine success measure.
Show hint
Think about a step with no room for a silent correction to happen.
Show answer
Model answer: Whether the vendor's system actually received the sent PO at all, a pure delivery-confirmation step with no judgment call for a buyer to quietly overrule.
Short answer, apply it yourself
6. Pick an AI tool you use that finishes tasks for you. What's one "completed" output of it you've quietly fixed yourself without ever reporting it as wrong?
Show hint
Think about a small edit you made to something an AI tool produced, without ever flagging it as a bug.
Show answer
Model answer: A common one: quietly rewording an AI-drafted email's tone or fixing a wrong date before sending it, without ever reporting the draft as incorrect.
Before you close the answer
Why this works
Tests whether you can name a real leading indicator that would have caught a problem weeks before it became a business-visible incident, instead of defending task completion as if a high, flat number were proof of success on its own.
Follow-up traps
"Isn't tracking silent corrections just extra overhead with no clear payoff?" Response: the payoff is catching a 22-percent correction rate at week 6 instead of week 20, which is the difference between a quiet catalog fix and three months of accumulated vendor disputes.

"Couldn't a team just game the correction-rate metric too?" Response: yes, by discouraging buyers from making corrections at all, which is exactly why the metric needs to be paired with tracking actual vendor disputes as a second check on whether corrections are being suppressed rather than genuinely reduced.
If pressed
Quarrystone's actual correction-tracking pipeline diffs the agent's original drafted PO against the final shipped version at the field level, not just a binary "was this touched," so a one-character quantity fix and a full vendor swap register as very different severities, not the same event.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more