CaseAdvancedResponsible AI & Advanced Practice / Agent product management specifics / #12

How do you communicate what an agent did after it finishes?

TRACE the product is the returns-and-exchange agent at Wick Outdoor, an outdoor gear retailer

The laptop Renata Kowalczyk works from has a crack across the hinge, held shut on one side with a strip of packing tape. Wick Outdoor's agent handles about 350 returns a day end to end: it checks eligibility, decides refund or store credit or exchange, prints the label, and tells the customer. Renata runs customer operations, and pulls last week's escalations every Thursday.

The direct answer
Don't summarize the outcome. Summarize the decision. Every entry needs what changed, the one-line reason behind it, the policy or rule it applied, how sure it was, and a link back to the real evidence. A log that only says what happened gives a reviewer nothing to defend when a customer, or a colleague, asks why.
Do this, in order
  1. Log the reasoning, not just the result.Why: "processed per policy" tells a reviewer nothing they didn't already know from the outcome itself.
  2. Break the daily summary out by category, not one flat list.Why: a skew hiding inside one product type is invisible in an undifferentiated pile of 350 entries.
  3. Never assume the internal log and the customer-facing message say the same thing.Why: they usually don't, and the gap is exactly where disputes live.
  4. Attach a confidence and a policy reference to every entry, not just the confident majority.Why: the borderline decisions are the ones anyone will actually ask about later.
  5. Make every claim checkable against real evidence, one click away.Why: a summary nobody can verify is just a more confident-sounding rumor.

How to answer this, stage by stage

Nobody's grading whether your summary is short. They're grading whether a person reading it later could actually defend the decision to someone else.

Stage 1
Scope it to one real agent
Say it like this
"I'll answer this for a returns agent that decides refund, credit, or exchange on its own, and needs to explain that decision after the fact."
Why this works
Turns "how do you communicate" into one concrete report you can actually design.
Stage 2
Say your structure out loud
Say it like this
"I'll use TRACE. Build the timeline for one decision, recut the volume by category, assume nothing about what the reader already knows, name the cause candidates for anything surprising, and run one evidence test."
Why this works
Signals you're about to reconstruct a real decision, not just describe a UI.
Stage 3
Build the timeline for one decision
Say it like this
"Take one disputed order. Eligibility checked at 9:02, decision made the same minute, label out at 9:03, customer told at 9:03, inventory adjusted at 9:04. All of that happened in two minutes, and the log only ever showed the last step."
Why this works
Shows how much happens, and how little of it usually gets written down.
Stage 4
Recut by category
Say it like this
"Don't just report 350 returns. Break it out by outcome and product type. Jackets got store credit 22% of the time, against a 6% average everywhere else. That gap was invisible in the flat total."
Why this works
A single daily number hides exactly the pattern worth reporting.
Stage 5
Assume nothing about what the reader knows
Say it like this
"Check whether the internal log and the customer's own message actually match. Ours didn't; both said 'processed per policy,' and neither one said anything more, even internally."
Why this works
Catches the case where nobody, not even the team, actually had the real reasoning.
Stage 6
Name the cause candidates
Say it like this
"Three reasons jackets could skew toward credit: the policy says worn items get credit by design, the classifier over-flags normal shelf wear as worn, or people genuinely wear jackets before returning them more than other items."
Why this works
Turns a strange number into named, testable explanations instead of a shrug.
Stage 7
Run the evidence test
Say it like this
"Pull twenty of those disputed jacket returns and check the reasoning snippet against the actual product photos. That's what told us it was the classifier, not real misuse."
Why this works
The single check that actually separates the real cause from the plausible-sounding guesses.
Stage 8
Close on the one line
Say it like this
"Communicate the reasoning, not the result. The result was never the part anyone needed defending."
Why this works
Leaves the interviewer with the design principle, not a recap of one incident.

Let's learn

Wick Outdoor's returns agent handles all 350 daily returns end to end: it decides refund, store credit, or exchange, prints the shipping label, and messages the customer, without a rep touching it first.

Before the agent, reps handled every return by hand, about four minutes each, and wrote a one-line note in the ticket: "Refunded, per policy." When a customer pushed back later, at least a person remembered the call.

Knowledge spark: what's a reasoning snippet? One short line, in plain words, saying why a decision landed where it did. Not the model's full internal working, just enough for a person to know what to check if they want to push further.

The agent's own log kept the exact same one-line habit: "Processed per policy," on all 350 entries a day, no matter which of the five pipeline steps actually decided the outcome.

Store credit share: all returns vs jacket returns
80% 40% 0 All returns Jacket returns 71% 6% 23% 54% 22% 24%
The store-credit bar nearly quadruples for jackets alone. That gap was sitting inside a daily total nobody had ever split open.

The turn: the extra decisions the agent made were never the real problem. Nobody could tell whether any single one of them was right, because "processed per policy" is what got written whether the reasoning was solid or a coin flip.

We didn't lose the ability to process returns fast. We lost the ability to ever check whether fast was also right.

At its worst: a customer disputed a store-credit decision on a barely-worn jacket, and the only way to answer was a rep manually pulling raw order data and photos, a 25-minute investigation for something the agent had "decided" in under a second.

The decision I would take back We designed the agent's action log to show only the outcome, since a clean one-line entry looked simpler and nobody had pushed back yet. That was fine while decisions were rarely questioned. It stopped being fine the moment customers started asking why, and reps had nothing real to point to.

What I would leave alone: the customer-facing message itself can stay short and plain. The reasoning belongs in the internal log, not in a paragraph a shopper has to read to get their refund status.

The lesson: a summary that only states the outcome isn't a smaller version of a good summary. It's a different thing entirely: a receipt, not an explanation.

Now here is the same thing as a story

The short version above is what you'd say defending this redesign to Wick's operations leadership. Read this one for how the gap actually got found.

Every Thursday, Renata pulls last week's escalations and reads through them at her desk, the cracked laptop propped at an angle that keeps the hinge from finally giving out.

Hand sketched timeline titled Order 48213, reconstructed. Five milestones: eligibility check 9:02am, decision made 9:02am highlighted, label generated 9:03am, customer notified 9:03am, inventory adjusted 9:04am.
Five real steps, all inside two minutes. The old log only ever showed the last one.

For months, Thursday reviews were quick. A handful of escalations, all resolved the same way: reissue, apologize, move on. Nobody dug into why the agent had made the call it made, since the outcome usually looked defensible enough on its face.

Then, on an ordinary Thursday, a senior rep reading over Renata's shoulder said something that stuck: "You're not just trusting that one-line summary, are you? None of us could actually defend these if a customer pushed harder."

Hand sketched comparison diagram titled The log entry, before and after. Left, a box icon labeled Before, caption processed per policy. Right, a document icon labeled After, caption credit item flagged worn 61 percent sure.
Same decision, same outcome. One version gives a reviewer nothing. The other gives them a place to start.

Renata pulled the raw numbers that week for the first time in months, split by product category instead of one running total.

Hand sketched decision tree titled Why more store credit on jackets. Root: jackets get credit 22 percent vs 6 percent average. Three branches: policy worn items get credit leads to expected by design, classifier over-flags shelf wear leads to the real cause, genuine wear-then-return pattern leads to true but a small part.
Three honest guesses, and only one evidence check to tell them apart.

Her team pulled 20 disputed jacket cases and checked the agent's internal confidence and the actual product photos side by side. Most of the "worn" flags were ordinary shelf wear, a slightly creased tag, a faint fold line, nothing close to real damage or use.

Hand sketched labeled parts diagram titled What's inside one action summary. Center icon a document labeled Summary. Four callouts: what changed, why, confidence, what to review.
The redesigned log carries all four of these. The old one carried exactly one: what changed.

Replayed with the redesigned log in place: the same disputed jacket return now shows "Store credit issued: item flagged as worn, 61% confidence, policy WR-4," with a link straight to the photo the classifier flagged. A rep answers the dispute in about three minutes instead of 25, and can actually say why, not just what.

Hand sketched flow diagram titled How the summary reaches a person. Four boxes: agent decides, structured log written highlighted, daily digest built, manager's inbox.
The fix wasn't a smarter agent. It was giving the decision somewhere real to land before it reached a person.

We kept one line per entry because it was simple to build and nobody had complained yet. It took a colleague's offhand question on an ordinary Thursday, not a dramatic failure, to see that "simple" and "defensible" had quietly stopped being the same thing.

TRACE, reading an agent's own trailNot a diagnosis of a metric drop. TRACE here reconstructs what a machine already did, one decision at a time.

T
Timeline. What happened, in order.
Five pipeline steps for one disputed order, all inside two minutes, only the last one ever logged before.
Shows how much detail a one-line summary was always discarding.
R
Recut. By category, not one pile.
Jackets: 22% store credit. Everything else: 6%. Invisible in the daily flat total.
A skew hiding in a segment stays hidden until someone splits the number open.
A
Assume nothing. Check what the reader actually has.
The internal log and the customer message said the exact same four words. Neither had the real reasoning.
The hard step: confirming the gap exists before proposing to close it.
C
Cause candidates. Three named guesses.
Policy design, classifier over-flagging, or a genuine usage pattern, each with a different fix if true.
Turns a strange number into testable explanations instead of a shrug.
E
Evidence test. The one check that decides it.
Twenty disputed jacket cases, reasoning snippet against real photos: mostly shelf wear, not real damage.
The single strongest move: it separates the real cause from the two plausible-sounding wrong ones.
Average time to resolve a customer dispute, week 1 to week 10
25min 12min 0 redesign ships wk1: 25min wk10: 3min
The redesign didn't make the agent's decisions better. It made them checkable, which is what actually cut the resolution time.

The recap, one line per letter: timeline is the five steps inside two real minutes, recut is the jacket-specific skew a flat total hid, assume nothing is catching that even the internal log had no more detail than the customer got, cause candidates is the three named guesses, and evidence test is the twenty-case photo check that found the real one.

And if you want to be sure it really works, try it somewhere elseSame five letters, an HVAC dispatch board instead of a returns queue. No packages in sight, and the same gap shows up anyway.

Dale Petrov supervises dispatch for an HVAC field service company where an agent assigns incoming repair calls to technicians and can reschedule or upgrade a job's priority on its own. Mapped onto TRACE: the timeline is one rescheduled emergency call, reconstructed minute by minute from dispatch to technician notification. The recut shows one technician's queue getting deprioritized more often than the team average, invisible in the daily dispatch count. Assume nothing checks whether the technician's own app and the dispatch office's internal note match, and they don't; the app just says "rescheduled." The cause candidates: that technician's territory has worse traffic data, the agent's routing model underweights his zone, or he's genuinely often unavailable. The evidence test: pulling two weeks of his actual GPS logs against the agent's assumed travel times shows the routing model, not the technician, was the real cause.

Hand sketched icon list titled What a real after-action summary needs. Five items: a document icon labeled category not just outcome, a question mark box icon labeled a one-line reasoning snippet, a scale icon labeled the policy it applied, a gauge icon labeled how confident it was, a funnel icon labeled a link to the real evidence.
The same five things belong in a dispatch summary as in a returns summary. Only the nouns change.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "log the reasoning, not the outcome, and make every claim checkable," and stop.
Cost: there's no budget this quarter for a fancy summary dashboard. Add one plain-text reasoning field to the existing log first; that alone closes most of the gap.
The model gets better, for real: if the agent's decisions get more accurate overall, that's still not a reason to drop the reasoning field. A more accurate agent is still an agent someone will eventually need to defend a specific call from.

Where people run it wrong.
They treat "the agent finished" as the whole story, when finishing and being defensible are different claims entirely.
They write the customer-facing message and assume the internal log says the same thing, without ever checking.
They report one big daily number instead of splitting it by category, hiding exactly the segment worth a second look.

How to use it live. When someone asks how you'd communicate what an agent did, don't describe a dashboard. Ask what a person would actually need to defend one specific decision to someone who disagreed with it, then build the summary around that.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you communicate what an agent did," and what's its one job?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Here its job is reconstructing and explaining an agent's own decision, not diagnosing a metric drop.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Renata Kowalczyk, who runs customer operations at Wick Outdoor and reviews escalations every Thursday from a laptop with a cracked hinge.
3 · THE HABIT
What did Renata's team stop doing because Thursday reviews looked fine?
Tap to flip
ANSWER
They stopped digging into why the agent made each call, since a defensible-looking outcome felt like enough on its own.
4 · THE GAP
What's the actual gap this answer uncovers?
Tap to flip
ANSWER
The internal log and the customer-facing message said the exact same four words. Nobody, not even the team, had the real reasoning on file.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Designing the action log to show only the outcome, since it looked simpler and nobody had pushed back on it yet.
6 · THE NUMBER
Fill in the blank: jackets received store credit ___% of the time, against a ___% average everywhere else.
Tap to flip
ANSWER
22% for jackets, 6% company average. The gap turned out to be a classifier over-flagging normal shelf wear.
7 · THE REPLAY
Same disputed jacket return, redesigned log. What changes?
Tap to flip
ANSWER
The entry now shows the reasoning, confidence, and policy reference. A rep resolves the dispute in about 3 minutes instead of 25.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what did the evidence test uncover there?
Tap to flip
ANSWER
Dale Petrov's HVAC dispatch agent. The evidence test, comparing GPS logs to assumed travel times, showed the routing model, not the technician, was the real cause of his deprioritized queue.

Check yourself Score: 0 / 0

Multiple choice
1. Why did "processed per policy" fail as a log entry, even though it was technically true?
  • A. Because it was too long for the interface.
  • B. Because it gave a reviewer no way to check the reasoning, confidence, or evidence behind the specific decision.
  • C. Because customers found it rude.
  • D. Because the agent wasn't actually following any policy.
Show hint
Look at the direct answer and the comparison diagram.
Show answer
B. A true statement about the outcome still isn't an explanation of the decision behind it.
True or false
2. True or false: this answer recommends showing the full internal reasoning to the customer, not just to internal reviewers.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The customer message stays short and plain. The detailed reasoning belongs in the internal log for reviewers.
Fill in the blank
3. Fill in the blank: after the redesign, average dispute resolution time dropped from 25 minutes to about ___ minutes.
Show hint
Look at the line chart in the TRACE recap section.
Show answer
3 minutes. Reading a structured entry took a fraction of the time of rebuilding the case from raw data.
Short answer, where it wouldn't matter
4. Name a part of this system where a short, outcome-only message is still perfectly fine.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The customer-facing notification. It can stay a short status update; the customer doesn't need the internal reasoning to get their refund.
Short answer, apply it yourself
5. Pick a product you use yourself. Name a time it told you WHAT it did but never WHY, and you had to guess.
Show hint
Think of a spam filter, a credit score change, or a social feed ranking.
Show answer
Model answer: Many people name an email marked as spam with no explanation, leaving them to guess which part of the message triggered it.
Short answer, the number
6. If the jacket skew had been 8% instead of 22%, would this still be worth surfacing in the summary? Why or why not?
Show hint
Think about what actually made the 22% number worth acting on.
Show answer
Model answer: Probably not urgently. Close to the 6% baseline, it likely reflects normal variation rather than a real classifier problem worth an investigation.
Before you close the answer
Why this works
Tests whether you think "communicating what happened" means restating the outcome, or actually building something a person could use to defend a specific decision under real pushback.
Follow-up traps
"Isn't a detailed log just more noise for reviewers to wade through?" Response: no, because the detail is only surfaced when someone actually needs it, through the review flow; the daily digest itself stays a short, categorized summary.

"Doesn't this slow the agent down, generating all that extra detail?" Response: no, the reasoning snippet is generated as part of the same decision the agent already makes; it's writing it down that changed, not the decision itself.
If pressed
The confidence field isn't just cosmetic. Wick's redesign automatically routes anything under 70% confidence into a same-day human review queue, so the number in the summary is also the number driving what gets a second look before a customer ever disputes it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more