CaseAdvancedResponsible AI & Advanced Practice / AI product case study teardowns / #9

Tear down an agent product: what does it let the user approve and what does it not?

GUARD the product is Wayfare Ops, an AI agent that handles supply procurement for a municipal permits office

Wayfare Ops reorders office and field supplies for a city permits department, picks among pre-approved vendors, and places orders under a set dollar cap, mostly without asking anyone first. Emeka Nwosu manages procurement for the department and set the agent up himself.

The direct answer
Wayfare Ops lets the department approve almost nothing directly. It reorders, picks vendors, and places orders under cap entirely on its own, and its logs never mark whether a given purchase was the agent's call or a human's. Fix the log first: tag every action with its real initiator, and give both Emeka's team and outside vendors a visible way to flag and reverse a decision the agent made alone.
Do this, in order
  1. Tag every purchase in the log with its real initiator, agent or human.Why: without this, nobody, including the person running the agent, can prove who actually made a given call.
  2. Give outside vendors a visible channel to flag a decision they believe was automated and wrong.Why: right now, a vendor silently dropped by the agent has no way to know that happened, let alone contest it.
  3. Track how often a previously recurring vendor quietly stops receiving any orders at all.Why: that's the one signal that would catch silent vendor attrition before an outside audit does.
  4. Leave true routine reorders of standard, uncontroversial supplies fully automated.Why: that category carries little real risk and doesn't need a human in the loop.
  5. Route vendor-selection decisions above a certain dollar amount through a human check.Why: that's where the agent's silent cost-optimization does the most quiet damage to long-standing vendor relationships.
  6. Audit a sample of agent-attributed purchases weekly to confirm the attribution log itself is accurate.Why: a mislabeled log is worse than no log, since it creates false confidence.

How to answer this, stage by stage

Seven moves. Handle this plainly. A real harm sitting inside a product decision doesn't need dramatic language, it needs a name.

Stage 1
Scope it to one real agent product
Say it like this
"I'll take Wayfare Ops, an AI agent handling supply procurement for a city department, and look at what it approves on its own versus what it never lets anyone weigh in on."
Why this works
Grounds "tear down an agent product" in one real system with real purchases behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, unequal, ability to contest, reduce, detect."
Why this works
Signals a risk-and-power analysis, not just a feature inventory of what the agent can do.
Stage 3
Name both groups affected
Say it like this
"There's Emeka, who set the agent up and technically holds the lever. And there's a small local vendor the department has ordered from for a decade, who has no login, no account, and no idea an algorithm is deciding whether they still get orders at all."
Why this works
Names the operator and the subject, not just the operator, which is the move most answers skip.
Stage 4
Ask who never gets to push back
Say it like this
"The vendor obviously can't contest a decision they don't know is automated. But even Emeka can't always prove which purchases were his own calls, since the log doesn't separate his approvals from the agent's."
Why this works
The strongest move in GUARD, and here it applies to two very different people for two different reasons.
Stage 5
Name the specific design fix
Say it like this
"Tag every purchase with its real initiator in the log, and give vendors a simple way to flag an order that looks automated and wrong, one that actually routes to a person."
Why this works
A concrete product decision, not a policy document or a training session.
Stage 6
Say how you'd detect it in production
Say it like this
"Track how often a previously recurring vendor's order share quietly drops to zero, and audit a sample of the agent's own attribution logs weekly to make sure they're actually accurate."
Why this works
Shows you'd catch this before an outside audit or a news story does.
Stage 7
Close on the one line
Say it like this
"An agent that acts without a visible lever isn't dangerous because it makes mistakes. It's dangerous because nobody, including the person running it, can prove afterward who actually decided what."
Why this works
Restates the harm in one breath, without softening it into process language.

Let's learn

What happens when nobody, not even the person who set the agent up, can prove which purchases were actually theirs?

Wayfare Ops handles procurement for Emeka's department: reordering standard supplies, choosing among a pre-approved vendor list, and placing orders under a set dollar cap.

Knowledge spark: what's an attribution log, and why would a company skip building one? It's a record of exactly who, or what, made each decision inside a system. Building one that separates human actions from an agent's actions takes real engineering work, and it's easy to skip when a product first launches, since everything looks the same either way until something goes wrong.

For routine categories, standard paper, basic janitorial supplies, this works cleanly. The agent reorders, the department never runs out, and nobody thinks twice about it.

Share of monthly purchases initiated by the agent, by category
100% 50% 0 94%Routine reorder 71%Vendor selection 88%Under $500 cap 4%Needs sign-off
Vendor selection, the category with the most real judgment involved, is still 71 percent decided by the agent alone.

At its worst: the agent's vendor selection quietly, gradually, stops choosing a family-run janitorial supplies vendor the department has used for a decade, in favor of a marginally cheaper option, with no one flagging the change and no way for the old vendor to know or ask why.

Groups affected, named plainly Emeka is the operator: he configured the agent and, on paper, holds the lever to stop it. The vendor is the subject: a small business with no account, no login, and no visibility into whether a person or a program decided to stop ordering from them. Emeka also becomes a subject in a different moment, when his director later asks why a purchase looks unusual and the log can't say whether it was his call or the agent's.

What I would leave alone: fully automating routine, low-stakes reorders of standard supplies is genuinely fine. Nobody's relationship or budget line is at real risk over a paper reorder, and requiring sign-off there would just slow down something that was never actually risky.

An agent that acts without a visible lever isn't dangerous because it makes mistakes. It's dangerous because nobody, including the person running it, can prove afterward who actually decided what.

The lesson: the real question for an agent product isn't just "what can it do without permission." It's "when something goes wrong, can anyone, inside or outside the system, actually point at the decision and say who made it."

Now here is the same thing as a story

The short version above is what you'd say defending this teardown to a director. Read this one for how the gap actually surfaced.

Emeka Nwosu has managed procurement for the city's permits department for six years, and he was the one who configured Wayfare Ops when it rolled out, setting the vendor list and the dollar cap himself.

Hand sketched icon list titled What Wayfare Ops approves without asking. Three rows: a document icon labeled reorder standard supplies, a gauge icon labeled pick among approved vendors, a box icon labeled place orders under a set cap.
Three categories, and all three run without anyone weighing in first. Only one of them is genuinely low-stakes.

For most of a year, this worked well enough that Emeka barely checked the logs. The agent reordered supplies, picked vendors from the approved list, and the department never ran short on anything.

Hand sketched comparison diagram titled Two people one purchase order. Left panel a gauge icon labeled Wayfare Ops, caption holds the button. Right panel a person icon labeled Emeka, caption sees it after the fact.
Emeka's name sits on every purchase either way. Only one side of this panel actually decided which vendor got picked.

A finance audit, a routine one, flagged a purchasing pattern: a family-run janitorial supplies vendor, ordered from monthly for a decade, had received almost nothing in the last six months.

Hand sketched flow diagram titled Where the appeal step should sit. Four boxes: agent drafts order, vendor picked, no contest step highlighted, order placed.
Four steps in every purchase, and the third one, where a contest or appeal should sit, is simply missing.
The vendor's share of monthly orders, six months
20% 10% 0 18% 3% Month 1 Month 6
No dramatic single drop, just a steady decline nobody flagged for six straight months, because nothing about it looked like an error.

Emeka pulled the purchase logs to understand what happened and found he couldn't tell, from the log alone, whether he had personally approved any given switch away from that vendor, or whether the agent had simply chosen the marginally cheaper option every single time on its own.

Hand sketched quadrant titled Sorting agent actions by who should approve. Axes reversible if wrong from hard to undo to easy to undo, and dollar amount from small to large. Reorder paper sits top right, easy to undo and small. Switch vendor sits middle, moderate on both. New supplier contract sits top left, hard to undo and large. Rush shipping fee sits lower right.
Vendor switching sits in genuinely uncertain territory: not the safest corner, and not clearly flagged as needing sign-off either.

The vendor themselves had noticed the drop in orders months earlier, assumed the department had simply found a cheaper supplier through a normal decision, and never once suspected an algorithm had made that call with nobody reviewing it.

Someone building Wayfare Ops decided, at launch, that every purchase would log under one procurement account, whether a human clicked approve or the agent acted alone, since building a separate attribution trail felt like extra engineering effort for a feature nobody had explicitly asked for. That made sense when the agent only handled the smallest, safest reorders.

It stopped making sense the moment vendor selection, a decision with real relationship and budget consequences, moved onto the same unmarked log. I would take that decision back and tag every single action with its real initiator, even if it meant a bit more engineering work up front.

We merged the logs because separating them felt like solving a problem nobody had asked us to solve yet. It took a finance audit catching a six-month vendor decline, with no way to say who or what actually caused it, to understand that the missing lever wasn't a minor gap. It was the entire accountability structure of the product.

GUARD, run on one procurement agentNot a feature checklist. GUARD exists to find who has the lever, who doesn't, and what happens when nobody can prove it either way.

G
Groups.
Emeka, the operator who configured the agent. The vendor, the subject with no account, no visibility, and no idea a machine made the call.
Names both sides, not just the person who set the system up.
U
Unequal.
The harm lands hardest on the vendor, who loses business silently, and on Emeka, whose name sits on every purchase regardless of who actually decided it.
Shows the harm isn't evenly spread, and isn't limited to the obvious outsider either.
A
Ability to contest.
The vendor can't contest a decision they don't know is automated. Emeka can't prove, after the fact, which purchases were genuinely his own calls.
The hardest, strongest step, and it applies to two very different people here.
R
Reduce.
Tag every purchase with its real initiator, and give vendors a visible, simple way to flag a decision they believe was automated and wrong.
A specific design change, not a policy memo or a training session.
D
Detect.
Track silent vendor attrition monthly, and audit a sample of the attribution log itself weekly to confirm it's actually accurate.
Catches the pattern before an audit or a vendor complaint does.

The recap, one line per letter: groups is Emeka and the vendor, unequal is the harm landing on both in different ways, ability to contest is that neither one could actually push back, reduce is tagging real initiators in the log, and detect is watching for silent vendor attrition.

And if you want to be sure it really works, try it somewhere elseSame five letters, a sports team's travel agent instead of a procurement agent. This time the subject who can't contest is the player, not a vendor.

A minor-league sports team uses an AI agent to book travel and hotels for road games, choosing among approved travel partners and charging the team card, mostly without asking a coach or player first.

Mapped onto GUARD: groups are the team's operations manager, who set up the agent, and the players themselves, who show up to whatever hotel and flight the agent booked with no say in the choice. Unequal: players with dietary needs or an old travel-related injury inherit the consequences of a booking they never saw, while the operations manager's name is attached to every booking whether they personally reviewed it or not. Ability to contest: a player who ends up in a hotel far from the team facility has no channel to flag it before travel, only after. Reduce: let players flag a real, standing preference or need that the agent must route through a human check, not just log and ignore. Detect: track how often a booking gets silently changed after a player complaint, since that's the signal a preference is being routinely overridden.

Hand sketched labeled parts diagram titled A team ops agent with the same gap. Center gauge icon labeled Roster agent, with four callouts: books travel, picks hotel, no player sign-off, charged to team card.
A different sport, a different agent, and the same missing lever: a real person absorbs the consequences of a decision they never got to see.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "it approves almost everything itself, and the log can't even say who decided what, fix the attribution first" and stop.
Cost: no engineering budget to rebuild the logging system this quarter. Say so, and start with the cheapest version: a simple manual flag a human adds whenever they personally override or approve something, even without full automated tagging.
The model gets better, for real: even if Wayfare Ops picks better vendors on average, the missing attribution and appeal path remain the same risk. A better model just makes the gap harder to notice, not smaller.

Where people run it wrong.
They ask what an agent can technically do, and stop there, without asking who can push back when it does something wrong.
They assume the "subject" of a risk analysis is always an outsider, missing that the person who set the system up can also lose their own ability to account for it.
They treat a missing log as a technical detail instead of the actual accountability structure of the product.

How to use it live. When asked to tear down an agent product, ask yourself one question before anything else: when this agent does something wrong, can anyone actually point at the decision and say, plainly, who made it? If the honest answer is no, that's the finding, not a footnote.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "tear down an agent: what can it do without approval, and what happens when it's wrong"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Built for risk, safety, and power-imbalance questions.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Emeka Nwosu, who manages procurement for a city permits department and configured Wayfare Ops himself.
3 · THE GROUPS
Who are the two groups this answer names?
Tap to flip
ANSWER
Emeka, the operator, and a small local vendor, the subject, who has no visibility into whether a person or the agent stopped ordering from them.
4 · WHO CAN'T CONTEST
Who never gets to push back in this story, and why?
Tap to flip
ANSWER
The vendor, who doesn't know a decision was automated. And Emeka himself, since the log can't prove which purchases were genuinely his own calls.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Logging every purchase under one shared account with no attribution, reasonable when the agent only handled the smallest, safest reorders.
6 · THE NUMBER
Fill in the blank: the local vendor's monthly order share fell from 18 percent to ___ percent over six months.
Tap to flip
ANSWER
3 percent. A steady decline, not a single dramatic drop, which is exactly why nobody flagged it sooner.
7 · THE FIX IN ACTION
With the reduce and detect steps in place, what changes for the vendor?
Tap to flip
ANSWER
The declining order share gets flagged automatically within weeks instead of six months, and the vendor has a channel to ask why, routed to an actual person.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and who's the subject there?
Tap to flip
ANSWER
A sports team's travel-booking agent. The subject there is the player, who absorbs a travel decision they never got to see or approve.

Check yourself Score: 0 / 0

Short answer, name the groups
1. Who are the two groups this answer names as affected by Wayfare Ops's design, and how is each one harmed?
Show hint
Look at the "groups affected, named plainly" key point.
Show answer
Model answer: Emeka, whose name sits on purchases he may not have personally decided, and the vendor, who loses business with no visibility into why.
Multiple choice
2. Why does this answer say Emeka himself can become a "subject," not just the operator, of this design gap?
  • A. Because he was fired for the vendor decline.
  • B. Because the log can't distinguish his own approvals from the agent's actions, so he can't prove which purchases were genuinely his calls.
  • C. Because he personally chose to stop ordering from the vendor.
  • D. Because he has no access to the purchase logs at all.
Show hint
Look at the "ability to contest" step.
Show answer
B. The missing attribution log hurts the operator too, not only the outside vendor.
True or false
3. True or false: this answer recommends requiring manual sign-off on every single purchase the agent makes, including routine paper reorders.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Routine, low-stakes reorders can stay fully automated. The fix targets attribution and vendor selection, not every action.
Fill in the blank
4. Fill in the blank: vendor selection, the category with the most real judgment, is still ___ percent agent-initiated.
Show hint
Look at the bar chart of purchases by category.
Show answer
71 percent. Higher than most people would expect for a decision with real relationship and budget consequences.
Short answer, name the reduce step
5. What specific design change does this answer propose to reduce the harm, and why is it a real product decision rather than a policy fix?
Show hint
Look at the "reduce" step.
Show answer
Model answer: Tag every purchase with its real initiator and give vendors a visible flag channel, both concrete features, not a training or a memo.
Short answer, apply it yourself
6. Think of an app or service that acts on your behalf automatically, like an auto-renewal or an auto-invest feature. Could you actually tell, from its history, which decisions were yours and which were made for you?
Show hint
Think about a subscription, a smart thermostat, or an auto-investing app.
Show answer
Model answer: Many automated features log an action the same way regardless of who triggered it, exactly the gap this answer names in Wayfare Ops.
Before you close the answer
Why this works
Tests whether you can find the actual power imbalance behind an agent product, including one that hits the person who set the system up, instead of treating "who can push back" as a question only about outside users.
Follow-up traps
"Isn't it unrealistic to expect a vendor to monitor an internal purchasing log?" Response: yes, which is why the fix isn't asking vendors to watch a log, it's giving them a simple channel to flag something that looks wrong, with the monitoring happening on the department's side.

"Wouldn't tagging every action slow the agent down?" Response: no, attribution tagging is a logging change, not a decision-making one. It doesn't add friction to what the agent does, only to how clearly its actions get recorded.
If pressed
The real fix also needs a decay rule: if a previously recurring vendor's order share drops below a set threshold for two consecutive months, the next vendor-selection decision in that category automatically requires a human sign-off, rather than waiting for someone to notice on their own.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more