ConceptAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #20
Explain how feedback design differs for an agent that acts versus an assistant that suggests.
FLIPS the product is FlowPilot, an irrigation control tool from Tanoak Agritech that can suggest a schedule change or, in agent mode, open and close valves on its own
Tanoak Agritech's FlowPilot reads soil-moisture sensors and weather data to manage irrigation. In assistant mode it suggests a schedule change for a person to approve. In agent mode, turned on to keep up during heat waves, it opens and closes valves by itself, inside set limits. Adaeze Mwangi has run irrigation for Elsworth Valley Farms for eight years.
The direct answer
An assistant needs feedback captured at the moment of suggestion, since a person is right there to accept, edit, or reject it. An agent needs feedback captured after the fact, on a much shorter clock, because the action already happened before anyone looked. That means an agent's feedback design has to do two things an assistant never had to: flag anything outside its normal bounds before it acts, not after, and turn its after-the-fact log into something a person actually checks on a schedule, not just whenever they happen to remember.
Do this, in order
Build the after-the-fact action log before agent mode ever ships, not after.Why: once an action already happened, the log is the only feedback channel left; there's no draft to edit anymore.
Hold anything outside the agent's normal bounds for a person, instead of acting and logging it.Why: the actions worth stopping for are exactly the ones nobody would think to go looking for in a log.
Give the log a real schedule, a same-hour alert for big actions, not just a page someone can ignore.Why: a log nobody's told to check is a log nobody checks, no matter how complete it is.
Design for the moment checking stops feeling necessary, not just the moment the agent launches.Why: this flip has no single dramatic trigger, it drifts, so the design has to assume the drift will happen.
Keep a real undo window on every action, sized to how long a mistake takes to show up in the field.Why: feedback after the fact is only useful if there's still time left to act on it.
How to answer this, stage by stage
Nobody is grading whether you know agents act and assistants suggest. They're grading whether you can say what that difference actually does to feedback.
Stage 1
Ground it in one tool, two modes
Say it like this
"I'll answer this for FlowPilot, an irrigation tool that has an assistant mode, suggesting a schedule change, and an agent mode, opening and closing valves on its own."
Why this works
One real tool with both modes lets the comparison stay concrete instead of theoretical.
Stage 2
Say your structure out loud
Say it like this
"I'll use FLIPS. Find the person, locate the habit, identify the flip, pinpoint the old decision, show the replay."
Why this works
Names a method up front, so the comparison doesn't turn into a loose list of pros and cons.
Stage 3
Name the habit that existed under the assistant
Say it like this
"Under assistant mode, Adaeze glanced at FlowPilot's suggested schedule most mornings before approving it. That glance was the feedback loop, built in for free."
Why this works
You can't show what breaks until you've shown what the habit was doing right in the first place.
Stage 4
Name the flip agent mode causes
Say it like this
"Once the valve already moved by the time anyone would look, checking stopped feeling useful in the moment. So she checked less, and less, until she mostly stopped opening the log at all."
Why this works
This is the I step, the hard one: over-trust, and it happened with no single dramatic moment.
Stage 5
Name the design decision that removed the natural pause
Say it like this
"We removed the morning approve screen when agent mode launched, on purpose, to make it feel truly autonomous. We never replaced it with anything that would catch her eye before something went wrong."
Why this works
The P step. Naming a specific, sensible-at-the-time decision, not "add more oversight."
Stage 6
Give the fix, and prove the replay
Say it like this
"Hold anything outside normal bounds for a person before it acts, and page same-hour on anything large. Run the same faulty sensor through that design, and it gets caught within an hour instead of eleven days."
Why this works
The S step. A countable replay is what separates a real fix from a wish that people would just check more.
Stage 7
Close on the actual contrast
Say it like this
"An assistant's feedback lives in the moment of suggestion, because a person is standing right there. An agent's feedback has to live after the fact, on a clock, because by the time anyone looks, it already happened."
Why this works
Restates the direct answer, in the shape of the actual question asked.
Let's learn
For eight years, Adaeze Mwangi could tell an underwatered block from the cab of her truck, just by how the leaves held themselves in the afternoon heat.
Before FlowPilot, adjusting irrigation across Elsworth Valley Farms' twelve blocks took a full morning of manual valve checks. In assistant mode, FlowPilot cut that to twenty minutes: it suggested the day's schedule, and Adaeze approved or edited it over coffee.
Five steps, and the third one is the only genuinely hard one.
In agent mode, FlowPilot opens and closes valves on its own, within limits the farm set, no approval screen involved. That single change is the whole difference this question is really about.
Days the action log went unopened, assistant era versus agent era
The habit of checking didn't vanish overnight. It just had nothing left to check against once the valve had already moved.
The turn: the extra unopened days weren't the real problem. The real problem was that FlowPilot's own reliability was what let Adaeze stop checking, and the one time it was confidently wrong, nothing in the design gave her a reason to look sooner.
The decision I would take back
We removed the morning approve screen the day agent mode launched, on purpose, to make the tool feel genuinely autonomous during heat waves when speed mattered most. That made sense: the approve screen was the whole bottleneck agent mode existed to remove. It stopped making sense once nothing else was built to catch Adaeze's eye before a bad call had already run for days.
At its worst: a soil-moisture sensor on a young orchard block failed, reading "wet" while the ground actually dried out, and FlowPilot confidently reduced irrigation there for eleven straight days during a heat wave, before Adaeze happened to notice leaf wilt on a drive-by. That season's yield on that block never fully recovered.
What I would leave alone: routine, small, reversible valve adjustments, a few minutes either way on a well-established block, don't need a heads-up first. Flagging every small action would just rebuild the bottleneck agent mode was built to remove.
The lesson: giving a tool the power to act doesn't just speed up the good days. It quietly removes the one habit that used to catch the bad ones, and something has to be built on purpose to replace it.
Now here is the same thing as a story
The short version above is what you'd say defending the redesign to Tanoak's product team. Read this one for how the gap actually got found.
Adaeze checks the fields most afternoons, phone in hand, moving block to block faster than any dashboard could keep up with her.
For FlowPilot's first year in assistant mode, her mornings had a rhythm: coffee, open the app, glance at the suggested schedule, approve it in under two minutes, and go check the fields herself anyway, out of habit more than need.
Three branches, and only the third one existed anywhere before the redesign.
When agent mode launched, that morning ritual had nowhere to go. There was no draft to approve anymore, just a log of what had already happened. Adaeze checked it daily at first, out of the same habit. Slowly, with nothing new to catch her attention and nothing ever going wrong, the daily check became every other day, then whenever she thought of it.
Knowledge spark: why does this count as over-trust, not just getting busy?
Over-trust fires when a tool gets better or more capable, not worse. Agent mode was a genuine improvement, faster response during heat waves. The checking habit didn't fade because FlowPilot got worse. It faded because the tool getting more capable removed the one moment that used to remind her to look.
Nobody can point to the exact day Adaeze stopped opening the log. It just thinned out, over weeks, the way habits do when nothing rewards keeping them.
Eleven days, and the middle stretch is the one nobody was watching.
The eleven days weren't eleven days of a broken valve. They were eleven days of a checking habit that had already quietly stopped existing, weeks before the sensor ever failed.
A faulty soil-moisture sensor on Block 9, a young orchard planted the previous spring, began reporting "wet" the same week its wiring corroded. FlowPilot, trusting its own sensor, reduced irrigation there, confidently, correctly following its own bad data, for eleven days straight during the season's hottest stretch.
We assumed her attention would fade like a dial, a little at a time. It didn't. It stayed off until the day she happened to drive past.
Adaeze caught it on day eleven, not from the log, but from the leaves, the same way she'd caught problems for eight years before FlowPilot existed at all.
Block 9's soil moisture: what FlowPilot believed, versus what was real
FlowPilot was never lying. It was reporting a broken sensor with the same confidence it reports a working one, for eleven straight days.
Run the same sensor fault forward under the redesign: a reading that flat for that long, on a block that hot, crosses the "hold and ask" bound by day two. Adaeze gets an alert the same afternoon, checks the block in person within the hour, and finds the corroded wire before a single extra day of underwatering happens.
I removed the approve screen the day agent mode shipped because it was, genuinely, the whole point: speed during a heat wave when a slow human approval could cost a block its whole season. It took eleven quiet days and one young orchard block to see that removing the pause without building anything in its place doesn't just speed up the good days. It removes the only thing that used to catch the bad ones.
The five steps, if you want to remember itNot a checklist for irrigation. FLIPS is what tells you where a real habit thinned out, and what decision to take back for it.
F
Find the person. Whose morning is this?
Adaeze Mwangi, eight years running irrigation by feel before FlowPilot, now trusting a tool she can no longer see approve anything.
A specific person with a real prior skill, not a generic "the farmer."
L
Locate the habit. What did she stop doing?
She stopped opening the daily approve screen, because agent mode removed it, and had no equivalent morning ritual to replace it.
The habit that used to catch problems, named precisely, not as a character flaw.
I
Identify the flip. Checks sometimes, then stops entirely.
Over-trust, fired by an improvement, not a failure. No single dramatic moment, just weeks of the check thinning out.
The hard step, and the reason agent-mode feedback design can't just copy assistant-mode's.
P
Pinpoint the old decision. Removed affordance.
We removed the morning approve screen on purpose to make agent mode feel autonomous, and never built anything to replace the pause it removed.
A specific, sensible-at-the-time decision, not a vague call for "more oversight."
S
Show the replay. Same sensor fault, new design.
The flat reading crosses the hold-and-ask bound by day two. Adaeze checks the block within the hour instead of eleven days later.
A countable ending: hours instead of days, not just "much better."
Four fields on every logged action, and the fourth one is what makes checking late still worth something.
The recap, one line per letter: find the person is Adaeze losing her morning ritual, locate the habit is the daily glance that agent mode removed, identify the flip is over-trust with no single trigger, pinpoint the old decision is removing the approve screen with nothing to replace it, and show the replay is catching the same fault in an hour instead of eleven days.
And if you want to be sure it really works, try it somewhere elseSame five letters, a port gate instead of a farm valve. Nothing else about the two jobs is alike.
Duskferry Port Authority uses a gate-control tool that can suggest which trucks to hold for inspection, or, in agent mode, open and close inspection lanes on its own based on cargo risk scores. Selin Aydemir supervises the gate floor.
Mapped onto FLIPS, with a different flip family this time, an input flip instead of over-trust: find the person is Selin, who used to type detailed manual notes on unusual cargo for the next shift. Locate the habit is those handoff notes. Identify the flip: once agent mode auto-clears most lanes without her involvement, she stops writing detailed notes at all, since most of what she'd note now clears itself, and starts only noting the rare holds, in a much thinner, less useful shorthand. Pinpoint the old decision: the system was built to only prompt her for a note on a hold, never on an auto-clear, so ordinary but informative cargo details, useful context for spotting a pattern later, stopped getting written down anywhere. Show the replay: a redesign that asks one thin, optional note on auto-clears too restores enough context that the next shift catches a smuggling pattern across normal-looking auto-cleared trucks within a week, instead of after a full season of silence.
A different flip family, a different port entirely, and the same shape: a habit that used to write things down, quietly stopped.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "assistants get feedback in the moment, agents need it after the fact, on a clock, plus a before-the-fact bound for anything unusual," and stop.
Cost: building a real hold-and-ask bound takes engineering time the team doesn't have this sprint. Say so, and ship a same-hour alert first; a late catch still beats no catch at all.
The model gets better, for real: if FlowPilot's sensor accuracy improves overall, agent-mode feedback still matters, because a rarer sensor fault is even easier to write off as a one-time fluke instead of the pattern it might be.
Where people run it wrong.
They copy an assistant's feedback design onto an agent and call it done, assuming a log is the same thing as a review step.
They assume attention fades gradually, like a dial, instead of planning for it to just stop, like a switch.
They build the after-the-fact log but never give anyone a reason or a schedule to actually open it.
How to use it live. When someone asks about agents versus assistants, ask yourself first: by the time a person could give feedback, has the thing already happened? If yes, your whole feedback design has to move to before the action and after it, never during.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip. Adaeze checked the action log daily at first, then less and less, until she mostly stopped, as agent mode kept working without incident.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Adaeze Mwangi, who has run irrigation for Elsworth Valley Farms for eight years, since before FlowPilot existed.
3 · THE HABIT
What did Adaeze stop doing because FlowPilot's agent mode worked?
Tap to flip
ANSWER
She stopped opening the action log on a daily rhythm, since there was no draft left to approve and nothing rewarded keeping the habit.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Checking the log most mornings, versus not opening it for days at a time, once nothing ever seemed to go wrong.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Removing the morning approve screen when agent mode launched, without building anything to replace the pause it used to give her.
6 · THE NUMBER
Fill in the blank: the faulty sensor caused reduced irrigation on Block 9 for ___ straight days before it was caught.
Tap to flip
ANSWER
Eleven. Adaeze caught it from the leaves on day eleven, the same way she'd caught problems for eight years before FlowPilot existed.
7 · THE REPLAY
Same sensor fault, new design. What changes?
Tap to flip
ANSWER
The flat reading crosses the hold-and-ask bound by day two, and Adaeze checks the block in person within the hour instead of eleven days later.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and which flip family?
Tap to flip
ANSWER
Duskferry Port Authority's gate-control agent. Input flip this time: Selin stops writing detailed handoff notes once most lanes auto-clear without her.
Check yourself Score: 0 / 0
Short answer, apply it yourself
1. Think of a tool in your own life that switched from suggesting something to just doing it automatically. What did you stop checking once it did?
Show hint
Think about an automatic payment, an auto-reply, or a scheduling tool that stopped asking first.
Show answer
Model answer: Many automated tools remove a review step the moment they're trusted enough to act alone, and the person rarely notices the review step is gone until something goes wrong.
Multiple choice
2. Why can't an agent's feedback design simply reuse the same "approve before it happens" step an assistant uses?
A. Because agents are always less accurate than assistants.
B. Because the action already happened by the time anyone could review it, so feedback has to move to before the fact (bounds) and after it (a checked log), never during.
C. Because agents don't produce any log of what they did.
D. Because assistants only work for irrigation, not for other domains.
Show hint
Look at the direct answer.
Show answer
B. There's no draft left to approve once the agent has already acted, so the feedback moment has to shift to before it (bounds) and after it (a log someone actually checks).
True or false
3. True or false: routine, small, reversible valve adjustments should be flagged for a heads-up before FlowPilot acts on them.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Flagging every small, reversible action would rebuild the exact bottleneck agent mode was built to remove. Only actions outside normal bounds need a heads-up first.
Fill in the blank
4. Fill in the blank: in assistant mode the action log went unopened for an average of 0.5 days. In agent mode, before the redesign, it went unopened for an average of ___ days.
Show hint
Look at the first chart.
Show answer
6. The checking habit didn't disappear on purpose, it just had nothing left to check against once the valve had already moved on its own.
Short answer, name the reversal
5. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Removing the morning approve screen the day agent mode launched. It made sense because that screen was the exact bottleneck agent mode existed to remove; it broke because nothing replaced the pause it used to give Adaeze.
Short answer, where it wouldn't matter
6. Name a kind of valve action in this story that would NOT need a before-the-fact heads-up.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine, small, reversible adjustment on a well-established block. Those don't need a pre-action flag; only actions outside normal bounds do.
Before you close the answer
Why this works
Tests whether you understand that autonomy doesn't just remove work from a person, it removes the moment they used to catch a mistake, and something has to replace it on purpose.
Follow-up traps
"Isn't a same-hour alert on big actions basically the same as requiring approval?" Response: no, the action already happened, the alert is a chance to catch and reverse it fast, not a gate that slows the agent down before it acts.
"What if the agent is right almost every time, does this level of feedback design still matter?" Response: yes, because the rarer the miss, the more over-trust has already set in by the time it happens, which is exactly what made the eleven-day miss possible in the first place.
If pressed
Tanoak's actual redesign sizes the undo window per block, based on how many days a given crop can tolerate reduced water before real damage starts, so a young orchard gets a much tighter bound than an established row crop.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.