ConceptIntermediateModel Fluency & the AI PM Role / The AI literacy baseline every PM needs / #19

Explain what an agent is in a way that distinguishes it from a chatbot with tools.

PICK · Thrumcastle Apparel calls Weftline "your AI stylist" and promises a full outfit, built and checked out for you, but it only searches once and waits

Thrumcastle Apparel sells clothes through an app. Weftline is the feature inside it that answers a shopper's question with real matches from the catalog. Idris Bramante owns it. When the marketing team renamed it "your AI stylist" and promised it would build a whole outfit and check it out for you, nobody changed what Weftline actually does under the hood. Three weeks later, that gap had a dollar figure attached to it.

The direct answer
A chatbot with tools answers one turn, then waits: it calls a tool once and stops until you ask again. An agent keeps going without you: it picks an action, reads the result, and picks the next action from that result, several times in a row, with nobody approving each step. Weftline is the first kind. Thrumcastle's ad promised the second, and that gap is what broke.
In this order
  1. Run the one test: does it take a second action on its own, deciding it from the result of the first, with no one approving each step?Why: that single test is the whole distinction. Everything else is commentary.
  2. Never let a product's name promise more autonomy than its mechanism actually has.Why: Weftline's ad promised multi-step outfit building and checkout it structurally could not do.
  3. If you do build a real agent, put a human check before its most expensive, hardest to reverse step.Why: Duesend's collections escalation needed a pause. Nothing earlier in its chain did.
  4. Fix an undersold chatbot with an afternoon of honest copy before reaching for a bigger rebuild.Why: it is the cheap, visible mistake, and rewriting a promise is almost always cheaper than rebuilding a mechanism.
  5. Watch the complaint rate against a real number, every week, not after a review goes wide.Why: the crossing happened a full week before anyone actually looked.
  6. Weigh which mistake you are actually risking before you pick the word.Why: underselling costs a little shine. Overselling costs trust a refund cannot buy back.

How to answer this, stage by stage

Nobody is grading whether you can define "agent" like a glossary entry. They are grading whether you can hold the line between two words that sound almost the same, under a real, dated number.

1
Anchor the question in one product, one promise
Say it like this
"Let's make this real. Thrumcastle Apparel sells clothes through an app. Weftline is a feature inside it, and the ad campaign for it says 'your AI stylist,' the kind that builds your whole outfit and checks it out for you. I want to test that promise against what Weftline actually does."
Why this works
Pins an abstract word, agent, to one product's real claim, so the answer cannot drift into a dictionary definition.
2
Name the structure before the reasoning
Say it like this
"I'll use PICK. Position: what a chatbot with tools actually is, next to what an agent actually is. Impact: what breaks when a team blurs that line. Cost asymmetry: which mistake is cheap and which one hides. Kill criteria: the one test that tells you which one you actually built."
Why this works
Two seconds of structure tells the interviewer you have a method, not a hunch about a buzzword.
3
State the position, unhedged, one line each
Say it like this
"Here's the split. A chatbot with tools calls a tool once, per message, then it waits for you. An agent chains actions: it decides the next move from the result of the last one, on its own, for more than one step in a row, with nobody approving each step. Weftline searches once and stops. It's a chatbot with tools wearing an agent's name."
Why this works
This is the direct answer, said before any story gets a chance to blur it.
4
Show what actually broke, with real numbers
Say it like this
"Three weeks after 'Your AI Stylist' launched, sixty one percent of sessions from that campaign ended after one screen of search results, almost double the normal rate. One star reviews saying it doesn't do what it says climbed from two a week to forty one. A shopper named Odalys posted the line that made it spread: 'This isn't a stylist. It's a search bar in a chat bubble.'"
Why this works
A real, dated number is the thing this whole argument would fall apart without.
5
Point straight at the cost asymmetry
Say it like this
"If Thrumcastle had undersold Weftline, called it 'AI powered search' and left it there, the fix is a copywriter and an afternoon, about four hundred dollars. They oversold it instead, and by week three that cost them eighty six thousand dollars in refunds plus a forty thousand dollar ad buy they had to pull mid flight. Underselling costs you some shine. Overselling costs you the one thing a refund can't buy back: trust. I'm accepting less flashy marketing on purpose, in exchange for a promise the product can actually keep."
Why this works
Naming which mistake is cheap and which one hides, and naming the trade you're accepting, is the hardest, most convincing move in PICK.
6
Give the kill criteria, as a test anyone can run
Say it like this
"Before I call anything an agent, I ask one question. Does it take more than one action in a row, deciding the next one from the result of the last, with nobody approving each step? If yes, it's an agent. If it does one thing and waits for you, it's a chatbot with tools, no matter what the landing page calls it."
Why this works
A claim with no way to be checked is just a word choice. Naming the exact test is what makes it a real answer.
7
Say what you'd watch, and close on one line
Say it like this
"I'd watch the weekly complaint rate against a real number, not wait for a review to go wide before anyone looks. And I'd leave the plain 'similar items' strip alone, it never promised more than one match, so nobody's confused by it. So: a chatbot with tools waits for you after one step, an agent keeps going without you, and Weftline needed honest words, not a rebuild, to close that gap."
Why this works
Shows judgment past the headline promise, then restates the direct answer in one breath.

Let's learn

Weftline is a feature inside Thrumcastle Apparel's shopping app. You type what you need, a wedding outfit, a rain coat, and it searches the catalog and shows you real matches you can buy.

Hand sketched comparison diagram titled One turn and stop, or keep going without you. Left panel, a plain box icon labeled Chatbot with tools, caption Calls one tool, shows you the result, then waits for your next message. Right panel, a funnel icon labeled Agent, caption Picks an action, reads the result, picks the next one itself, several times in a row.
One box waits for you. The other keeps feeding its own last answer into its next move. That's the whole difference, drawn.

For its first year, Weftline's own copy matched what it did: "AI powered search. Finds pieces that match what you asked for." Modest, accurate, and nobody complained, because nobody expected more than a search. Reviews mentioning confusion sat around two a week, background noise for an app this size.

Knowledge spark: what's a tool call? A model reaches outside itself to get something it doesn't already know, a live search, a price, a stock count, then answers using what came back. One tool call is one question asked and one answer returned. It doesn't decide to ask a second question on its own.
Knowledge spark: what makes something an agent? It keeps going without a person telling it to. It takes an action, checks what happened, and picks its next action from that, on its own, more than once in a row. A loop with nobody's hand on it between steps.

Then a rival app started calling its own feature an "AI stylist," and it tested well. In a launch planning meeting, Thrumcastle's marketing lead, Piers Callendar, pushed to rename the campaign to match: "Your AI Stylist. Tell us the event. We build the whole look. We check it out for you." Idris signed off. Nobody changed what Weftline's system actually did. It still called one search tool per message, showed a handful of matches, and waited.

Hand sketched comparison diagram titled Same feature, two very different promises. Left panel, a document icon labeled Before, caption AI powered search. Finds pieces that match what you asked for. Right panel, a question mark icon labeled After the rebrand, caption Your AI Stylist. Builds the whole look. Checks it out for you.
Same code underneath both panels. Only the promise changed, and the promise is the part a shopper actually reads.

The extra confusion wasn't really about Weftline getting worse at search. It was about what Thrumcastle let a shopper believe would happen next. A shopper who reads "we build the whole look and check it out for you" reasonably expects Weftline to also pick shoes and a bag, put them in a cart, and charge her card, no further typing needed. Weftline had never done any of that, not once, for anyone.

Cost of the two mistakes, three weeks after launch
$130k $65k $0 $400 Underselling honest copy, an afternoon $126,000 Overselling refunds plus pulled ad spend
Underselling, fixed with wordsOverselling, fixed with money
The underselling bar is real, it's just too small to draw at this scale. Four hundred dollars against one hundred twenty six thousand, in three weeks alone, before counting what a shopper who leaves angry is worth over years of repeat orders.
We didn't build a worse chatbot. We built an honest one, then lied about what it would do next.

What it cost at its worst: sessions that came from the campaign ended after one search screen sixty one percent of the time, against a normal rate of thirty three percent for shoppers who never saw the "stylist" promise. Refunds and goodwill credits tied to "not what the AI promised" hit eighty six thousand dollars across three weeks. Piers pulled the remaining ad spend, forty thousand dollars already committed, once the pattern was undeniable.

The choice I would take back In that launch planning meeting, the team decided to change what Weftline's ad promised without changing what Weftline's system did. It made sense at the time: a rival's "AI stylist" language was testing well, and nobody in the room was thinking about it as a claim about how many decisions the system would make without a shopper's help. They were thinking about it as a better sentence.

What I would leave alone: the plain "similar items" strip at checkout, which shows one related item and has never claimed to do anything more. Nobody's ever filed a ticket about it, because its words match its behavior exactly.

The lesson: the word "agent," like "stylist," is a promise about how many decisions a system will make without you. It isn't a mood word for "has AI in it." Say the real mechanism, not the aspiration, and the mechanism will earn the mood word back on its own, later, once it's true.

Now here is the same thing as a story

The short version above is what you actually say out loud. Read this one for the meeting that started it, and the one review that ended it.

Idris Bramante had run Weftline for two years by the time any of this happened. He liked the feature precisely because it didn't overpromise: type what you need, get real matches from a real catalog, no magic. Return rates on items bought through Weftline sat lower than the app average, because the copy set the right expectation and shoppers who found a genuine match kept it.

The trouble started with a competitor. A rival shopping app rebranded its own search feature as an "AI stylist" that spring, and its download numbers jumped. Piers Callendar, who ran marketing for Thrumcastle, brought a slide to the quarterly planning meeting: "We keep calling this thing search. Everyone else calls theirs a stylist. Let's stop underselling what we built." Nobody in the room thought they were promising anything untrue. They thought they were finally describing Weftline the way it deserved.

Hand sketched horizontal timeline titled Three weeks at Thrumcastle Apparel. Four milestones: Your AI Stylist launches, caption week 0, same chatbot underneath. Sessions stall, caption one search then nothing, over and over. Odalys posts her review, this milestone emphasized in amber, caption week 3, a search bar in a chat bubble. Piers pulls the campaign, caption week 3, ad spend already spent.
Nobody decided, on any single day, to promise something Weftline couldn't do. A better sentence in a slide deck just never got checked against the actual code.

"Your AI Stylist" launched on a Monday. The new copy read: "Tell us the event. We build the whole look. We check it out for you." Weftline's code changed not at all. It still called one search tool per message, returned a handful of matching items, and stopped.

For the first week, the numbers looked fine, because nobody was watching the right one. Sessions that came in through the campaign were ending fast, but a fast session had always read as a good session on Thrumcastle's dashboard. One star reviews mentioning confusion crept from two a week to nine. Nobody had a number in mind that would have made nine feel alarming.

Then, on a Tuesday in the third week, Odalys Marchetti opened Weftline looking for something to wear to her sister's outdoor wedding in June. She typed exactly what the ad told her to type. Weftline came back with four dresses that matched. She waited for it to also suggest shoes, a bag, and a way to buy the whole thing at once, the way "we build the whole look and check it out for you" had told her it would. Nothing else happened. It was just waiting for her to type again, same as it always had.

She left a review that afternoon: "This isn't a stylist. It's a search bar in a chat bubble." A shopping forum picked it up by evening. By Wednesday morning, forty one reviews that week used some version of the same complaint, and Piers was the one who found Odalys's post, forwarded to him by a friend outside the company who had no idea he'd had anything to do with the campaign.

We spent six figures finding out that a name is a promise, and a promise gets tested by the first person who takes it literally.

The decision Idris would take back sits in that quarterly planning meeting. Renaming Weftline felt like giving it the credit it deserved. Nobody asked the one question that actually mattered: does the system take more than one action on its own, or does it wait for a person every single time? Weftline waited every time. It had always waited every time. The new name didn't change that. It just changed what a shopper was allowed to expect.

One alternative got seriously discussed before the copy went back to something honest, and it's worth naming, because on paper it looked like the braver fix: build Weftline into an actual agent, one that really does assemble a full outfit and check it out, so the ad would finally be true. It lost, for now. A system that decides on its own to charge a shopper's card needs real guardrails first, a required check before it ever spends money, a hard cap on what it can spend without asking. Thrumcastle hadn't built either one. Shipping a fast version to match a marketing line was a worse bet than telling the truth immediately and building the real thing properly, later, with the safety checks in place before the first autonomous charge, not after.

Run the three weeks again, with the copy left honest from day one: "Smart search, picked for your event." Sessions end after one search screen at the normal thirty three percent, because nobody was promised a second act. One star reviews mentioning confusion hold near two a week. Odalys still gets four good dress options. She still has to pick her own shoes. She never has a reason to write the line that found its way to Piers's inbox.

What Idris would tell himself, back in that planning meeting: a name isn't marketing polish sitting on top of a product. It's a claim about how much work the product will do without you. We changed the claim and left the product exactly where it was, and the gap between the two was never going to stay invisible for long.

PICK, argued from both sides of one word

Not a way to decide whether "agent" is a good word to use. PICK is what forces you to test the word against the actual mechanism, twice, once for a team that oversold a chatbot, and once for a team that undersupervised a real agent.

PPosition. The split, in one line each.
Chatbot with tools: answers one turn. It calls a tool once, shows the result, and waits for the next message. A person approves every single step by sending it.
Agent: keeps going without a person between steps. It picks an action, reads the result, and picks the next action from that result, more than once in a row, on its own.
Weftline is the first kind, dressed in the second kind's name.
Say both positions before any story. A position that only shows up after the evidence sounds reverse engineered from it.
Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a document icon labeled Undersell it, caption Call it search. A copywriter fixes it in an afternoon, about four hundred dollars. Right panel, a question mark icon labeled Oversell it, caption Call it an agent when it isn't. Costs one hundred twenty six thousand dollars and trust no refund buys back.
One mistake is small enough to redraw by hand. The other one needed a chart to hold its own number.
IImpact. What breaks, in both directions.
Call a chatbot an agent, and shoppers expect it to run a whole multi-step goal unattended. They get one search screen, and they feel lied to, exactly what happened to Odalys.
Build a real agent and skip the guardrails for its own real risk, and it takes several unsupervised actions that add up to something nobody would have approved if asked, exactly what happened at Tallymere Bookkeeping's Duesend, in Section 4 below.
Naming both directions of the mistake, not just the flashy one, is what keeps this from turning into a lecture about marketing honesty.
CCost asymmetry. The heart of it.
Underselling a chatbot with tools, calling it "just search," costs you some marketing shine, cheap and visible, fixed with a better sentence. Overselling it as an agent costs you money you can trace and trust you cannot: eighty six thousand dollars in refunds, forty thousand in pulled ad spend, and a review that outran the company's own ability to respond to it. Start with the honest, modest claim. Earn the bigger word once the product can actually do the thing.
Hand sketched decision tree titled The one test that settles it. Root box reads Does it take a second action on its own. Two branches: No, it waits for you to ask again, leading to Chatbot with tools. Yes, it decides step two from step one's result, leading to Agent.
One question, two branches, no third option. Weftline sits on the left branch no matter what its ad said.
KKill criteria. The one test that decides it.
Does the system take more than one action in a row, deciding the next one from the result of the last, with nobody approving each step? If yes, it's an agent, and the job is to guard its riskiest step, not to rename it. If no, it's a chatbot with tools, and the job is to say so plainly, not to borrow a bigger word.
A claim with no way to be checked is an opinion wearing a definition's clothes. Naming the exact test, before anyone asks for one, is what makes this a real answer.
One star reviews saying "doesn't do what it says," by week since launch
45 22 0 kill line: 15/week 2 9 22, crossed here 41 Week 0 Week 1 Week 2 Week 3
Weekly one star reviews, "doesn't do what it says"First point past the kill line
Support had once mentioned fifteen a week as a rough line worth worrying about. Nobody had it on a dashboard. The count crossed it in week two, and it took Odalys's review, a full week later, to make anyone actually look.

Three things worth naming directly, since the real judgment sits here. The alternative worth naming and rejecting isn't only "keep the flashy name," it's also "build the real agent fast to make the ad true," and that one loses for a specific reason: a system that autonomously charges a shopper's card needs a confirm before charge checkpoint and a hard spend cap before it ever gets to act unsupervised, and Thrumcastle hadn't built either. The AI specific failure mode worth naming by name is an under scoped autonomous agent taking a real world, hard to reverse action, spending money, with nobody approving that specific step. The guardrail is exactly those two things: a required human confirmation before any autonomous charge, and a cap on what it can spend in a session without asking. And the trade off is real and accepted on purpose: honest, modest copy gives up some of the urgency that drives clicks, in exchange for never promising a capability the system doesn't have yet.

And if you want to be sure it really works, try it somewhere else

Same four letters, invoices instead of outfits, and this time the honest label was never the problem.

Tallymere Bookkeeping runs small business bookkeeping software. Its tool Duesend chases overdue invoices: it sends a reminder, and if nobody replies within five days, it sends a follow up, applies a late fee, and escalates to a collections referral that cc's the client's own finance team, each step decided from what happened at the last one, with no person approving any single step in between. Casimir Oyelaran, who owns Duesend, built it to be exactly what its name claims. It genuinely is an agent.

Hand sketched flow diagram titled Duesend's chain, five steps, no human turn between them. Five connected boxes reading Invoice overdue, Reminder sent, Day 5 no reply, Late fee applied, Escalate to collections, this last box emphasized in rust to show the riskiest, least reversible step.
Every arrow in this chain fires without anyone's approval. That's what makes Duesend a real agent, and also what made it dangerous once, without a single word changing.

In its first quarter, Duesend ran this full chain, unattended, across 1,140 overdue accounts. Ninety six percent of the time it was exactly right, the client genuinely still owed the money at every step. But in forty six accounts, four percent, a fee or an escalation fired after the client had already paid, because the payment posted to the bank feed a day or two after Duesend had already queued its next step, and nothing in the chain ever paused to check again before acting. One of the forty six was a thirty eight thousand dollar enterprise account that nearly canceled after its own finance team got cc'd on a collections notice for an invoice it had settled four days earlier.

Hand sketched labeled parts diagram titled Same test, two products. A gauge icon in the center labeled One test, with three callouts: Weftline, one call, waits. Duesend, five steps, no pause. Ask, who approves step two.
Run the same test on both products and it splits them cleanly. Weftline fails the test and shouldn't be called an agent. Duesend passes it and should be, but it still needed a guard Weftline never did.

Mapped onto PICK: the position doesn't change, Duesend really does chain more than one action from its own last result, with no person between steps, so calling it an agent was accurate the whole time. The impact is the opposite shape from Weftline's: the mistake here isn't a false promise, it's a true promise with no guard on the one step that could embarrass a paying customer. The cost asymmetry lands just as sharp: adding a pause and recheck before the fee and escalation steps costs a few engineering hours and a slightly slower average collection cycle, cheap and visible. Skipping that pause cost a thirty eight thousand dollar account nearly walking, hidden until the exact week it happened. And the kill criteria still answers it: yes, it takes more than one action in a row on its own, so the fix was never "stop calling it an agent." It was "put a human check before the two steps that are hardest to take back."

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: run the one test, more than one action in a row, deciding the next from the last, no approval in between, that's an agent, anything else is a chatbot with tools, no matter what the landing page says.
Cost: no time to rebuild the guardrail this sprint. Fine, but pause the riskiest step manually until the check exists. Don't let a real agent's most expensive action run unsupervised just because the fix isn't built yet.
The model got better, for real: say Weftline's underlying search model genuinely got sharper at matching. Still doesn't make it an agent. A better single answer to one question is not the same claim as deciding what to do next on its own.

Where people run it wrong.
They let a marketing team pick the word before an engineer checks it against what the system actually does.
They assume "it uses AI" and "it's an agent" are the same claim, when the second one is a specific, checkable claim about who approves each step.
They build a real agent, get the definition right, and stop there, without asking which of its own steps is too expensive to run unsupervised.

How to use it live. Before you answer, ask yourself one thing: does this system take a second action on its own, deciding it from the result of the first, with nobody approving each step? If you can't point to that second, unapproved action, you haven't found an agent yet, you've found a chatbot with a tool and a good name.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question asking you to defend one definition against a nearby, easily confused one?
Tap to flip
ANSWER
PICK: state your position plainly, name the impact of confusing the two, find which mistake is cheap and which one hides, then name the one test that tells them apart.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Idris Bramante, who owns Weftline, the search feature Thrumcastle Apparel's marketing lead, Piers Callendar, renamed "your AI stylist."
3 · THE POSITION
State the position in one line each, for a chatbot with tools and an agent.
Tap to flip
ANSWER
A chatbot with tools calls one tool per message, then waits for you. An agent decides its next action from the result of its last one, more than once in a row, with nobody approving each step.
4 · THE COST ASYMMETRY
Which mistake is cheap to fix, and which one is expensive and hidden?
Tap to flip
ANSWER
Underselling a chatbot as "just search" costs about four hundred dollars, an afternoon of honest copy. Overselling it as an agent cost Thrumcastle a hundred twenty six thousand dollars in three weeks, and trust a refund can't buy back.
5 · THE KILL CRITERIA
What is the one test that tells you which one you actually built?
Tap to flip
ANSWER
Does it take more than one action in a row, deciding the next one from the result of the last, with nobody approving each step? Yes means agent. No means chatbot with tools, no matter what it's called.
6 · THE NUMBER
Fill in the blank: sessions from the campaign ended after one search ___ percent of the time, against a normal rate of ___ percent. One star reviews climbed from two a week to ___ by week three.
Tap to flip
ANSWER
Sixty one percent, against a normal thirty three percent. Forty one reviews a week by week three.
7 · THE OLD DECISION
What decision would Idris take back?
Tap to flip
ANSWER
Changing what Weftline's ad promised without changing what its system did. It felt like finally giving the feature credit it deserved. It was actually a claim about how many decisions the system would make without a shopper's help, and that claim was never true.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's different about its mistake?
Tap to flip
ANSWER
Tallymere Bookkeeping's Duesend, an invoice collection agent. It really is an agent, correctly named. Its mistake was skipping a human check before its riskiest step, not mislabeling itself.

Check yourself Score: 0 / 0

True or false
1. True or false: Weftline is not really an AI product at all, it's a fake.
  • True
  • False
Show hint
Look at the direct answer and the P step of the PICK recap.
Show answer
False. Weftline is a real, useful chatbot with tools, it genuinely searches a real catalog. The flaw was the marketing promise, not the underlying search being fake.
Multiple choice
2. Which of these is the actual test for whether something is an agent rather than a chatbot with tools?
  • A. It uses a large language model somewhere in the pipeline.
  • B. It was marketed using the word "agent" or "stylist."
  • C. It takes more than one action in a row, deciding the next one from the result of the last, with no human turn in between.
  • D. It has a chat interface instead of a search bar.
Show hint
Check the K step in the PICK recap, and stage 6 of the walkthrough.
Show answer
C. The other three describe surface features, the model used, the marketing word, the interface. None of them tell you who approves each step.
Fill in the blank
3. The one star review rate crossed the informal kill line of fifteen a week during week ___. It took until week ___, when Odalys's review spread, for anyone to actually look at the number.
Show hint
Check the line chart under the K step, and the chart note beneath it.
Show answer
Week two, week three. The gap between the crossing and the discovery was a full week, because nobody had the number on a dashboard anyone was watching.
Short answer, where it wouldn't matter
4. Name a place in Thrumcastle's app where calling a feature "AI powered" but not "a stylist" would NOT cause any confusion.
Show hint
Check "what I would leave alone" in Let's learn.
Show answer
Model answer: The plain "similar items" strip at checkout. It only ever shows one related item and has never claimed to do more, so nobody expects more from it, and nobody's confused when that's all it does.
Short answer, apply it yourself
5. Think of an AI tool you use yourself. Does it ever take more than one action in a row without you approving each one? What's the one test you'd run to check?
Show hint
Look for a moment where the tool did something, then did something else, without you sending a new message in between.
Show answer
Model answer: A code assistant that reads a file, decides to run a test based on what it read, and then edits the file again based on the test result, all without you typing anything in between, is acting as an agent for that stretch. A tool that only ever waits for your next message after each step is a chatbot with tools, no matter how capable that one step is.
Short answer, work the number
6. If Weftline's session abandonment rate for campaign traffic had only been forty percent instead of sixty one, barely above the normal thirty three percent baseline, would the overselling mistake still be the same size of problem? Why or why not?
Show hint
Think about what the whole cost argument actually rests on.
Show answer
Probably not the same size, but not fixed either. A smaller gap between promised and actual behavior means fewer angry sessions and a smaller refund bill, but the underlying problem, a chatbot doing the work of one turn while being marketed as something that does the work of several, is exactly as real at forty percent as it is at sixty one. It would just take longer to notice.
Before you close the answer
Why this works
Tests whether a candidate can tell a real capability gap from a marketing word, in both directions, a chatbot dressed as an agent, and a real agent nobody guarded. Most candidates can recite a textbook definition of "agent." Few can apply the same one line test to catch both failures.
Follow-up traps
"Isn't looping through a tool multiple times enough to call something an agent?" Response: no. Looping across separate messages, where you ask again each time, is still a chatbot with tools, because a person approved every step by sending it. The test is whether it decides step two on its own, from step one's result, with nobody approving in between.

"If Duesend really is an agent, was building it the mistake?" Response: no, the label was right the whole time. The mistake was shipping a genuine multi-step agent with no pause before its two riskiest, hardest to reverse steps, the late fee and the collections escalation.
If pressed
The kill test has a concrete, checkable form in an actual execution trace: look for more than one assistant issued action back to back, with only tool result messages between them and zero human messages in between. Find two actions chained like that in one trace, it's an agent. One action per human turn, however many turns happen, is still a chatbot with tools.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more