Artifact critiqueAdvancedResponsible AI & Advanced Practice / Agent product management specifics / #19

Critique an agent product that gives no visibility into intermediate steps.

AUDIT the product is LaneCast, Duskrail Logistics' freight-booking agent used by Northgale Freight Brokers

Interviewer's question: "Critique an agent product that gives no visibility into intermediate steps." LaneCast, built by Duskrail Logistics, selects a carrier and books a freight lane automatically, showing brokers only a final line: carrier and rate. Bastian Okwuosa manages operations at Northgale Freight Brokers, one of LaneCast's customers.

The direct answer
Don't take Duskrail's "94 percent on-time, cost-optimal" claim at face value. Ask who ran that eval, on what named set of lanes, and on which exact model version, then replicate it yourself on fifty of your own real bookings before trusting it further. And refuse to keep running a booking agent with no per-booking reasoning trail in production at all, since a black box you can't audit after the fact is a liability the day one booking goes wrong.
Do this, in order
  1. Ask who ran the 94 percent eval and whether Duskrail has any reason to inflate it, before trusting the number at all.Why: the vendor selling the agent is the same party who ran the test, and nobody else has checked their work.
  2. Get Duskrail to name the eval set behind that number, which lanes, which carriers.Why: a score with no named test set is a marketing line dressed up as evidence.
  3. Demand a version pin on every booking, so a bad call can be traced to the exact model version that made it.Why: models get updated quietly, and a report with no version pin can never be reproduced or checked.
  4. Replicate the claim yourself on fifty of your own real lanes before trusting it on the rest.Why: a vendor's aggregate score can hide a huge split between easy lanes and the rare ones where judgment actually matters.
  5. Require a visible reasoning trail on every booking going forward.Why: a booking agent with no way to see why is not fit to run unsupervised, whatever its score claims.

How to answer this, stage by stage

Nobody is grading whether you can spot that "no visibility" sounds bad. They're grading whether you can name exactly what evidence is missing and how you'd go get it yourself.

Stage 1
Scope it to one real product
Say it like this
"I'll critique LaneCast, a freight-booking agent that shows Northgale Freight Brokers one line, carrier and rate, with nothing showing how it got there."
Why this works
A critique answered against "agents in general" turns into vague caution. One real product with a real claim keeps it checkable.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself."
Why this works
Signals you're judging evidence, not just reacting to a design choice with a gut feeling.
Stage 3
Question who ran the number
Say it like this
"Duskrail's own marketing says 94 percent on-time, cost-optimal bookings. That's their internal test, not an independent one, and faster adoption means more revenue for them."
Why this works
The A step. Names the incentive behind the number before accepting the number itself.
Stage 4
Demand the eval set and version pin
Say it like this
"Which lanes was that 94 percent measured on, the easy high-volume ones brokers already book confidently, or the messy rare ones? And which model version produced it, since LaneCast has changed at least three times since Northgale signed?"
Why this works
The U and D steps together. A number with no named test set and no version pin can't be trusted or even reproduced.
Stage 5
Name what's missing
Say it like this
"There's no per-booking reasoning trail, no confidence signal, no published failure cases, and no record of how often a broker would have picked differently. The headline number has no denominator anyone can actually see."
Why this works
The I step. What a claim leaves out is usually more informative than what it shows.
Stage 6
Say how you'd test it yourself
Say it like this
"Before trusting the 94 percent, I'd run fifty of last quarter's real bookings back through LaneCast in a shadow test and compare its pick to what an experienced broker actually chose."
Why this works
The T step, the strongest move in the whole framework. Replaces trust in a vendor's claim with your own evidence.
Stage 7
Close on the actual verdict
Say it like this
"If you can't say who ran a number, on what set, on which exact model version, that number is decoration. A booking agent needs a visible reasoning trail before it's fit to run unsupervised, whatever its score claims."
Why this works
Restates the direct answer, in the shape of an actual verdict, not just a list of concerns.

Let's learn

LaneCast reads an incoming freight request, picks a carrier, negotiates a rate, and books the lane automatically. The broker sees one line: carrier and rate, booked.

Before LaneCast, a Northgale broker spent about twenty minutes per lane manually checking a carrier's safety score, on-time history, and rate against two or three alternatives before booking.

Knowledge spark: what's a reasoning trail? A record of why a system chose what it chose. Not just the final answer, the carriers it considered, what it weighed, and how confident it was. Without one, a bad outcome can't be told apart from a fluke, because there's nothing to look at except the result itself.

Now LaneCast books a lane in under a minute, one line, no visible reasoning behind it at all.

Here's the turn: the speed was never the actual problem. The problem shows up the first time a bad pick happens, a carrier with a poor safety record gets booked, and Northgale has no way to know why it was chosen, whether that was a fluke or a pattern, or whether whatever caused it has even been fixed since.

Hand sketched flow diagram titled Where LaneCast's proof runs out. Three steps: rate quoted, carrier booked, no reasoning shown, the last one emphasized in orange.
Two ordinary steps, and then the step that never happens at all.

At its worst: a driver-safety incident happens on a lane LaneCast booked with a carrier that had a flagged record, and Northgale can't produce a single document showing why that carrier was chosen, exposing them to real liability with no paper trail to defend the process at all.

What I'm actually challenging Not necessarily LaneCast's carrier-matching itself, which might be perfectly fine. The challenge is that nobody, Northgale included, can currently tell whether it's fine, and that not-knowing is the real problem, whatever the headline number claims.

What I would leave alone: the speed of the booking itself isn't the issue. A one-minute booking with a visible reasoning trail attached would solve this completely; the critique is about the missing trail, not about automating the task in the first place.

The lesson: a headline accuracy number with no eval set, no version pin, and no reasoning trail isn't proof a system works. It's a claim waiting to be tested, and nobody had tested it yet.

Now here is the same thing as a story

The short version above is what you'd say critiquing this product to Duskrail's own account team. Read this one for how Northgale actually started asking the hard questions.

Bastian Okwuosa had run operations at Northgale Freight Brokers for eleven years, long enough to remember booking every lane by phone, carrier by carrier, checking safety scores by hand.

LaneCast made that entire process disappear. A lane came in, LaneCast booked it, and the confirmation read like a receipt: carrier name, rate, done. For the first year, Bastian mostly stopped thinking about it. The bookings looked fine. Nothing had gone visibly wrong.

Then, at an industry meetup, Wystan Larrabee, who ran operations at a competing brokerage using a similar booking agent, mentioned almost in passing that his own system had booked a lane with a carrier flagged for a safety violation two months earlier, and nobody at his company had known until a client asked about it directly.

Wystan wasn't describing a bug. He was describing exactly what LaneCast did every single day, at every brokerage running it, and neither of them had ever noticed because there was nothing to notice it with.

Bastian went back and asked Duskrail's account team a simple question: what's actually behind the "94 percent on-time, cost-optimal" number on the sales deck. He got a paragraph about "rigorous internal testing" and nothing else. No named lanes. No named carriers. No date. No model version.

Hand sketched labeled parts diagram titled The 94 percent claim, taken apart. Center document icon labeled The 94% Claim, with four callouts: who ran it, what set, which version, what is missing.
Four questions. Duskrail's account team could not answer any of them.

He pushed further and learned LaneCast had been updated at least three times since Northgale signed, informally, through a changelog email nobody at Northgale had actually read closely. There was no record anywhere connecting a specific booking to the specific model version that made it. If a bad booking had happened in March, nobody, not even Duskrail, could say for certain which version of the agent made that call.

Hand sketched decision tree titled Should you trust this vendor's number. Root: vendor claims 94 percent. Three branches: eval set named leads to worth checking further, eval set not named leads to treat as marketing, you replicated it yourself leads to now you actually know.
Northgale had been living entirely in the middle branch for over a year without realizing it.

Rather than keep arguing with Duskrail's account team, Bastian pulled fifty of Northgale's own real bookings from the last quarter and ran them back through LaneCast in a shadow test, comparing its carrier pick to what an experienced human broker had actually chosen for the same lane at the time.

LaneCast's pick versus a broker's pick, by lane type, in Northgale's own shadow test
100% 50% 0% 91% Simple, high-volume lanes 58% Rare, complex lanes
The vendor's 94 percent was true and also misleading, since it averaged across a mix dominated by exactly the lanes where an agent's judgment barely matters.

The decision I'd challenge here is Duskrail's, not Northgale's: choosing to show one line, "Booked," instead of any reasoning, because it made the product feel simpler and more impressive in the sales demo. That looked good in a five-minute pitch. It stopped looking good the first time a bad booking needed an explanation nobody could produce.

Replayed with a required reasoning trail in place: the flagged-carrier booking that worried Wystan would have shown its own confidence level and the carriers it considered instead, and Northgale's monthly audit would have caught the pattern within weeks instead of relying on a chance conversation at an industry meetup.

I used to think a clean, simple confirmation screen was a sign of a mature product; less clutter, more trust. It took one flagged carrier and a conversation with a competitor to see that the missing clutter was actually the evidence, and nobody had ever been able to check the work at all.

AUDIT, spelled out for one claimNot a takedown of LaneCast specifically. AUDIT is what tells you whether any claim like this deserves your trust yet.

A
Ask who paid for it.
Duskrail ran its own 94 percent test. The company selling the agent is the same party who measured how good it is, with every incentive to make the number look strong.
The step most people skip, and the reason a vendor's own number is evidence of intent, not proof of quality.
U
Uncover the eval set.
No named lanes, no named carriers, no published test set behind the headline number at all.
A score means nothing if the eval set behind it isn't named.
D
Demand the version pin.
LaneCast changed at least three times since Northgale signed, with no record connecting any single booking to the model version that made it.
Models update quietly. A claim with no version pin can never be reproduced or checked.
Hand sketched icon list titled What a trustworthy claim needs. Four items: a named eval set, a version pin on the model, published failure cases, replicated on your own lanes.
LaneCast's marketing deck had none of the four.
I
Isolate what's missing.
No per-booking reasoning trail, no confidence signal, no published failure cases, no denominator behind the headline number.
What a report leaves out is usually more informative than what it shows.
T
Test it yourself.
Fifty of Northgale's own real bookings, run back through LaneCast, compared against what an experienced broker actually chose for the same lane.
The strongest move in the whole framework. Real evidence instead of a vendor's word for it.
Hand sketched quadrant titled How much should you trust a claim. Axes specificity of evidence and how much to trust it. LaneCast's 94 percent sits low specificity low trust. Version pinned score sits mid. Shadow tested claim sits high specificity high trust.
Duskrail's marketing deck lived in the bottom left corner the entire time.
Bookings with a flagged-safety-record carrier, per month
8 4 0 meetup incident Jan Feb Mar Apr May Jun
The rate didn't drop because LaneCast got smarter. It dropped once a required reasoning trail let Northgale actually see a flagged pick coming before it got booked.

The recap, one line per letter: ask who paid for it is Duskrail's own internal test, uncover the eval set is no named lanes or carriers anywhere, demand the version pin is three quiet updates with no record, isolate what's missing is no reasoning trail and no denominator, and test it yourself is the fifty-booking shadow test that found the real split.

And if you want to be sure it really works, try it somewhere elseSame five letters, a hospital's supply-reordering agent instead of a freight broker. A different vendor claim, the same shape of question.

Wrenhollow Health Systems uses a supply-reordering agent from a vendor claiming "97 percent stockout prevention" across its hospital customers, with no visible reasoning behind any single reorder decision.

Mapped onto AUDIT: ask who paid for it is the vendor's own case study, published on their own site, with no independent party involved. Uncover the eval set: the 97 percent doesn't say whether it was measured across all supply categories or only the easy, high-volume ones like gauze and gloves, leaving out specialty surgical items where a stockout is far more dangerous. Demand the version pin: the vendor has pushed at least two silent model updates in the past year, and no hospital using it can say which version produced last quarter's numbers. Isolate what's missing: no breakdown by supply category, no confidence signal on individual reorders, no record of near-misses caught by a human instead of the agent. Test it yourself: Cedarline should pull six months of its own specialty-item reorders and check the agent's actual performance there specifically, rather than trusting a blended hospital-wide average that hides exactly the category where a miss matters most.

Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "no eval set, no version pin, no reasoning trail means the number is decoration, test it yourself before trusting it," and stop.
Cost: if running your own fifty-lane shadow test feels expensive this quarter, start with ten from the highest-value lane category, since that's where a wrong pick costs the most anyway.
The model gets better, for real: even if LaneCast's real accuracy climbs to 99 percent, the case for a reasoning trail doesn't weaken, it strengthens, because the rarer a miss becomes the more it looks like a fluke instead of the pattern it might actually be.

Where people run it wrong.
They read a vendor's headline number as if it were an independent finding instead of a claim made by the party selling the product.
They accept a blended average without asking whether it hides a much worse number on exactly the cases that matter most.
They wait for something to go wrong before ever checking a claim, instead of testing it on their own data before it reaches production.

How to use it live. When someone hands you an agent's accuracy claim, ask yourself one thing out loud: if I asked for the eval set and the model version behind this number right now, could anyone in the room actually produce them.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that hands you a claim and asks if it's trustworthy?
Tap to flip
ANSWER
AUDIT: ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself.
2 · THE PEOPLE
Who are the two people this answer centers?
Tap to flip
ANSWER
Bastian Okwuosa, who runs operations at Northgale Freight Brokers, and Wystan Larrabee, a competitor whose remark started the whole audit.
3 · THE CLAIM
What claim is this answer actually auditing?
Tap to flip
ANSWER
Duskrail's marketing claim that LaneCast produces "94 percent on-time, cost-optimal bookings," with no named eval set behind it.
4 · WHAT IS MISSING
Name two things missing from LaneCast's headline claim.
Tap to flip
ANSWER
Any two of: a named eval set, a version pin, a per-booking reasoning trail, published failure cases.
5 · WHAT I'M CHALLENGING
What is this critique actually targeting, and what is it not saying?
Tap to flip
ANSWER
It's not saying LaneCast's matching is bad. It's saying nobody can currently tell whether it's good, and that not-knowing is the real problem.
6 · THE NUMBER
Fill in the blank: on rare or complex lanes, Northgale's shadow test found LaneCast agreed with a broker's pick only ___ percent of the time.
Tap to flip
ANSWER
58 percent, against 91 percent on simple lanes. The blended 94 percent hid that gap completely.
7 · THE REPLAY
Same flagged-carrier booking, with a required reasoning trail in place. What changes?
Tap to flip
ANSWER
The booking would show its own confidence level and the carriers it considered, and Northgale's monthly audit would catch the pattern within weeks instead of relying on a chance remark at a meetup.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what claim gets audited there?
Tap to flip
ANSWER
Wrenhollow Health Systems' supply-reordering agent, auditing a vendor's "97 percent stockout prevention" claim that hides performance on specialty items.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: LaneCast books a freight lane in under ___ , showing only a final carrier and rate.
Show hint
Look at the opening of "Let's learn."
Show answer
A minute. Speed was never the problem this critique targets. The missing reasoning trail is.
Multiple choice
2. Why does the 94 percent figure deserve suspicion, according to this answer?
  • A. 94 percent is a suspiciously low number for a booking agent.
  • B. Duskrail ran its own test with no named eval set or version pin, and the number blends easy and hard lanes together, hiding a much lower agreement rate on complex ones.
  • C. LaneCast has never been tested on real bookings at all.
  • D. Northgale's own brokers refuse to use the tool.
Show hint
Look at the shadow-test bar chart.
Show answer
B. The number is real but misleading. It averages across a mix dominated by simple lanes, hiding a much worse agreement rate on the rare, complex ones.
True or false
3. True or false: this critique concludes that LaneCast's carrier-matching model is definitely bad and should be replaced.
  • True
  • False
Show hint
Look at "what I'm actually challenging."
Show answer
False. The model might be fine. The critique is that nobody can currently tell, because there's no reasoning trail, no version pin, and no independent test behind the claim.
Short answer, where it wouldn't matter
4. Name something in this story that isn't actually part of the problem being critiqued.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The one-minute booking speed itself. A fast booking with a visible reasoning trail attached would fully resolve this critique.
Short answer, apply it yourself
5. Think of a product or service that advertises an accuracy or satisfaction statistic. What would you want to know before trusting it?
Show hint
Think about a "98 percent customer satisfaction" claim or a health app's accuracy number.
Show answer
Model answer: Who measured it, on what sample, and whether it was checked by anyone other than the company making the claim.
The number question
6. If the shadow test had shown 91 percent agreement on both simple and complex lanes instead of a split, would the critique still stand? Why or why not?
Show hint
Think about whether the critique is about the number or about the missing evidence.
Show answer
Model answer: Yes, largely. Even a genuinely strong, consistent number wouldn't fix the missing reasoning trail, version pin, or eval set. The critique is about what can be verified, not just what the score happens to be.
Before you close the answer
Why this works
Tests whether you can judge a vendor's evidence instead of reacting to a product design choice with a gut feeling. Most candidates say "that seems risky" without naming what specifically would make it trustworthy.
Follow-up traps
"What if Duskrail refuses to share their eval set, calling it proprietary?" Response: then Northgale runs its own shadow test instead, since a vendor's refusal to show evidence is itself evidence worth weighing.

"Isn't a reasoning trail just going to slow the booking down?" Response: no, generating the trail is a side effect of a decision the model already made, it doesn't require redoing the work, just showing it.
If pressed
Northgale's shadow test specifically stratified by lane type before averaging, rather than pooling all fifty bookings into one number, which is the exact step that revealed the 91-versus-58 split Duskrail's own blended figure had been hiding the whole time.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more