Critique an agent product that gives no visibility into intermediate steps.
Interviewer's question: "Critique an agent product that gives no visibility into intermediate steps." LaneCast, built by Duskrail Logistics, selects a carrier and books a freight lane automatically, showing brokers only a final line: carrier and rate. Bastian Okwuosa manages operations at Northgale Freight Brokers, one of LaneCast's customers.
- Ask who ran the 94 percent eval and whether Duskrail has any reason to inflate it, before trusting the number at all.Why: the vendor selling the agent is the same party who ran the test, and nobody else has checked their work.
- Get Duskrail to name the eval set behind that number, which lanes, which carriers.Why: a score with no named test set is a marketing line dressed up as evidence.
- Demand a version pin on every booking, so a bad call can be traced to the exact model version that made it.Why: models get updated quietly, and a report with no version pin can never be reproduced or checked.
- Replicate the claim yourself on fifty of your own real lanes before trusting it on the rest.Why: a vendor's aggregate score can hide a huge split between easy lanes and the rare ones where judgment actually matters.
- Require a visible reasoning trail on every booking going forward.Why: a booking agent with no way to see why is not fit to run unsupervised, whatever its score claims.
How to answer this, stage by stage
Nobody is grading whether you can spot that "no visibility" sounds bad. They're grading whether you can name exactly what evidence is missing and how you'd go get it yourself.
Let's learn
LaneCast reads an incoming freight request, picks a carrier, negotiates a rate, and books the lane automatically. The broker sees one line: carrier and rate, booked.
Before LaneCast, a Northgale broker spent about twenty minutes per lane manually checking a carrier's safety score, on-time history, and rate against two or three alternatives before booking.
Now LaneCast books a lane in under a minute, one line, no visible reasoning behind it at all.
Here's the turn: the speed was never the actual problem. The problem shows up the first time a bad pick happens, a carrier with a poor safety record gets booked, and Northgale has no way to know why it was chosen, whether that was a fluke or a pattern, or whether whatever caused it has even been fixed since.
At its worst: a driver-safety incident happens on a lane LaneCast booked with a carrier that had a flagged record, and Northgale can't produce a single document showing why that carrier was chosen, exposing them to real liability with no paper trail to defend the process at all.
What I would leave alone: the speed of the booking itself isn't the issue. A one-minute booking with a visible reasoning trail attached would solve this completely; the critique is about the missing trail, not about automating the task in the first place.
The lesson: a headline accuracy number with no eval set, no version pin, and no reasoning trail isn't proof a system works. It's a claim waiting to be tested, and nobody had tested it yet.
Now here is the same thing as a story
The short version above is what you'd say critiquing this product to Duskrail's own account team. Read this one for how Northgale actually started asking the hard questions.
Bastian Okwuosa had run operations at Northgale Freight Brokers for eleven years, long enough to remember booking every lane by phone, carrier by carrier, checking safety scores by hand.
LaneCast made that entire process disappear. A lane came in, LaneCast booked it, and the confirmation read like a receipt: carrier name, rate, done. For the first year, Bastian mostly stopped thinking about it. The bookings looked fine. Nothing had gone visibly wrong.
Then, at an industry meetup, Wystan Larrabee, who ran operations at a competing brokerage using a similar booking agent, mentioned almost in passing that his own system had booked a lane with a carrier flagged for a safety violation two months earlier, and nobody at his company had known until a client asked about it directly.
Bastian went back and asked Duskrail's account team a simple question: what's actually behind the "94 percent on-time, cost-optimal" number on the sales deck. He got a paragraph about "rigorous internal testing" and nothing else. No named lanes. No named carriers. No date. No model version.
He pushed further and learned LaneCast had been updated at least three times since Northgale signed, informally, through a changelog email nobody at Northgale had actually read closely. There was no record anywhere connecting a specific booking to the specific model version that made it. If a bad booking had happened in March, nobody, not even Duskrail, could say for certain which version of the agent made that call.
Rather than keep arguing with Duskrail's account team, Bastian pulled fifty of Northgale's own real bookings from the last quarter and ran them back through LaneCast in a shadow test, comparing its carrier pick to what an experienced human broker had actually chosen for the same lane at the time.
The decision I'd challenge here is Duskrail's, not Northgale's: choosing to show one line, "Booked," instead of any reasoning, because it made the product feel simpler and more impressive in the sales demo. That looked good in a five-minute pitch. It stopped looking good the first time a bad booking needed an explanation nobody could produce.
Replayed with a required reasoning trail in place: the flagged-carrier booking that worried Wystan would have shown its own confidence level and the carriers it considered instead, and Northgale's monthly audit would have caught the pattern within weeks instead of relying on a chance conversation at an industry meetup.
I used to think a clean, simple confirmation screen was a sign of a mature product; less clutter, more trust. It took one flagged carrier and a conversation with a competitor to see that the missing clutter was actually the evidence, and nobody had ever been able to check the work at all.
AUDIT, spelled out for one claimNot a takedown of LaneCast specifically. AUDIT is what tells you whether any claim like this deserves your trust yet.
The recap, one line per letter: ask who paid for it is Duskrail's own internal test, uncover the eval set is no named lanes or carriers anywhere, demand the version pin is three quiet updates with no record, isolate what's missing is no reasoning trail and no denominator, and test it yourself is the fifty-booking shadow test that found the real split.
And if you want to be sure it really works, try it somewhere elseSame five letters, a hospital's supply-reordering agent instead of a freight broker. A different vendor claim, the same shape of question.
Wrenhollow Health Systems uses a supply-reordering agent from a vendor claiming "97 percent stockout prevention" across its hospital customers, with no visible reasoning behind any single reorder decision.
Mapped onto AUDIT: ask who paid for it is the vendor's own case study, published on their own site, with no independent party involved. Uncover the eval set: the 97 percent doesn't say whether it was measured across all supply categories or only the easy, high-volume ones like gauze and gloves, leaving out specialty surgical items where a stockout is far more dangerous. Demand the version pin: the vendor has pushed at least two silent model updates in the past year, and no hospital using it can say which version produced last quarter's numbers. Isolate what's missing: no breakdown by supply category, no confidence signal on individual reorders, no record of near-misses caught by a human instead of the agent. Test it yourself: Cedarline should pull six months of its own specialty-item reorders and check the agent's actual performance there specifically, rather than trusting a blended hospital-wide average that hides exactly the category where a miss matters most.
Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "no eval set, no version pin, no reasoning trail means the number is decoration, test it yourself before trusting it," and stop.
Cost: if running your own fifty-lane shadow test feels expensive this quarter, start with ten from the highest-value lane category, since that's where a wrong pick costs the most anyway.
The model gets better, for real: even if LaneCast's real accuracy climbs to 99 percent, the case for a reasoning trail doesn't weaken, it strengthens, because the rarer a miss becomes the more it looks like a fluke instead of the pattern it might actually be.
Where people run it wrong.
They read a vendor's headline number as if it were an independent finding instead of a claim made by the party selling the product.
They accept a blended average without asking whether it hides a much worse number on exactly the cases that matter most.
They wait for something to go wrong before ever checking a claim, instead of testing it on their own data before it reaches production.
How to use it live. When someone hands you an agent's accuracy claim, ask yourself one thing out loud: if I asked for the eval set and the model version behind this number right now, could anyone in the room actually produce them.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a reasoning trail just going to slow the booking down?" Response: no, generating the trail is a side effect of a decision the model already made, it doesn't require redoing the work, just showing it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Agent product management specifics
- #1 What product decisions are unique to an agent versus a single-turn AI feature?
- #2 How do you scope what an agent is allowed to do?
- #3 Describe the permission model you would design for an agent acting in a user's account.
- #4 What does success look like for an agent, and why is task completion insufficient?
- #5 How do you evaluate an agent's trajectory rather than its final answer?
- #6 Explain the product implications of an agent that takes 40 steps instead of 4.