How do you choose pilot customers, and what makes a bad one?
- Run the frozen model cold on a random sample of sellers who don't look like the pilot customer, before signing off on general availability.Why: it's the one check that would have caught the gap before every seller on the platform felt it.
- Recut the pilot's own numbers by how closely the pilot customer matches the real seller base, not just by the headline rate.Why: a blended number that looks fine can be one seller type running fine and everyone else drowning.
- Pick a pilot customer who looks like the messy middle of the real base, not the cleanest account on the books.Why: the cleanest account is exactly the one whose data was never going to expose a fraud model's blind spots.
- Separate the pilot's white-glove support channel from what a general rollout actually gets.Why: a direct line to the build team resolves every borderline case fast, and nobody else gets that line.
- Check whether the pilot customer's own team pre-cleaned every submission before the model ever saw it.Why: a sophisticated team can make an average model look flawless just by never sending it a messy case.
- Rule out the model itself before blaming the new population.Why: confirms whether the fix is retraining the model or picking a better pilot next time, so effort lands in the right place.
How to answer this, stage by stage
Six moves. This question tempts a general lecture about picking friendly pilot customers, so most of these stages exist to prove exactly how one clean seller hid a fraud tool's real weak spot, and to name the test that would have caught it.
Let's learn
What does a fraud tool need to see before it can tell risky from ordinary? Vetlock is Brackenford Marketplace's answer. It watches every sale between two regular people and holds the ones that look risky before the seller gets paid.
Before Vetlock, Brackenford's small trust and safety team read a sample of sales by hand, maybe 200 a week, guessing which ones felt off. Real fraud slipped past more often than anyone wanted to admit.
Teodor Kessling piloted Vetlock for six weeks on one seller: Dorotea Voskuijlen, who runs a professional watch-resale shop, about 45 sales a week, every watch photographed, every serial number logged, every shipment insured. Vetlock held only 1 percent of her sales for review, and it never once held a clean one or missed a bad one. Baltasar Kessendra, Brackenford's head of trust and safety, signed off on the full rollout that same afternoon.
Here is the turn. That climb from 3 to 19 percent was never the real problem, and it never was. The real problem showed up the moment somebody split that number apart by seller type instead of reading it as one number for the whole marketplace.
At its worst, this costs Brackenford about 2,600 sellers abandoning a listing rather than wait days for a payout hold to clear, pulling roughly $180,000 of inventory off the marketplace in Vetlock's first month live, and it nearly gets the whole tool switched off two weeks before quarter close, taking Vetlock away from the sellers it was actually protecting.
The choice I would take back. Teodor's team picked Dorotea because she was Brackenford's easiest seller to work with, high volume, fully verified, a direct line to his own team, and never wrote down that this made her the cleanest account on the books, not a fair test. I would take that back. I would put one line in the go/no-go review: this pilot ran on our most verified seller; a random batch of ordinary sellers, run cold, is the real bar, before anyone signs off on rollout.
What I would leave alone. Dorotea's own account doesn't need any of this. Her hold rate stayed at 1 percent the whole five weeks, exactly what the pilot promised. Slowing down her account to fix a problem she never caused would just cost the sales Vetlock is already getting right.
The lesson. A pilot that never once looks wrong isn't proof a tool is ready. It's proof nobody has shown it a seller it doesn't already recognize, and the room signing off has no way to tell the difference.
The morning Baltasar said yes, and the five weeks after
Read the short version above if you're short on time. This is the long version, for the part where you feel exactly how close it came to shipping broken.
The support queue at Brackenford Marketplace usually quiets down by midnight. In Vetlock's fifth week live, it didn't.
Teodor Kessling had spent the better part of a year listening to the trust and safety team complain that they were guessing. Two hundred sales a week, read by hand, and everyone knew the real number of scams was higher than what they caught. Vetlock was supposed to fix that: score every sale, hold the risky ones, let the rest go straight through.
He needed one seller to pilot it on, and Dorotea Voskuijlen was the easiest choice in the building. Her shop, Ambergate Timepieces, moved fast. She answered every message within minutes. Every watch she sold already had its serial number logged and a photo of the authentication mark before Vetlock ever asked for one. For six weeks, Vetlock ran quietly against her account. It caught two real attempts, both confirmed fraud, and cleared everything else clean. When something looked borderline, Dorotea's team texted Teodor directly, and it got sorted within the hour.
By the readout meeting, Teodor had a deck with one line that mattered: zero missed fraud, zero false holds, six weeks straight. Baltasar Kessendra, watching from the head of the table, didn't ask a single follow-up question. He approved the full build before the meeting even ended. Vetlock would go live for every seller on the platform, three weeks out.
For the first two weeks after launch, the numbers looked fine. The hold rate, the share of sales Vetlock stopped for review, sat at 3 percent, then 6. Nobody was watching closely. It read exactly like a new tool finding its feet.
By week four it was 15. By week five, 19. Still, on paper, a number you could explain away. Nobody had split it by seller type, because nobody had a reason to look.
Then, at 11:40 that Thursday night, the support lead pinged Teodor. A spike of tickets, all some version of "why is my payout on hold," more in one evening than the whole first month combined. Teodor's first instinct was to wonder if sellers were just adjusting to a new tool. So he checked the easy thing first, the model version. Same frozen build, running everywhere, nothing had changed since the pilot ended.
So he split the number by seller type instead. Verified business sellers, the pattern Dorotea's shop matched, held at 1 percent, exactly the pilot's number. Everyone else, 94 percent of every seller account on the platform, held at 34 percent.
He pulled up the recording of his own readout to Baltasar. Twelve minutes in, he heard himself say it: "Vetlock is ready to run on every seller, out of the box." He'd meant it about the account he'd tested. Nobody in that room had any way to know that.
He took one real transaction, still sitting in the queue, a first-time seller's guitar listing that had never come anywhere near the pilot, and ran it live in front of his own team. It queued, the same way Dorotea's near-misses never had, and came back four days later, cleared, no fraud found at all. Then he reshaped the same charges to carry Dorotea's pattern, verified ID, a sale history, insured shipping selected by default, and ran it again: cleared in 0.4 seconds, no queue.
So here is the decision I would take back. When Teodor's team picked a pilot customer, they picked the seller who made the pilot easiest to run, fast replies, clean paperwork, a direct line back to the build team. What nobody did was write down that this made her the least representative account on the platform, and run even one seller who didn't look like her, cold, before a whole marketplace's trust in the tool was riding on it.
And the part I'd want to tell myself, if I could go back: we tested Vetlock against the seller who would never once make it look bad. We never once tested it against the seller it was actually built to protect.
What the recut actually showed
Before trusting the seller-type gap, Teodor's team checked whether Vetlock's own scoring was even right. Two trust and safety reviewers hand-checked 20 of the held individual-seller transactions against what they'd have flagged themselves. They agreed with Vetlock's own risk score on 19 of 20. The model wasn't the problem. That left the population it had learned to expect.
Three reasons a pilot customer goes wrong, and the one that was true
Not because anyone was careless. Each of these, on its own, looks like a sensible way to pick a pilot customer. Together, they're why a pilot that never once looked wrong can still make a promise the real product can't keep.
Dorotea's shop moved roughly fifteen times more volume than a typical account, every watch photographed and logged before Vetlock even needed it, every shipment insured by default. Nothing about a couch sold once by someone who's never listed anything before looks anything like that.
Dorotea's team had a direct line to Teodor's group the whole six weeks. Any hold that looked borderline got a reply and a fix within the hour. A first-time seller waiting in the regular queue gets none of that.
Dorotea's own staff authenticated every watch, logged every serial number, and picked insured shipping before a single listing ever reached Vetlock. The model rarely had to work hard, because it was never shown a messy case.
Running TRACE against the seller Vetlock never had to worry about
This reads like a question that wants a general rule about friendly pilot customers, but the real job is diagnosis: work out why a pilot that never once looked wrong could still walk a marketplace into a tool that broke for everyone else, and prove exactly where that promise broke.
Same blind spot, a contractor who never once needed managing
Tannerbrook Homeworks, a home-repair booking platform, pilots CrewCheck: an AI that screens new contractor sign-ups for fraud, forged licenses, stolen insurance certificates, before they can accept a job. Rune Bergeron runs Bergeron Fireplace & Chimney, a long-established company with a full-time office admin who submits pristine paperwork every time. CrewCheck cleared every one of Rune's applications instantly across an eight-week pilot, and rolled out to all new contractor sign-ups four weeks later.
T. Rune's pilot ran eight weeks; the platform approved full rollout the same week it ended. New contractor sign-ups started going through CrewCheck three weeks later. The blended hold rate crept from 4 to 22 percent over the next five weeks, and nobody split it apart until a regional manager escalated a wave of stalled sign-ups.
R. Recut by company size. Established, multi-crew companies like Bergeron's: held at 2 percent, steady the whole time. Solo contractors, who make up most new sign-ups: held at 41 percent.
A. Same frozen model scored both groups. A manual review of 15 held solo-contractor applications agreed with CrewCheck's own risk score on 14. The grading held up. Rune's own team had pre-scanned and reformatted every document before it ever reached CrewCheck; a solo contractor photographing a license on a truck dashboard gets no such help.
C. Three candidates, the same shape as before: only an established, fully-staffed company was ever piloted, chosen because its paperwork was cleanest; no time was set aside to test a solo contractor's messier submission; and the demo audience was the platform's regional directors, deciding whether to fund the tool company-wide, and a visible false flag would have read as "not ready."
E. One real solo contractor's application, never in the pilot, queued and came back six days later, cleared with no fraud found. Reformatted with Bergeron's pattern, office-scanned documents, an established company number, an admin's cover note, and rerun: cleared in under a second. Same model, only the paperwork's shape changed, which pointed straight at who CrewCheck had been piloted on, not at a model that can't verify a license.
Swap the trigger and it still runs
- Speed: Baltasar could have pushed Vetlock to all 14,000 sellers in one week instead of three, to hit a fraud-loss target before quarter close. TRACE still starts by asking what shipped at pilot time and when the real friction reached someone, not by how fast the rollout ran.
- Cost: the team could have skipped the cold test to save two days before the go/no-go review. The check still has to happen eventually, just after $180,000 of inventory walks off the marketplace instead of before it does.
- The model really did get better: say Vetlock's next version genuinely got sharper at catching real fraud, company-wide, in the very same stretch a new seller pattern still caught it flat-footed. TRACE still finds the gap, because the recut isolates one seller group even while the overall trend looks like good news.
Where people run it wrong
- Trusting a blended hold rate that's still technically "mostly fine," without ever cutting it apart by seller type.
- Treating a wave of false holds as proof the model needs more training, before checking whether the sellers even matched who got piloted.
- Fixing the visible symptom, retraining on more fraud examples, instead of the actual gap: who the pilot ran on, and who it never touched.
How to use it live
Buy yourself ten seconds by naming the split out loud. "So there's the pilot everyone remembers, and there's whoever never got tested. Let me say how I'd check whether that gap is already showing up." That's not stalling. That's where the real answer starts.
Flashcards (click a card to flip it)
This is a question about picking pilot customers, worked as a diagnosis, so these eight test the TRACE moves and the real numbers behind them.
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pilot design and POC-to-production
- #1 Design a four-week pilot for an AI feature with one enterprise customer.
- #2 What success criteria should be agreed before a pilot begins?
- #3 Explain the difference between a pilot and a beta.
- #5 Describe the pilot-to-production gap and the work that lives in it.
- #6 Why do most AI POCs fail to reach production? Give four reasons.
- #7 What data do you need to collect during a pilot that you would not otherwise?