CaseAdvancedShipping & Model Lifecycle / Pilot design and POC-to-production / #8

How do you handle a pilot that succeeds on quality but fails on cost?

The direct answer
Push forward on the current setup and fix cost in parallel when the overrun is small, roughly under three times what you budgeted. Once it crosses that line, the way one pilot did at six times budget, pause the rollout and fix the cost path before anyone else gets the same deal. The business is the one who quietly eats a bad cost surprise at scale, not the customer who notices a pause.
How to handle a quality-win, cost-miss pilot, in order
  1. Set a hard multiple as the pause line, and pause now if you're past it.Why: a pick with no threshold number is just a mood, not a decision anyone can act on.
  2. Write a hard cost ceiling into the pilot's exit criteria, the same way quality got one.Why: quality had a number that would fail a report; cost had no number to fail against, so nothing forced anyone to look.
  3. Name who feels each kind of miss before you commit.Why: a paused rollout costs the customer a few visible weeks; an unpaused one costs the business a quiet loss that compounds with every deal signed after.
  4. Run a time-boxed cost audit before assuming a full redesign is needed.Why: the kill criteria only means anything if there's a real deadline on finding out whether the overrun is fixable or built into the architecture.
  5. Leave the already-priced, already-proven parts of the product alone.Why: not every capability needs the same gate; gating a proven cost model wastes the pause on the wrong target.
  6. Reset the price or the pause the moment the audit's answer comes back, whichever the evidence points to.Why: a pick that can't be reversed by real evidence is stubbornness wearing a framework's clothes.

How to answer this, stage by stage

This is a straight tradeoff between two responses to the same overrun, not a step-by-step build, so PICK carries the weight here.

1
Scope it to one real decision
Say it like this
"Let's make this concrete. Say Loomwell Labs builds ReelText, a tool that watches a video, transcribes what's said, and writes the closed captions. The team just finished a new capability that can caption real documentary footage, people talking over each other, music under the narration, cutaways, not just clean interview audio. Before selling it to anyone else, we ran a pilot with one customer, and now I've got two very different numbers on my desk."
Why this works
Stops the answer floating at "it depends" and gives the interviewer one real feature to push on.
2
Say your structure out loud
Say it like this
"I'll give you my pick first, then who actually feels each kind of miss, then which one costs more even though it's the quiet one, then what would change my mind. That's PICK, and I'll go in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Reframe what the question is really testing
Say it like this
"This isn't really 'do you ship a pilot that ran over budget.' It's 'which failure would you rather manage in the open: a customer noticing you paused, or your own company quietly losing money on every minute it processes.' Those two need completely different responses."
Why this works
Shows the interviewer you see past the surface ask to the judgment actually being tested.
4
State the position, with the real magnitude in it
Say it like this
"My pick: if the pilot's real cost lands under roughly three times what we budgeted, keep shipping on the current setup and fix cost as a fast follow, the quality bar was the harder problem and I don't want to reopen it. This pilot came in at six-point-four times budget, nine cents a minute planned, fifty-eight cents actual, so I'd pause the rollout now and fix the cost path before we sign the four studios waiting behind this one."
Why this works
PICK rewards a real threshold and a real number, not a vague promise to "watch spend more closely."
5
Name who feels each kind of miss
Say it like this
"Here's the split. If we pause, the customer notices, their launch slips maybe three weeks, their post-production lead asks me why, and I have a real answer. If we push forward instead and just promise to fix cost later, nobody notices anything, right up until finance reconciles the cloud bill at quarter's end and finds we're paying twenty-three cents more per minute than we're charging, on every studio signed since."
Why this works
Turns "cost asymmetry" from a phrase into two people who would actually feel it.
6
Say what you'd leave alone
Say it like this
"I wouldn't gate the version of ReelText we already sell, clean-dialogue captioning, its cost model has been proven and billed for a year. And I wouldn't hold up something small, like a caption font tweak, behind any of this. That's cheap to change wherever it lands."
Why this works
Shows judgment instead of freezing the whole product the moment one number looks bad.
7
Name the kill criteria and close
Say it like this
"I'd flip this back to push-forward-and-fix if a two-week cost audit shows the overrun is mostly one fixable thing, like redundant per-frame model calls we can batch or cache, instead of something baked into the architecture. Until that evidence exists, six times budget doesn't get patched on the fly. It gets paused."
Why this works
Ends on the line the interviewer remembers, and shows the pick can be reversed by evidence, not stubbornness.

A last note before the walkthrough ends: this pick is about a threshold and a cadence, not a rule that cost always wins over quality. Most candidates hear "succeeds on quality, fails on cost" and reach for a vague "balance the two." Name the magnitude and who feels each kind of miss, and you've shown judgment instead of reciting a platitude.

Let's learn

ReelText is a tool that watches a video and writes its captions, matching every word to the moment someone actually says it.

Knowledge spark: what's actually driving the cost? Clean, single-speaker audio is cheap to caption, the model listens once and moves on. Real documentary footage has people talking over each other, music under the narration, and cutaways to something else while a voice keeps going. Getting that right means checking a lot more of each frame, not just the audio track, and that checking is what runs the bill up.

Before ReelText, an editor at a media company captioned footage by hand: watch it, type it, fix the timing, watch it again. For a documentary-style hour of footage, the kind with crosstalk and music and cutaways, that took about three hours.

The version of ReelText already sold to everyone cuts clean-dialogue captioning to about 40 minutes an hour. It's been out a year and sells fine, but it chokes on real documentary footage, the overlaps and the music confuse it, so a media company can only trust it on a fraction of what they actually shoot.

Then Loomwell Labs built something harder: a capability that captions real documentary footage, crosstalk, music, cutaways, all of it, correctly. Isidro Renner, the senior product manager who owns ReelText, picked one customer, Hollow Pine Studios, to try it first, on real editorial footage, before anyone else got the same deal.

Here is the important part. Getting the words right was never the hard part of this test. Isidro's team cleared that bar in the first two weeks, 98.7 words out of 100 correct, caption timing tight enough nobody noticed a lag. The real test was whether ReelText could do all that watching and matching without costing more than the caption was worth. It couldn't. Not yet.

We didn't lose the pilot on the captions. We lost the budget on a number nobody was watching.

At its worst: Loomwell planned to charge about thirty-five cents a minute of processed footage, priced against a budgeted cost of nine cents. The pilot's real cost came in at fifty-eight cents a minute, six-point-four times over. Four more documentary studios are waiting behind Hollow Pine for the same deal. Sign all five at that price and Loomwell isn't selling a caption tool anymore. It's paying customers to use one, about $3,400 a month once the queue signs, and climbing every time sales closes another account.

Budgeted cost vs. the pilot's real cost, per minute processed
$0.09 / min $0.58 / min Budgeted Hollow Pine's actual pilot cost
The green bar is what Loomwell priced the new capability against: nine cents a minute. The red bar is what the pilot actually cost: fifty-eight cents a minute, six-point-four times over, and that gap is what gets multiplied across every studio signed at the old price.
The choice I would take back Quality had a hard number written into the pilot's exit criteria, 97 percent accuracy, checked by an automatic weekly report. Cost had no number at all, just an assumption that caching would bring it down. There was nothing for a report to fail against, so nothing forced a look until finance caught it by accident. I'd write a cost ceiling into the exit criteria next time, with the same automatic check quality already had.

What I would leave alone. The version of ReelText we already sell, clean-dialogue captioning only, doesn't need a second look. Its cost model has been proven and billed for a year. This is only a live question for the new capability nobody's priced yet.

The lesson. A pilot only proves the number that has a bar written against it. Quality had one and it held. Cost never got one, so nothing forced anyone to look until the bill did it for us. Next time both numbers get a real ceiling, or neither one is really proven.

Now here is the same thing as a story

You don't need this to answer the question. Read it if you want to feel why the pick has to hang on a written threshold, not on how good the calls were going.

Every Monday, before anything else, Isidro pulled up two numbers: how many words ReelText got right that week, and how much it cost to get them.

He'd run ReelText at Loomwell Labs for five years, since before it could do much more than swap a name into a template line. He knew what a good pilot looked like, and this one looked good. Junie Sanborn, who runs post-production at Hollow Pine Studios, joined every Monday call. She'd flag two or three lines that clipped a word during a crosstalk moment, Isidro's engineer would patch it by Thursday, and by the second week the caption accuracy number read 98.7 and held there.

At the kickoff meeting, someone on Isidro's own team had asked whether they should track the processing cost every week too, the same way they were about to track accuracy. Isidro remembers saying, "let's not slow down the quality work chasing a number that should settle once the caching layer warms up." It felt reasonable. The team was small, and quality was the thing everyone was nervous about.

So the cost number went unwatched. Week two. Week three. Junie's notes kept getting shorter, which read as progress, because it was progress, on the one number anyone was checking.

Then, on a Thursday in the sixth week, Loomwell's finance controller sent the routine monthly cloud reconciliation to the whole product team, the same email that went out every month, nothing dramatic about it. Isidro opened it between meetings. The processing cost for Hollow Pine's pilot wasn't the nine cents a minute it had been budgeted at. It was fifty-eight cents.

It wasn't the fifty-eight cents that got him. It was that nobody had looked at this number since the day the pilot started.
Hand-sketched comparison in colour pencil. Left, a small green document icon labeled Hollow Pine's launch slips, captioned seen right away, account team manages it. Right, a red-orange gauge icon labeled Loomwell eats the cost gap, captioned hidden until the cloud bill lands.
Same overrun, two very different sizes of miss

He pulled up the caption accuracy chart out of habit, the way you check a wound you already know isn't the problem. It still read 98.7. It had read 98.7 for a month. That number had never once been in danger. The danger had been sitting in a spreadsheet nobody opened.

So here's the decision Isidro would take back. Not the choice to pilot the new capability with one customer first, that part was right. It's the choice, made in that kickoff meeting, to write a hard number into quality's exit criteria and give cost nothing at all, no ceiling, no bar, no number a report could ever fail against. "It should settle once caching warms up" sounds like a technical judgment. It's actually a guess standing in for a number that was never written down.

He paused the rollout to the four studios waiting behind Hollow Pine that same afternoon, and gave his team a two-week audit before committing to anything bigger: find out whether fifty-eight cents was a fixable mistake or a structural one.

It was mostly fixable. ReelText's new capability had been calling the vision model fresh on every single frame, even during long stretches of clean dialogue where nothing in the picture was actually helping the caption. Batching those calls and caching repeated visual context, the same cutaway reused across three edits, brought week one down to forty-one cents. Reserving the expensive full check for only the hard stretches, the actual crosstalk and music, brought week two down to twenty-four cents, under the three-times line that would have let them push forward in the first place. By week three it was eleven cents, just above target, and accuracy had dipped only to 98.3, still comfortably clear of the 97 percent bar.

And the thing I'd tell myself, back at that kickoff meeting: a number you don't check isn't a number that's fine. It's just a number you haven't met yet.

PICK, priced out to the minute

This is a tradeoff about which failure to manage, not a rule for every capability Loomwell ever ships, so PICK carries the weight here.

P, position. What's the pick, stated first? Commit before the reasoning, since "watch it closely" isn't a real pick. → Push forward and fix cost in parallel under roughly three times budget. Past that line, the way this pilot ran at six-point-four times budget, pause the rollout and fix the cost path before signing anyone else.
I, impact. Who feels each kind of miss, and in what units? Both sides get named, not just the louder one. → A paused rollout is felt by Hollow Pine, whose launch slips about three weeks and whose post-production lead asks a question Isidro can actually answer. A pushed-forward rollout is felt by Loomwell's own finance side, absorbing a loss on every minute processed, invisible until the cloud bill lands.
C, cost asymmetry. Which miss is cheap and which is hidden? Optimise against the one that's quiet and expensive, not the one that's loud and cheap. → The paused launch is cheap and visible, a customer notices, an account team manages it, it costs weeks. The unaudited cost overrun is hidden and expensive, it doesn't show up as one bad invoice, it shows up as a business quietly losing money on every studio it signs.
K, kill criteria. What evidence flips the pick? This is what separates a confident answer from a stubborn one. → Flip pause back to push-forward-and-fix if a two-week cost audit gets the overrun under roughly three times budget, proving it was one fixable thing, not the architecture. Flip a live capability back into a paused pilot the moment a new client's stakes are different enough that a hidden cost gap would do real damage before anyone caught it.

Recap, one line per letter: P name the threshold and the actual number. I name who feels each side. C the quiet miss usually costs more than the loud one. K let a real, time-boxed audit change your mind.

Knowledge spark: why not just lower the price instead of pausing? Because a lower price hides the same problem instead of fixing it. Loomwell would still be running an expensive pipeline, just charging less to cover it, which loses money faster, not slower. Fixing the cost, not the price, is what actually lets ReelText scale to more studios.
Cost per minute during the two-week audit and the fix that followed it
Actual cost per minute, by week
Kill line: $0.27, three times budget
$0.00 $0.30 $0.60 $0.27 kill line $0.58 $0.41 $0.24 (under line) $0.11 Week 0, pause Week 1 Week 2 Week 3
Cost per minute fell from fifty-eight cents at the moment of the pause to forty-one cents after batching redundant frame calls, then twenty-four cents once the full model only ran on the hard stretches, crossing under the twenty-seven cent kill line in week two. By week three it settled at eleven cents, just above the nine cent target, with accuracy still holding at 98.3 percent.

And if you want to be sure it really works, try it somewhere else

Naveen Pillai runs product at Barkline Diagnostics, a company that builds ScanTriage, a tool that flags anomalies on veterinary X-rays before a vet even opens the file. About 40 referral clinics use it.

Most X-rays ScanTriage sees are routine, a healthy hip, a clean chest film. Nobody's hurt if that triage call takes a beat longer. Naveen priced that tier at forty cents a scan a year ago and it's held steady since. But the team also built a harder capability: full high-resolution review for scans that might hide something small, an early fracture line, a mass the size of a coin. That capability just ran its own pilot, and it came in at two dollars and ten cents a scan, three-and-a-half times its sixty-cent budget.

P. Push forward, but split ScanTriage into two tiers immediately: a cheap first-pass model on every scan, the full high-resolution model only on the roughly 30 percent it flags as uncertain. At three-and-a-half times budget, close enough to the pause line, a targeted routing fix beats stopping everything.
I. A vet who gets a fast, clean triage call on a routine scan never notices anything unusual. Barkline's finance side feels an unrouted rollout the same way Loomwell would, a quiet monthly loss that only shows up once all 40 clinics are billed at the old rate.
C. A wrongly cheap-tiered easy scan is cheap and visible, worst case a vet asks for a second look and gets one within the hour. A silently expensive rollout across every clinic is hidden until the invoice lands, and by then contracts are already signed at the wrong price.
K. Flip straight back to a full pause the moment the cheap first pass misses a real anomaly, not just costs too much. The tiering only works while the cheap tier's mistakes are false alarms, never missed cases.

What I would leave alone, at Barkline The layout of the triage alert itself, whether it's a banner, a badge, or a line in a report, stays whatever's cheapest to change. That's a question about how a vet likes to see a flag, and it has nothing to do with whether the flag itself can be trusted.
What one wrongly cheap-tiered scan costs, caught by the routing check vs. missed entirely
Caught by the routing check, same day
$140
Missed, found at a client clinic's own audit
$8,600
Both start with a similar base cost, about $140 in review time either way. The green sliver is what the routing check catches early, a vet asks for a second look the same day. The red portion is what shows up instead if nothing's watching the tiering: a client audit, a review of every scan that month, and an account that came close to leaving Barkline over one missed case.

Swap the trigger and it still runs

  • Speed: if Hollow Pine needed the new capability live in two weeks instead of a quarter, the pick doesn't move, a two-week cost audit is exactly what a rushed launch still has time for; a full quarter spent hoping the cloud bill would settle never did.
  • Cost: if the overrun had been two times budget instead of six, the pick flips, push forward and fix cost as a fast follow, because a gap that small is closer to a tuning problem than an architecture one.
  • The model got better: if Loomwell's own general-purpose ReelText version already proved the same frame-grounding could run on a cheaper base model with no quality loss, that's exactly the evidence that sends the new capability straight to a wide rollout, no pause needed at all.

Where people run it wrong

  • Treating "quality passed" as the whole test, and shipping cost unaudited because the harder-looking number already cleared.
  • Checking cost once at kickoff and assuming it'll settle as the pipeline "warms up," instead of watching it on the same cadence as quality.
  • Pausing everything, including the parts of the product that were never expensive, instead of naming exactly which capability needs the gate.

Buy yourself two seconds, out loud

Say the reframe before answering with a verdict. "Give me a second, I want to separate two different failures here: the quiet one nobody notices until the bill arrives, and the loud one a customer notices in a week. Those need different responses." That's true, it's already stage three of the walkthrough, and it buys you the time to find the real asymmetry instead of just picking a side.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's the hardest step to nail?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a business quietly absorbing a bad unit-economics surprise at scale is worse than a customer noticing a pause, even though the pause is the one people feel first.
2 · THE PERSON
Who decides how to handle ReelText's over-budget pilot, and what has he done for five years?
Tap to flip
ANSWER
Isidro Renner, the senior product manager who has run ReelText at Loomwell Labs for five years.
3 · THE HABIT
What number did quality get written into the pilot's exit criteria, and what did cost get instead?
Tap to flip
ANSWER
Quality got a hard 97 percent bar with an automatic weekly report. Cost got no number at all, just an assumption that caching would bring it down on its own.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A paused rollout: Hollow Pine's launch slips about three weeks, seen and managed. An unpaused rollout: Loomwell absorbs a hidden loss on every minute processed, caught only when finance reconciled the cloud bill in week six.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Push forward and fix cost in parallel under roughly three times budget; past that line, the way this pilot ran at six-point-four times, pause the rollout and fix the cost path first.
6 · THE NUMBER
Loomwell budgeted ReelText's documentary-grade captioning at $0.09 a minute. The pilot's real cost came in at $______ a minute.
Tap to flip
ANSWER
$0.58 a minute, about six-point-four times target, the number that pushed the pick past the pause line.
7 · THE KILL CRITERIA
What evidence would flip this pick?
Tap to flip
ANSWER
A two-week cost audit showing the overrun is mostly one fixable thing, like redundant per-frame model calls that can be batched or cached, rather than something built into the architecture.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the position land there?
Tap to flip
ANSWER
ScanTriage, a vet X-ray triage tool at Barkline Diagnostics. At a smaller three-and-a-half times overrun, the position lands on push forward with an immediate routing fix, not a full pause.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Loomwell budgeted ReelText's new documentary-grade captioning at about $______ a minute.
Show hint
It's the number from the original plan, before the pilot's real cost came in six times higher.
Show answer
$0.09. Nine cents a minute was the target the whole pilot was measured against, and it's the number the pilot missed by six-point-four times.
Multiple choice
2. Which of these is the actual mechanism behind this answer's pick?
  • A. Lower the price so the unit economics work out even at the higher cost.
  • B. Push forward and fix cost in parallel under roughly three times budget; past that line, pause the rollout and fix the cost path before anyone else gets the same deal.
  • C. Cancel the new capability and keep selling only the clean-dialogue version.
  • D. Keep signing new studios at the pilot price while engineering works on cost "in the background."
Show hint
Three of these either give up on the new capability entirely or let the hidden loss keep growing while nobody's watching it.
Show answer
B. A gives up margin instead of fixing anything. C throws away a capability that already cleared the hard bar. D is exactly the "push forward and hope" path that let the cost go unwatched for six weeks the first time. Only B ties the response to a real threshold.
True or false
3. True or false: this position means Loomwell should never run this capability on the current architecture again, even after the fix.
  • True
  • False
Show hint
Think about what the two-week audit actually found, and what happened by week three.
Show answer
False. The audit found the overrun was mostly fixable, batching redundant frame calls and reserving the expensive check for hard stretches brought cost down to eleven cents by week three. The position is about pausing until there's evidence either way, not banning the architecture outright.
Multiple choice
4. Why didn't Isidro catch the cost overrun until week six, even though quality had already cleared the bar by week two?
  • A. He checked quality every week but treated cost as something that would settle once the pipeline's caching warmed up, so nobody watched it the same way.
  • B. Loomwell's finance team never had access to the cloud billing dashboard.
  • C. The cost overrun didn't actually start until week five.
  • D. Cost can't be measured until a pilot fully ends.
Show hint
Think about which number got a weekly check and which one got a guess at kickoff instead.
Show answer
A. The overrun was there from the start, it just wasn't being watched. A quiet week of quality notes never meant cost was fine too, it just meant nobody had looked.
Short answer
5. If Hollow Pine's overrun had been one-and-a-half times budget instead of six-point-four times, would the same position still hold? Walk through it.
Show hint
Think about whether the position depends on the exact multiple, or on having a threshold at all.
Show answer
No, not the same position. The threshold in this pick puts "push forward and fix in parallel" under roughly three times budget. One-and-a-half times sits comfortably under that line, closer to a tuning gap than an architecture problem, so the call would be push forward and fix cost as a fast follow, no pause. The shape of PICK doesn't change, name the threshold, name who feels each side, but the actual position flips because the magnitude is different. That's exactly why the position names a number instead of a rule.
Short answer, apply it yourself
6. Pick a product you use yourself. Name one part where the business would rather quietly absorb a cost than pause and let a customer notice, and one part where the reverse is true.
Show hint
Look for the part where a delay is annoying but visible, versus the part where a hidden ongoing cost could quietly pile up before anyone checks it.
Show answer
Model answer: "A photo-editing app. If a new background-removal feature runs a little slow, that's cheap and visible, users notice a spinner and shrug. If the same feature's cloud processing cost quietly runs three times over budget per photo, the company should pause wide rollout and fix the cost path, because a slow feature costs seconds; an unpriced one costs real money on every photo edited, hidden until the bill arrives." Any answer works if it names the part where a miss stays hidden until it's already expensive.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more