CaseAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #8
Your spike produces 70 percent accuracy. How do you decide whether that is promising or fatal?
PICKone number, two very different bets
Haulbridge is a marketplace where people rent tools, tents, and camera gear from each other. ClaimLens is the photo-based tool that scores a returned item as clean or damaged at the drop-off counter. Bertrand Okwuosa is the AI PM who had to decide what a 70 percent spike actually meant.
The direct answer
Seventy percent is promising, not fatal, but only once you stop reading it as one number. Split it by what a wrong call actually costs: on a forty-dollar tent deposit, ClaimLens is right 91 percent of the time, and a miss just means a clerk glances twice. On a six-hundred-dollar camera lens, it's right barely half the time, and a miss means Haulbridge eats the cost of real damage nobody caught. Ship it as the first read on cheap gear now. Do not let the same 70 percent auto-decide anything with real money behind it.
Do this, in order
Split the accuracy number by what a miss actually costs, before judging it as one figure.Why: a blended 70 percent can hide a 91 percent case sitting right next to a 52 percent one.
Name who absorbs each kind of error, in dollars, not just a percent.Why: a cheap, visible mistake and a hidden, expensive one are not the same failure, even at the same error rate.
Ship the model where the wrong call is cheap, and hold it back where the wrong call is expensive.Why: this turns one ambiguous number into two clear decisions instead of one shaky compromise.
Set a kill line, a specific accuracy floor per category, not a gut feeling.Why: a stated floor is something you can be held to, and something you can act on later without relitigating the whole decision.
Never let an operational quota push staff to trust the model equally across categories.Why: a fixed cap on manual review time will get spent wherever feels urgent, not where the model actually needs it.
Say plainly when the blended number really is good enough, like a catalog with mostly low-value gear.Why: shows judgment about where a single accuracy figure is fine, instead of splitting everything by habit.
How to answer this, stage by stage
Nobody is scoring whether you can call 70 percent good or bad. They're scoring whether you can turn one ambiguous number into two different, defensible decisions.
Stage 1
Scope it to one real spike
Say it like this
"Let's ground this in ClaimLens at Haulbridge. That's the spike where 70 percent turned out to mean two very different things depending on what was being rented."
Why this works
Keeps the answer from becoming an abstract opinion about whether 70 percent is a good score.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as PICK. Position, my call, before the reasoning. Impact, who feels each kind of error. Cost asymmetry, which error is the hidden, expensive one. Kill criteria, what would change my mind."
Why this works
Shows a repeatable way to judge a spike number, not a one-off gut call.
Stage 3
Reframe: it isn't "is 70 percent good," it's "70 percent of what, and wrong on whom"
Say it like this
"This isn't really a question about whether 70 percent clears some bar. It's a question of which claims are inside that 70, and what it costs somebody the day it's wrong."
Why this works
This is where a strong answer separates from someone who just calls 70 percent "decent" and moves on.
Stage 4
Give the position, committed, before the reasoning
Say it like this
"My call: promising for tents, coolers, and kids' bikes. Fatal, as a blanket policy, for power tools and camera gear. I'm not averaging those into one verdict."
Why this works
Commits to an answer instead of hedging, which is exactly what PICK is testing.
Stage 5
Prove it with the segmented evidence
Say it like this
"On our own spike data, ClaimLens hits 91 percent on low-value gear, 68 percent on power tools, and 52 percent on camera gear. That 70 percent headline is just those three numbers blended together, weighted by how many tents we rent compared to lenses."
Why this works
Compresses the whole argument into the one split that the headline number was hiding.
Stage 6
Name the kill criteria, the AI-specific reasoning, and close
Say it like this
"My kill line is 60 percent per category, not overall. Below that, no auto-decide, full stop. The reason this isn't a normal quality bar is that a wrong 'clean' call here doesn't throw an error, it just quietly releases someone's deposit on a broken lens. I'm accepting slower review on expensive gear in exchange for never eating that cost silently again."
Why this works
States a real, checkable bar and names the trade-off, speed against the cost of a confidently wrong release.
Let's learn
What happens the first time a single accuracy number gets used to decide something it was never built to decide?
Before ClaimLens, a clerk at Haulbridge's drop-off counter checked every returned item by hand against the intake photo, about three minutes an item, roughly 200 returns a day across four hubs. With ClaimLens, a photo gets scored in about two seconds, clean or damaged, and clerks were told to trust a clean call outright to keep the counter line moving.
The four letters, held up as one page. Cost asymmetry is the one a blended number skips past.
Here's the turn: ClaimLens's spike hit 70 percent accuracy overall. That number blended three very different results: 91 percent on cheap, common gear, 68 percent on power tools, and 52 percent on camera gear, the smallest share of returns but the largest deposits. The 70 percent felt borderline-promising. It was actually two decisions wearing one number.
Both are "ClaimLens got it wrong." Only one of them costs six hundred dollars.
At its worst, a clerk releases a deposit on a camera lens with a cracked mount because ClaimLens called it clean, and the damage only surfaces when the next renter opens the case, weeks after any photo evidence window has closed.
Seventy percent was never one number. It was three numbers standing in a line, hoping nobody asked which was which.
The choice I would take back
Haulbridge onboarded every clerk with one line: "ClaimLens is right about 7 times out of 10, use it as your first read." That made sense when the catalog was mostly tents and coolers, where a miss barely mattered. It stopped making sense once camera gear and power tools joined the catalog, at three to ten times the deposit value, with no separate guidance at all.
What I would leave alone: for a catalog that stayed mostly low-value gear, tents, coolers, folding chairs, the blended 70 percent would already be a fine number to act on without splitting it further.
The lesson: a single accuracy number is a average wearing a verdict's clothes. The real decision lives one level down, in which claims are inside it.
Now here is the same thing as a story
The short version above is what you'd say defending the rollout plan in a Monday review. Read this one for what it felt like the week a colleague's remark turned a comfortable number into an uncomfortable one.
The drop-off hub gets busy right around five, when everyone returning gear on their way home hits the counter at once.
The third step is where a blended 70 percent quietly decided far more than anyone meant it to.
Bertrand launched ClaimLens with the 70 percent number and a rule: manual holds could not exceed 15 percent of a day's returns, since a slower counter line was its own real cost. Clerks stayed under quota fine at first, mostly by trusting ClaimLens across the board, tents and camera gear alike.
Camera gear sits exactly where a quota-driven clerk has the least room to double-check by hand.
A colleague reviewing weekly disputes said it plainly in a hallway, not unkindly: "You're still trusting that 70 percent thing on camera gear?" Bertrand hadn't split the number by category since the spike readout. He didn't have an answer ready.
Knowledge spark: why would one accuracy number hide such different results?
A blended accuracy number is a weighted average across everything the spike was tested on. If most of your returns are cheap, common items, the model's strength there can carry the overall score even while it quietly fails on the rare, expensive items sitting inside the same average.
Six weeks in, a renter disputed a released deposit on a six-hundred-dollar lens with a cracked mount, caught only when the next renter opened the case. The clerk had followed policy exactly. The policy was the problem.
Four things a single clean-or-damaged call has to get right, on a much harder item than a tent pole.
The real question was never whether ClaimLens was 70 percent accurate. It was whether that 70 percent had ever been asked to do the same job on a forty-dollar cooler and a six-hundred-dollar lens.
What a wrong call costs, by claim type
A false "damaged" flag costs a few dollars in staff time either way. A false "clean" call costs almost nothing on a cooler and 650 dollars on a lens.
When the 15 percent manual-hold quota was set, someone said, "let's keep the line moving, clerks will use their judgment," and it sounded reasonable, since at the time camera gear barely existed in the catalog.
Accuracy by deposit value, and where the kill line sits
Camera gear crosses under the kill line well before its own deposit value peaks.
Rerun the same six weeks with claims split by category from day one: camera gear gets a mandatory manual hold regardless of quota, the lens with the cracked mount gets caught at the counter, and the renter dispute never happens.
What I'd tell myself, hearing "you're still trusting that on camera gear": 70 percent was never the number I should have been defending. The number that mattered was 52, and it was sitting inside my own headline the whole time.
PICK, in one screenNot a script for making every spike number look fine. PICK is what tells you exactly which slice of it isn't.
P
Position. The pick, before any reasoning.
Promising for low-value gear, fatal as a blanket policy for camera gear and power tools. Two decisions, not one average.
Committing to a split call, instead of one verdict on 70 percent, is the whole point of this letter.
I
Impact. Who feels each kind of error.
A false damage flag costs a clerk a few extra minutes. A false clean call costs Haulbridge the deposit and the renter a dispute nobody can resolve weeks later.
Naming both sides in real units, minutes against dollars, is what makes the asymmetry visible.
C
Cost asymmetry. Which error is hidden and expensive.
A false "clean" on camera gear is the hidden one, invisible until a renter disputes it weeks later, while a false "damaged" flag is cheap and gets caught at the counter immediately.
This is the hardest step, and the one a single blended accuracy number was built to hide.
K
Kill criteria. What would flip the pick.
Below 60 percent accuracy in any category, no auto-decide, full stop, regardless of the manual-hold quota.
A stated floor is a decision you can be held to later, not a feeling you'll relitigate every time a dispute lands.
The recap, one line per letter: position is promising for cheap gear, fatal as policy for expensive gear, impact is minutes against dollars, cost asymmetry is the hidden clean-call miss on camera gear, and kill criteria is the 60 percent floor per category that overrides any staffing quota.
And if you want to be sure it really works, try it somewhere elseSame four letters, a shipping terminal instead of a rental counter. The blended number changes shape, the split still saves it.
Henrike Solmundsdottir runs product at Port of Kallabrekka Terminal Ops, where CrateLens scores photos of unloaded shipping crates for visible damage before a claim gets filed against the carrier. A spike hit 74 percent overall, comfortable on paper. Mapped onto PICK: position is promising for standard dry-goods crates, fatal as a blanket call for refrigerated and hazardous cargo. Impact means a missed dry-goods dent costs a shrug and a reshelving, while a missed refrigeration seal breach costs a spoiled load worth tens of thousands. Cost asymmetry is the same shape as Haulbridge's, a rare, expensive category hiding inside a comfortable average. Kill criteria sets a hard floor for cold-chain cargo specifically, regardless of the terminal's own dock-speed targets.
A different flip than Haulbridge's: once cold-chain misses got expensive, a senior inspector reclaimed every call a junior clerk used to make alone.
Week 3 is where a blended number quietly became a decision about camera gear nobody had actually tested for.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "split the number by what a miss costs, ship where it's cheap, hold where it's not," and stop.
Cost: there's no time to build separate models per category before launch. Say so honestly, and set a manual-hold rule by category using the one blended number you have, rather than trusting it equally everywhere.
The spike genuinely improves, for real: if a retrain pushes camera gear from 52 to 75 percent, that's still worth testing against the same 60 percent kill line before loosening the hold rule, not assumed safe because the headline number went up.
Where people run it wrong.
They judge one blended accuracy number as if it applies equally everywhere it was measured.
They let an operational quota, like a manual-review cap, quietly decide which categories get real scrutiny.
They set a kill line after a dispute happens instead of before, so it reads as blame instead of a plan.
How to use it live. The moment someone hands you one accuracy number, ask yourself: what's the most expensive thing hiding inside that average, and what does it cost the day it's wrong. Split the decision there, and the rest of the answer follows.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Substitution flip: once a fixed daily quota on manual holds made checking scarce, clerks trusted ClaimLens roughly equally across all claim types, including camera gear, exactly the case it handled worst.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bertrand Okwuosa, the AI PM at Haulbridge, who launched ClaimLens off a blended 70 percent spike result.
3 · THE HABIT
What did clerks stop doing once the 15 percent manual-hold quota existed?
Tap to flip
ANSWER
They stopped manually double-checking camera gear and power tools at a higher rate than cheap gear, since checking every expensive item would have blown the daily quota.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Reserving manual scrutiny for the cases that actually needed it, versus trusting ClaimLens at roughly the same rate everywhere to stay under the quota. No middle setting once the quota existed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Onboarding every clerk with one blanket line, "ClaimLens is right about 7 times out of 10, use it as your first read," with no calibration by claim type.
6 · THE NUMBER
Fill in the blank: the blended 70 percent broke down into ___ percent on low-value gear, 68 percent on power tools, and ___ percent on camera gear.
Tap to flip
ANSWER
91 percent, and 52 percent.
7 · THE REPLAY
Same six weeks, claims split by category with a mandatory hold on camera gear. What changes?
Tap to flip
ANSWER
The lens with the cracked mount gets caught at the counter by a manual hold instead of released clean, and the renter dispute that took six weeks to surface never happens.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Port of Kallabrekka Terminal Ops' CrateLens. The flip is delegation: a senior inspector reclaimed every damage call a junior clerk used to make alone, once cold-chain misses got expensive.
Check yourself Score: 0 / 0
True or false
1. True or false: a 70 percent blended accuracy number means ClaimLens is equally reliable on every kind of claim.
True
False
Show hint
Look at "here's the turn" in the first section.
Show answer
False. The 70 percent blended 91 percent on cheap gear with 52 percent on camera gear, two very different levels of reliability hiding in one average.
Multiple choice
2. Why is a false "clean" call on camera gear more dangerous than a false "damaged" flag on a tent?
A. Tents are made of cheaper material, so ClaimLens scans them faster.
B. A false clean call is invisible until a renter disputes it weeks later, while a false damage flag on a tent just costs a clerk a quick re-check.
C. Camera gear photos are always blurrier than tent photos.
D. Clerks are less experienced with camera gear.
Show hint
Look at the grouped bar chart, "what a wrong call costs, by claim type."
Show answer
B. The cost asymmetry is about where the error hides and how much it costs when it surfaces, not about image quality or experience.
Fill in the blank
3. Fill in the blank: the kill line in this answer is set at ___ percent accuracy, per category, not as one overall average.
Show hint
Look at stage 6 of the walkthrough, or the K step in the PICK recap.
Show answer
60 percent. Below that floor in any single category, ClaimLens does not get to auto-decide, regardless of how the overall blended number looks.
Short answer, where it wouldn't matter
4. Name a version of Haulbridge's catalog where the blended 70 percent would already be a fine number to act on, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A catalog that stayed mostly tents, coolers, and folding chairs. With no expensive category hiding inside the average, the blended number and the real number would be close to the same thing.
Short answer, apply it yourself
5. Think of a rating or score you trust for a product or service. What expensive, rare case might be hiding inside that one number?
Show hint
Think about whether the rating is a single blended average across very different situations.
Show answer
Model answer: A restaurant's overall rating might be carried by strong scores on common weeknight orders while hiding a much worse record on rare, complicated catering orders that almost nobody leaves a review for.
Short answer, work the number
6. If camera gear made up 40 percent of returns instead of a small share, would the same 70 percent headline number still be believable?
Show hint
Think about how a blended average shifts as the mix of categories changes.
Show answer
Model answer: No. With camera gear at 52 percent making up 40 percent of returns instead of a small slice, the blended number would drop well below 70, closer to the low 60s, and the problem would have been obvious much sooner.
Before you close the answer
Why this works
Tests whether you'll accept a single accuracy number at face value, or ask which categories are hiding inside it and what a miss costs each one.
Follow-up traps
"Isn't splitting the number just cherry-picking a worse result?" Response: no, the split is by real cost, not by whichever slice looks bad, and it's the same split that would show a category doing better than the average too.
"What if you don't have enough camera-gear returns yet to trust a 52 percent figure?" Response: then that's the honest answer, the sample is too thin to trust any number on that category, which is itself a reason to hold it back from auto-deciding, not a reason to default to the blended one.
If pressed
The 60 percent kill line was set using the same margin-of-error logic as a sample-size check: below that floor, a category's own sample of past claims was too small for the accuracy figure to be trusted as more than noise.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.