CaseIntermediateAI Opportunity & Model Strategy / Model selection from a PM lens / #3
Describe how you would build a model bake-off for your specific use case.
SPARKthe vendor's demo answered the easy permit, Halvergate's counter gets the hard one
Here's what happens when a model gets chosen off a fifteen-minute demo instead of a real test. Halvergate County Permits Office built CodeReader, a tool that reads a submitted building permit application and flags which zoning and fire-code sections a reviewer should check first. Rosalind Achterberg runs the program, and inherited the job of picking CodeReader's model from a predecessor who left two weeks after the demo happened.
The direct answer
Build the bake-off around your own messiest real cases, not a vendor's clean demo set. Collect 150 to 300 of your own actual permit applications, including the confusing ones with hand-written additions and mismatched addresses. Write a short rubric with a real person who does this job today. Run every candidate model blind against the same set, score them against the rubric, and only then look at cost and speed. The bake-off's whole job is to make the messy case unavoidable before you sign anything.
Do this, in order
Build the test set from your own messiest real cases, not a vendor's demo examples.Why: a demo is built to look clean, and clean is exactly what your real backlog isn't.
Write the rubric with someone who actually does this job today.Why: only a real reviewer knows which mistake is dangerous and which one is cosmetic.
Run every candidate blind, scored the same way, before anyone sees a live demo again.Why: a live demo lets the vendor pick which case you see. A blind run doesn't.
Score against the rubric first, then bring in cost and speed.Why: a cheap model that fails your rubric isn't a bargain, it's a liability with a lower price tag.
Keep a visible reason on every flag, so a wrong one costs a minute, not a re-review.Why: it turns a miss into a fast fix instead of a reason to distrust the whole tool.
Say plainly what the bake-off won't decide yet, like rare permit types with too few real examples.Why: shows judgment instead of pretending one test set covers everything on day one.
How to answer this, stage by stage
Nobody is scoring whether you know the word "bake-off." They're scoring whether your test set would actually catch something a fifteen-minute vendor demo wouldn't.
Stage 1
Scope it to one real tool
Say it like this
"Let's ground this in a real case. Halvergate County's CodeReader reads a submitted building permit and flags which zoning and fire-code sections need a closer look. I'd build the bake-off around exactly that task, not a generic model comparison."
Why this works
Keeps the answer from turning into a generic list of bake-off best practices with nothing real underneath it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Situation, how the review gets done today without any tool. Payoff, the habit I actually want the bake-off to build. Anchor, the one concrete design decision behind the test itself. Risk, what breaks the first time a model's wrong. Keep out, what I won't try to settle on day one."
Why this works
Signals a repeatable design method for building the test, not just a gut feeling about which model looked good.
Stage 3
Reframe: it isn't "which model wins," it's "which test would catch a real miss"
Say it like this
"This isn't really about picking a winner from three demos. It's about building a test hard enough that a model can't pass it by only handling the easy, clean cases a vendor would choose to show you."
Why this works
This is where a strong answer separates from a candidate who just says "I'd run a bake-off" and stops there.
Stage 4
Give the one decision: the anchor
Say it like this
"Here's the anchor: pull 200 of Halvergate's own past permit applications, weighted toward the messy ones, hand-written additions, mismatched addresses, unusual zoning overlays. Build a rubric with a real plans examiner. Run every candidate model blind against the same 200, score against the rubric, then compare cost and speed only among the ones that clear the bar."
Why this works
This is the direct answer, stated as an actual test you could run next week, not a philosophy about evaluation.
Stage 5
Prove it with the compressed evidence
Say it like this
"Halvergate's original plan skipped a real bake-off entirely. The vendor's live demo ran five clean, single-family permit applications, and the model flagged all five correctly. Nobody ran it against a real mixed-use conversion with a hand-marked setback change until three months after go-live, when a reviewer noticed the model had cleared 40 percent of that permit type without a single flag."
Why this works
Compresses the whole case into the one gap between the demo's five cases and Halvergate's real 200.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't just a generic vendor evaluation is that a model doesn't fail loudly on a case it's never really been tested on, it answers fluently and confidently, and a five-case demo can't tell you that. We accepted three extra weeks before go-live, and about six thousand dollars in reviewer time, to build a rubric and a real test set instead of trusting a vendor's own chosen examples."
Why this works
This is the load-bearing, AI-specific judgment. A normal software demo doesn't quietly hide a whole category of case it was never actually tested against.
Stage 7
Say what you'd keep out
Say it like this
"I wouldn't try to settle rare permit types, like a historic-district variance, in this first bake-off, since Halvergate only has six of those on file, nowhere near enough to trust a score either way. I'd flag those for manual review and revisit once there's a real sample."
Why this works
Shows judgment instead of pretending one 200-case test set can settle everything on day one.
Stage 8
Close on the one line
Say it like this
"So the bake-off is built on Halvergate's own messiest 200 cases, scored against a real reviewer's rubric, blind, before cost or speed even enter the conversation. That's what would have caught the mixed-use gap before it shipped instead of three months after."
Why this works
Restates the direct answer in one breath, closing on the exact thing the question asked for.
Let's learn
Say we build a tool that reads a permit application and tells a reviewer which sections to check first.
Before CodeReader, a Halvergate plans examiner read every submitted permit cover to cover, checking it against zoning maps and fire code by hand, about 35 minutes for a straightforward single-family permit and well over an hour for anything with a variance or an addition.
With CodeReader, a submitted permit gets flagged in under 20 seconds, telling the reviewer which sections likely need a closer look.
Four manual tasks, all riding on one examiner's memory of the whole code book.
Here's the turn: the real risk was never that CodeReader would be obviously wrong. It's that a model can look perfect on five clean demo cases and still have never actually been tested on the messy ones that make up a real counter's daily backlog.
The demo answered the easy permit. Halvergate's real counter gets the hard one, every single day.
The actual design decision behind the bake-off, drawn as five steps you could run this week.
At its worst, a permit office trusts a model that quietly clears a whole category of case, mixed-use conversions with hand-marked changes, without a single flag, and nobody notices until a building goes up with the wrong setback.
Three finalists, scored against Halvergate's own rubric on 200 real permits
The demo pick's real score only surfaced once the test set stopped being five clean cases and became Halvergate's own messy 200.
The choice I would take back
Halvergate's original rollout plan merged "watch the vendor demo" and "sign the contract" into one meeting, with no separate testing step in between. That made sense when the office was moving fast to clear a permit backlog and the demo looked genuinely clean. It stopped making sense the moment a real mixed-use permit exposed a gap nobody had tested for.
What I would leave alone: for straightforward single-family permits with no variances, which make up most of Halvergate's volume, CodeReader's flags have been reliable from day one, and I wouldn't add extra review steps there.
The lesson: a bake-off isn't a formality before you sign. It's the only thing standing between a model that looks perfect and a model that's only ever seen five perfect cases.
Now here is the same thing as a story
The short version above is what you'd say in a vendor review. Read this one for what it felt like the months nobody had drawn a clear line between watching a demo and trusting a model.
Rosalind Achterberg had reviewed permits herself for nine years before moving into the program manager seat, and she still knew a dangerous mistake from a cosmetic one on sight.
When CodeReader's vendor demo ran, it looked clean: five straightforward single-family permits, flagged correctly in under 20 seconds each. For the first three months after go-live, feedback from the front counter was good, examiners liked how much reading it saved them.
No single moment marked when it started going wrong. It built up slowly, permit by permit, and nobody could later say exactly which week it changed. Examiners kept trusting the flags for permit types nobody had actually tested the model against, because nothing about it looked different from the types that were working fine.
A demo is built to look effortless. Your real counter was never going to be that clean.
Three months in, a senior examiner reviewing a mixed-use conversion, an old house being split into two rental units, noticed CodeReader had cleared it with no flags, despite a hand-marked setback change on page four that clearly needed a zoning check.
Knowledge spark: why would a model miss something a five-case demo never caught?
A model trained mostly on clean, well-formatted examples can answer fluently on a messy one too, it just won't necessarily answer correctly. It doesn't know it's out of its depth. It doesn't refuse or hesitate. It just produces a confident, wrong "no flags needed," which looks identical to a genuinely clean permit until someone checks by hand.
Rosalind pulled every mixed-use permit CodeReader had processed since launch, 40 of them, and had a senior examiner manually re-check each one. Sixteen, 40 percent, had at least one real issue CodeReader missed entirely.
We didn't lose 16 flags. We lost the one thing a permit office actually sells the public: that someone checked.
The real question was never whether CodeReader's model was good. It was whether anyone had ever tested it against a permit that looked like the ones actually piling up at Halvergate's counter, instead of the five the vendor chose to show.
Same mistake. One version costs three months. The other costs one afternoon of testing.
When the original rollout plan was drawn up, someone said, "the demo looked clean, let's move to contract," and it sounded reasonable, since nobody yet knew mixed-use conversions would be where the gap was hiding.
Cost per correctly flagged permit, three finalists, at Halvergate's real volume
The cheapest option wasn't the best deal. It was the one that missed 68 real issues out of 200.
Rerun the same 40 mixed-use permits with a real bake-off run first: the 16 missed issues surface in the first afternoon of testing, on a 200-case set built from Halvergate's own backlog, three months before any of them would have shipped silently.
What I'd tell myself, hearing that senior examiner describe the missed setback: the demo was never lying. It was just never asked the question that actually mattered.
SPARK, applied to designing the bake-off itselfNot a script for distrusting every vendor. SPARK is what makes sure the test, not the demo, is what earns your trust.
S
Situation. How does review happen today, without a model?
A plans examiner reads every permit cover to cover, checking zoning maps and fire code by hand, 35 minutes to over an hour depending on complexity.
Grounding in the real task keeps the bake-off from becoming an abstract accuracy exercise.
P
Payoff. What habit do you want the bake-off to build?
Trust a shortlist of flagged sections enough to act on it in under a minute, on any permit type, not just the clean ones a vendor happened to demo.
The habit, not the twenty seconds saved, is what a real bake-off has to protect.
A
Anchor. The one design decision everything hangs on.
200 of Halvergate's own real permits, weighted toward messy cases, scored blind against a rubric built with a real examiner, before cost or speed enter the picture.
This is the hardest step, and the answer to the question, stated as a real test you could run, not a philosophy.
R
Risk. What breaks the first time a model's wrong?
A missed flag on a real permit ships silently, since the model never hesitates or signals uncertainty. It just clears the case.
Designing the bake-off to hunt for exactly this kind of silent miss is the whole point.
K
Keep out. What you won't settle on day one.
Rare permit types, like historic-district variances, with only six examples on file, nowhere near enough to trust a score. Flag those for manual review and revisit once a real sample exists.
The recap, one line per letter: situation is the manual read, payoff is trusting the shortlist on any permit type, anchor is the 200-case blind test, risk is the silent miss, and keep out is admitting six examples isn't enough to score anything yet.
And if you want to be sure it really works, try it somewhere elseSame five letters, a recycling facility instead of a permits counter. What "messy" means changes, the design still turns on it.
Anders Kettering runs operations for Kettering Materials Recovery, which uses SortCheck, a model that reads a conveyor camera feed and flags contaminated bales before they're baled and shipped. Mapped onto SPARK: situation is a sorter manually pulling suspect items off the line by eye, missing about a fifth of contamination in fast-moving mixed loads. Payoff is trusting an automatic flag enough to pull an item without double-checking it first. Anchor is the bake-off itself: 300 of the facility's own recorded conveyor clips, weighted toward its messiest mixed-material loads, scored blind by a veteran sorter's rubric before cost per camera or processing speed even enter the conversation. Risk is a contaminated bale clearing the line silently and getting rejected by a buyer weeks later, with no record of which clip caused it. Keep out is settling scores on hazardous-material contamination, since the facility only has four recorded incidents, far too few to trust any model's number on that category yet.
Plotted together, the cheapest option and the best deal turn out to be two different points.
Naming what the bake-off won't decide yet is as much a design choice as naming what it will.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "build the test from your own messiest real cases, score blind against a rubric written with a real reviewer, then compare cost," and stop.
Cost: no budget for 200 real cases before a decision has to be made. Say so honestly, and commit to even 40 of the messiest cases you can pull cheaply before trusting a live demo alone.
The model got better, for real: if a vendor claims an updated model fixes the exact gap you found, that's still worth re-running against your own messiest cases before trusting it, since "fixed" and "fixed on your specific gap" aren't the same claim.
Where people run it wrong.
They let the vendor pick which cases the demo shows, instead of bringing their own.
They build the rubric alone in a meeting room instead of with someone who actually does the job.
They compare cost and speed before any candidate has cleared a real quality bar.
How to use it live. The moment an interviewer asks you to design a bake-off, ask yourself: what's the messiest real example this team has, and is it actually in the test set? If the honest answer is "no, we'd probably use the vendor's demo cases," that's the gap worth naming out loud.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family sits underneath this story?
Tap to flip
ANSWER
Substitution flip: examiners kept trusting the easy, well-covered permit types, while the model quietly handled the hardest, least-tested cases with no flag at all.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Rosalind Achterberg, program manager at Halvergate County Permits Office, a former examiner who inherited the CodeReader rollout from a predecessor who'd already left.
3 · THE HABIT
What did examiners stop doing because CodeReader seemed to work?
Tap to flip
ANSWER
Manually cross-checking a permit's flags against the code book themselves, even for permit types the model had never actually been tested on.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Trusting CodeReader's flags on the permit types it was actually shown in the demo, versus trusting it identically on types it had never once been tested against.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging "watch the vendor demo" and "sign the contract" into one meeting, with no separate real-case testing step in between.
6 · THE NUMBER
Fill in the blank: of 40 mixed-use permits CodeReader cleared since launch, ___ percent had a real issue it missed entirely.
Tap to flip
ANSWER
40 percent (16 of 40), on a permit type nobody had tested the model against before go-live.
7 · THE REPLAY
Same 40 mixed-use permits, bake-off run first. What changes?
Tap to flip
ANSWER
All 16 missed issues surface in the first afternoon of testing, on Halvergate's own 200-case set, three months before any of them would have shipped silently.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Kettering Materials Recovery's SortCheck, an over-trust flip: sorters stop double-checking a flagged item the moment the tool starts looking reliable on easy loads.
Check yourself Score: 0 / 0
True or false
1. True or false: since CodeReader's vendor demo flagged all five sample permits correctly, that was enough evidence to sign the contract.
True
False
Show hint
Think about who chose those five cases, and what they left out.
Show answer
False. The five demo cases were all straightforward single-family permits. Mixed-use conversions, a real category in Halvergate's own backlog, were never tested at all before go-live.
Multiple choice
2. Why should cost and speed be compared only after a rubric bar is cleared, not before?
A. Cost and speed never actually matter for a permits office.
B. A cheap, fast model that fails the rubric isn't a bargain, since a missed flag costs far more than it saves.
C. Vendors always lie about their own pricing.
D. Rubric scores and cost are always identical anyway.
Show hint
Look at the scatter chart comparing cost per permit to correct flags.
Show answer
B. The cheapest finalist also missed the most real issues, 68 out of 200, which costs far more in re-review and risk than the savings on price.
Fill in the blank
3. Fill in the blank: the chosen candidate scored ___ correct flags out of 200 real permits, at a cost of $___ per permit.
Show hint
Look at the scatter chart in "the same thing as a story."
Show answer
181; $2.35. Not the cheapest option and not the most expensive, but the one that actually cleared the rubric bar by the widest margin.
Short answer, where it wouldn't matter
4. Name a part of Halvergate's own permit volume where CodeReader's original flags stayed reliable, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Straightforward single-family permits with no variances, which make up most of Halvergate's volume and were exactly the kind of case the original demo actually tested.
Short answer, apply it yourself
5. Think of a tool at your own job or school that was chosen after a demo. What's the messiest real example you could have tested it on, that the demo probably never showed?
Show hint
Think about the edge case that would actually embarrass the tool, not the case that flatters it.
Show answer
Model answer: A scheduling tool demoed on a simple two-person calendar. The messy real case is a shared calendar with recurring conflicts, time zones, and last-minute cancellations, exactly what a real team's week actually looks like.
Short answer, work the number
6. If Halvergate could only afford to test 40 real permits instead of 200 before deciding, would the bake-off still be worth running?
Show hint
Compare 40 real, messy cases to zero real cases and a five-case vendor demo.
Show answer
Model answer: Yes. Even 40 of Halvergate's own messy cases, weighted toward mixed-use permits, would likely have surfaced the 40 percent miss rate the later full audit found, well before three months of silent misses.
Before you close the answer
Why this works
Tests whether you'd build a test hard enough to actually fail a model, instead of accepting a vendor's own chosen examples as sufficient proof.
Follow-up traps
"Isn't 200 cases still a small sample for a whole permit office?" Response: it's a floor, not a ceiling, weighted toward messy cases specifically because that's where a demo-only evaluation would never look. It's also 40 times bigger than the vendor's own five-case demo.
"What if the rubric itself has blind spots the reviewer didn't think of?" Response: that's exactly why the rubric gets built with someone doing the job today, and why the bake-off gets revisited each time a new gap like the mixed-use miss actually surfaces in production.
If pressed
The 200-case set was drawn with deliberate oversampling: mixed-use and variance permits made up only 12 percent of Halvergate's real volume but were weighted to 35 percent of the test set, specifically because rare-but-risky cases need more test coverage than their raw frequency alone would give them.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.