CaseIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #2

Design a two-week spike to test whether an AI feature is viable.

ORDERone estimate off by a hundred and twenty four thousand dollars

Casterbridge Realty Group is building RenoScope, a tool that reads a listing's photos and estimates renovation cost, so property investors can triage which listings are worth a bid before ever flying out to see one. Bartholomew Kase is the AI PM who has two weeks to prove it works before anyone commits real budget to it.

The direct answer
Design the two weeks around the order you test things in, not a list of things to test. Assemble a real golden set first, since nothing else can happen without it. Test raw accuracy against it privately, since a wrong number here just gets rerun. Only after that holds up do you run a narrow, blind trial with a few real investors, since a bad number shown live is the one mistake that doesn't come back. Never touch the live dashboard or the data feed until the first three are done. The two weeks are not fourteen days of building. They're four decisions, taken in the order that keeps the expensive mistake from happening first.
Do this, in order
  1. Sequence the spike by which mistake is hardest to undo, not by what's easiest to build first.Why: a wrong number shown to a real investor costs trust you can't get back; a wrong number in a private test just gets rerun.
  2. Assemble a real golden set of photos matched to actual costs before testing anything.Why: every later test depends on this existing first; it's the one forced step.
  3. Spend a single day on the cheapest possible signal before committing the full two weeks.Why: a same-day pass with an off-the-shelf tool tells you if the idea is even in the right neighborhood.
  4. Deliberately test the property type most likely to break the model, not the easiest one.Why: an average error rate hides exactly the segment that would sink the launch.
  5. Keep the live integration and dashboard out of the two weeks entirely.Why: building the wrapper before the number is proven is how a spike quietly becomes the real launch.

How to answer this, stage by stage

Nobody is scoring whether you can fill fourteen days with tasks. They're scoring whether the order you chose protects against the one mistake you can't take back.

Stage 1
Scope it to one real decision
Say it like this
"Let's ground this in RenoScope at Casterbridge, where a two-week spike had to decide whether cost estimates from listing photos were safe to show a real investor."
Why this works
Keeps the plan concrete instead of a generic checklist that would fit any AI feature.
Stage 2
State the order of attack
Say it like this
"I'll run this as ORDER. Outcome, what we're actually trying to protect. Reversibility, which mistake can't be undone. Dependency, what has to happen first. Evidence, what's cheap to learn early. Rank, the actual sequence."
Why this works
Signals a repeatable way to sequence any spike, not a one-off list built from memory.
Stage 3
Reframe: it isn't "what do we test," it's "what order keeps the worst mistake from happening first"
Say it like this
"The question isn't a list of things RenoScope should be tested on. It's which test, if skipped or done last, lets a genuinely bad number reach a real investor before anyone caught it."
Why this works
This is where a strong answer separates from a plan that's just a to-do list with a two-week label on it.
Stage 4
Give the ranked plan
Say it like this
"Days one through three, assemble 45 real photo-to-invoice pairs. Days four through seven, test raw accuracy privately. Days eight through ten, test it specifically on the property type most likely to hide problems. Days eleven through fourteen, one narrow, blind trial with real investors, and nothing else gets built."
Why this works
This is the direct answer, concrete enough that someone could actually run it starting Monday.
Stage 5
Prove it with the miss that almost shipped
Say it like this
"An earlier, unscoped attempt at this quoted 18,000 dollars in cosmetic repairs on a property whose real contractor invoice came to 142,000 dollars, because a foundation issue never showed up in the listing photos. That's the exact miss the ranked plan is built to catch before day fourteen instead of after a live launch."
Why this works
Turns the abstract "reversibility" idea into one number nobody in the room can wave away.
Stage 6
Say what you would track past day fourteen
Say it like this
"After launch, I'd watch accuracy by property age monthly, not just overall, since that's exactly the split the spike showed actually matters."
Why this works
Shows the spike's finding turns into an ongoing check, not a one-time fact that gets forgotten.
Stage 7
Say what you would not gate on
Say it like this
"I wouldn't spend spike time on the dashboard's look and feel. A number on a spreadsheet is enough to test whether investors trust the estimate at all."
Why this works
Shows judgment about where two weeks of scarce time is worth spending, and where it isn't.
Stage 8
Close on the one line
Say it like this
"A two-week spike isn't fourteen days of building. It's four decisions, made in the order that keeps the mistake you can't undo from happening first."
Why this works
Restates the direct answer in one breath, closing the loop on the whole plan.

Let's learn

Say we build a tool that reads a real estate listing's photos and estimates how much renovation the property will need, so an investor can decide which listings are worth pursuing before ever visiting one.

Before RenoScope, an acquisitions analyst spent about 20 minutes per listing eyeballing renovation scope from photos, across roughly 50 listings a week, nearly 17 hours weekly just triaging which properties were even worth a closer look.

Hand sketched labeled parts diagram titled What's in the golden set. A document icon at the center labeled Golden Set, with four labeled callouts around it: Real listing photos, Actual contractor invoice, Property age, Sale date.
Four things, pulled from real closed deals, before any accuracy number means anything at all.

Here's the turn: the team's first attempt at proving RenoScope out wasn't badly built, it was badly ordered. Everyone was ready to spend the whole two weeks wiring a rough model into the live listing feed and building a clean dashboard, before anyone had checked whether its cost estimate was even in the right neighborhood on a real, messy set of past deals.

Estimates within 25 percent of the real contractor invoice, by property age
100% 50% 0 86% Built after 1980 22% Built before 1980
Same model, same 45-property golden set. Splitting by property age is what surfaces the real gap.

At its worst, RenoScope ships on the strength of a demo built only from newer homes, an investor bids on an older property based on a rosy estimate, and a hidden foundation or wiring issue turns a projected 20,000 dollar flip into a six-figure loss, the kind of miss that ends the whole program's credibility in one deal.

The first attempt wasn't wrong about the model. It was wrong about which mistake to test for first.
The choice I would take back Early planning treated "build the integration" and "test the accuracy" as two tasks on the same list, in whatever order was convenient for engineering. That made sense when the team was optimizing for a fast demo. It stopped making sense the moment a real investor could see a live number before anyone had checked it against a real invoice.

What I would leave alone: for newer, post-1980 properties, I wouldn't hold the launch hostage to a perfect number on the oldest homes. RenoScope can ship for the segment it's already proven on while the older-home gap gets fixed separately.

The lesson: a two-week spike doesn't fail because the tasks were wrong. It fails when the order lets an unrecoverable mistake happen before a recoverable one gets the chance to catch it.

Now here is the same thing as a story

The short version above is what you'd say proposing the redo to leadership. Read this one for what the audit that forced the redo actually felt like.

Bartholomew Kase had been at Casterbridge two years, long enough to know which investors trusted a number and which ones wanted to see the receipts.

The first RenoScope attempt moved fast: a rough model, ten sample estimates, and a dashboard mockup ready to show at the next investor call. Nobody had checked the ten estimates against anything real yet, but the demo looked sharp, and the room was excited.

Hand sketched comparison titled Which mistake is easy to take back. Left panel, a gauge icon labeled Private accuracy test, caption wrong here, just rerun it. Right panel, a person icon labeled Live investor demo, caption wrong here, trust doesn't come back.
Two ways to be wrong. Only one of them can be quietly fixed afterward.

Persephone Ilić, who runs operations audits at Casterbridge, pulled ten of the sample estimates at random and checked each one against the actual signed contractor invoice from the same property, a habit she'd built long before RenoScope existed.

Knowledge spark: why would a model miss a foundation issue entirely? A model reading listing photos can only judge what's visible in the frame. Structural problems, bad wiring, or plumbing damage often show up only during an in-person inspection, so a photo-based estimate can look confident and still be blind to the single costliest issue in the house.

One property stood out: RenoScope estimated 18,000 dollars in cosmetic repairs. The actual signed contractor invoice was 142,000 dollars, almost entirely from a foundation issue no listing photo had ever captured. Persephone brought the number to the next planning meeting, not as a complaint, just as a fact nobody had checked yet.

Hand sketched quadrant titled What earns the first week. Axes, how reversible if wrong versus how much it unblocks. Items placed: assemble the data high unblock and easy to redo, raw accuracy test similar, live dashboard build low unblock and hard to undo, investor trust test moderate on both.
The dashboard build sat exactly where it shouldn't have been first: hard to undo, and unblocking nothing.

Bartholomew scrapped the plan to spend the two weeks polishing the dashboard and instead ranked the actual sequence: assemble a real golden set first, since nothing else could be trusted without it.

Hand sketched flow diagram titled The order that gets it right, fifth step emphasized. Five steps left to right: Assemble golden set, Test raw accuracy, Check the worst property type, Narrow blind trust test, Only then build it.
The live build sits last on purpose. Everything before it exists to make sure it's worth doing at all.

The real question was never whether RenoScope could produce a number. It was whether anyone had ordered the two weeks so the costliest kind of wrong got caught before an investor ever saw it, not after.

Hand sketched icon list titled The ranked order, one page. Five rows: Assemble 45 real cost to photo pairs first. Test raw accuracy before anyone sees a number. Check the worst property type on purpose. Run one narrow blind trust test. Leave the live integration for after.
Five items, in the order that matters. The last one is deliberately last.

When the first plan was drawn up, someone said, "let's get the dashboard looking right, that's what'll sell it in the room," and it sounded reasonable, since a sharp demo really had gotten leadership's attention before.

Hand sketched timeline titled Two weeks, the actual plan, second milestone emphasized. Four milestones: Golden set frozen, day 3. Accuracy number in hand, day 7. Worst case tested, day 10. Kill, scope down, or go, day 14.
Fourteen days, four real checkpoints. No day left open for "and also polish the demo."

Rerun the same two weeks with the ranked plan already in place: the 142,000 dollar gap surfaces on day nine, against a private golden set, at a cost of about 14,000 dollars, instead of surfacing after a real investor had already bid low on a hidden six-figure problem.

What I'd tell myself, hearing Persephone's number land in that meeting: a sharp demo proves the model can produce a confident answer. It was never proof that the confident answer was the right one.

ORDER, the sequence that protects the mistake you can't undoNot a script for slowing every spike down. ORDER is what tells you which test has to come first, and which one can safely come last.

O
Outcome. What are all the candidate tests actually competing to prove?
Whether an investor can safely triage listings by renovation cost, without missing a good deal or chasing a bad one.
Without a stated outcome, any order of tasks looks equally reasonable.
R
Reversibility. Which mistake is hardest to undo?
A wrong number shown to a real investor costs trust that doesn't come back. A wrong number in a private test just gets rerun.
This is the hardest step, and the one that decides the whole order of the two weeks.
D
Dependency. What has to happen before anything else can?
A real golden set of 45 photo-to-invoice pairs has to exist before accuracy can be measured at all.
Some order isn't a judgment call, it's just what reality allows.
E
Evidence. What's cheap to learn before committing the full two weeks?
A same-day pass on ten sample photos with an off-the-shelf vision tool, before writing any custom code.
Cheap early evidence keeps the team from committing two full weeks to an idea that was never in the neighborhood.
R
Rank. State the order, and defend the top pick.
Golden set, then raw accuracy, then the worst property type on purpose, then a narrow blind trial, and only then, if ever, the live build.
The top pick is defended by what it protects against, not by gut feel about what's most exciting to build first.

The recap, one line per letter: outcome is safe listing triage, reversibility is a live bad number versus a private one, dependency is the golden set unlocking everything else, evidence is a same-day cheap pass before full commitment, and rank is data, accuracy, worst case, trial, then build, in that order.

And if you want to be sure it really works, try it somewhere elseSame five letters, a fishing dock instead of a listing photo. The costliest mistake still isn't the one everyone expects.

Sigrun Haugland runs product at Nordveil Fisheries Co-op, testing CatchWeigh, a tool meant to estimate a haul's total weight from a deck photo, to speed up dockside sorting before the truck scale confirms it. Mapped onto ORDER: outcome is pricing a haul fairly and fast without slowing offloading. Reversibility is a wrong private test number, easy to rerun, versus a wrong price quoted to a fisherman at the dock, which is hard to undo once trust in the tool is gone. Dependency is needing real scale-weighed hauls matched to deck photos before any accuracy test means anything. Evidence is a same-day check against 20 archived photos and their known scale weights, before committing the two weeks. Rank is assemble the weighed golden set, test single-species accuracy, then deliberately test mixed hauls where fish overlap in the photo, then one narrow trial with a single dock crew, with the register system left untouched until all of that holds up.

Hand sketched decision tree titled Kill, scope down, or go, a fishing dock. Root, accuracy versus the dock scale. Three branches: holds up on mixed hauls too leads to go use it dockside, only holds up on single species leads to scope down flag mixed hauls for a manual weigh, misses badly everywhere leads to kill not ready for the dock.
A different dock, the same shaped decision: single-species accuracy alone was never going to be the whole answer.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "rank by which mistake can't be undone, put the live build last," and stop.
Cost: only a few days available, not two weeks. Say so honestly, and compress to the golden set and the accuracy test alone, skipping the trial rather than skipping the ordering discipline.
The model really is ready, for real: if the worst-case test also holds up, that's a genuine green light to build the integration next, and saying so plainly is what makes the ranked plan trustworthy instead of reflexive caution.

Where people run it wrong.
They order a spike by what's fastest or most exciting to build, not by which mistake is hardest to take back.
They test only the easy, representative case and report an average that hides the segment likely to fail.
They let the dashboard or integration sneak into the two weeks before the number behind it has ever been checked.

How to use it live. The moment you're asked to design a two-week spike, ask yourself: which possible mistake, if it reached a real user first, could I never fully undo? Put whatever prevents that mistake first, and the rest of the order tends to follow.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a prioritization or sequencing question, and its one-line job?
Tap to flip
ANSWER
ORDER: rank by what's hardest to undo. Outcome, Reversibility, Dependency, Evidence, Rank.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bartholomew Kase, the AI PM at Casterbridge Realty Group, who re-sequenced RenoScope's two-week spike after an audit caught a huge missed estimate.
3 · THE FORCED STEP
What had to happen before any accuracy test could mean anything?
Tap to flip
ANSWER
Assembling a real golden set: 45 listing photos matched to their actual signed contractor invoices.
4 · THE HARDEST MISTAKE
Which mistake in this story was hardest to undo?
Tap to flip
ANSWER
Showing a wrong, confident estimate to a real investor. A private test's wrong number just gets rerun; a live one costs trust that doesn't come back.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating "test the accuracy" and "build the dashboard" as two tasks in whatever order was convenient, instead of ranking by which mistake couldn't be undone.
6 · THE NUMBER
Fill in the blank: the missed estimate quoted ___ dollars in repairs; the real contractor invoice was ___ dollars.
Tap to flip
ANSWER
18,000 dollars quoted; 142,000 dollars actual.
7 · THE REPLAY
Same two weeks, ranked plan already in place. What changes?
Tap to flip
ANSWER
The 142,000 dollar gap surfaces on day nine against a private golden set, costing about $14,000, instead of surfacing after a real investor already bid on a hidden problem.
8 · CROSS PRODUCT TRANSFER
Section 4 runs ORDER again for a different product. Which one, and what's the reversibility call?
Tap to flip
ANSWER
Nordveil Fisheries Co-op's CatchWeigh. A wrong private test number is easy to rerun; a wrong price quoted at the dock costs trust in the tool that doesn't come back.

Check yourself Score: 0 / 0

True or false
1. True or false: the audit found that RenoScope's estimates were equally accurate across all property ages.
  • True
  • False
Show hint
Look at the bar chart comparing homes built before and after 1980.
Show answer
False. Estimates landed within 25 percent of the real invoice 86 percent of the time on newer homes, but only 22 percent of the time on homes built before 1980.
Multiple choice
2. Why is showing a wrong estimate to a real investor harder to undo than a wrong number in a private test?
  • A. Private tests are always more accurate than live ones.
  • B. Investors are legally entitled to a refund if an estimate is wrong.
  • C. A private test can just be rerun, but a bad live number costs trust in the tool that doesn't come back.
  • D. Live tests cost more in server time than private ones.
Show hint
Look at the "which mistake is easy to take back" diagram.
Show answer
C. Reversibility, not accuracy, is what makes the ordering decision: a rerun costs nothing, a broken trust relationship costs the whole program.
Fill in the blank
3. Fill in the blank: the earlier estimate quoted ___ in repairs on a property whose real contractor invoice was ___.
Show hint
Look at "prove it with the miss that almost shipped."
Show answer
$18,000; $142,000. A hidden foundation issue never showed up in any listing photo, so the model had no way to see it either.
Short answer, where it wouldn't matter
4. Name a part of RenoScope's launch that would NOT need to wait for the full spike order to finish, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Launching for post-1980 properties, the segment already proven within budget. It doesn't need to wait on fixing the older-home gap, since that gap is a separate, later fix.
Short answer, apply it yourself
5. Think of a project you've worked on with a fixed short deadline. What was the one task that, done wrong, couldn't be quietly redone later? Was it done first or last?
Show hint
Think about which mistake would have cost something beyond just time.
Show answer
Model answer: Sending a first customer email with the wrong pricing was hard to walk back once a customer had already seen it, so checking that copy went first, well before polishing the email template's design.
Short answer, work the number
6. If the golden set had only included 15 properties instead of 45, all built after 1980, would the two-week spike still have caught the foundation-issue problem?
Show hint
Think about what made the pre-1980 gap visible in the first place.
Show answer
Model answer: No. Without older properties in the golden set, the spike would have shown a strong average result and missed the exact segment where the model was actually blind.
Before you close the answer
Why this works
Tests whether you sequence a spike around which mistake is hardest to undo, or just fill two weeks with whatever's easiest to build first.
Follow-up traps
"Isn't checking the worst-case property type just extra work if the average looks fine?" Response: the average is exactly what hid the 86 versus 22 percent gap here; a segment that fails badly can sit underneath a perfectly healthy-looking overall number.

"What if there's no time for a same-day cheap check before the two weeks start?" Response: then the two weeks themselves absorb that risk, but skipping it entirely means committing the full budget before knowing if the idea is even in the right neighborhood.
If pressed
The narrow blind trial showed five real investors ten estimates without telling them which were model-generated versus analyst-generated; three of five could not reliably tell the difference on post-1980 properties, which is what finally justified showing a live number to a wider group.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more