ConceptIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #11

Describe how you would timebox exploration without killing genuine discovery.

LEADthe calendar was never the right clock

Picture a fabric mill deciding whether a camera bolted over the loom can actually catch weaving defects before they waste a whole roll. Warpline Textiles calls that spike LoomCheck. Priya Vantongeren is the AI PM who had to time it, without a fixed date deciding when real discovery had already ended.

The direct answer
Don't timebox exploration with a date. Timebox it with a signal: count how many genuinely new questions the team writes down each week, not how many days have passed. While that count stays healthy, two or more real new questions a week, let the spike run. The moment it flattens to the same one or two questions repeating, stop, whether that's week two or week nine. LoomCheck's own spike ran five extra weeks past the week its question count had already gone flat, because the only thing anyone was watching was the calendar.
Do this, in order
  1. Track new questions raised per week, not days elapsed, as the real stop signal.Why: a calendar can't tell you whether a team is still discovering or just still working.
  2. Write down what would count as "genuinely new," before the spike starts.Why: a vague signal gets gamed the moment a deadline starts to feel close.
  3. Cost each week of extension in real terms, hours or dollars, and say it out loud.Why: "just one more week" feels free until someone names what it actually costs.
  4. Keep a shared, standing log of every spike's question count, not a private tracker per engineer.Why: without shared history, nobody has a reference for what a normal exploration length even looks like.
  5. Stop and write the finding the moment the signal flattens, even if the calendar says "early."Why: real discovery ending sooner than expected is a result, not a failure to plan around.
  6. Say plainly when a fixed date really is fine, like a low-stakes probe with an obvious, narrow answer.Why: shows judgment about where a signal-based stop earns its cost and where it's overkill.

How to answer this, stage by stage

Nobody is scoring whether you can name a number of weeks that sounds reasonable. They're scoring whether you can name the signal that tells you discovery is actually over.

Stage 1
Scope it to one real spike
Say it like this
"Let's ground this in LoomCheck at Warpline Textiles. That's the spike where the real discovery had quietly ended weeks before anyone actually stopped it."
Why this works
Keeps the answer from becoming an abstract opinion about how long a spike should run.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as LEAD. Link, the business outcome I actually care about. Early signal, what moves before that outcome does. Abuse, how the signal gets gamed. Decision, what I'd do at each level of it."
Why this works
Shows a repeatable way to time a spike, instead of a gut feeling about how many weeks "sounds right."
Stage 3
Reframe: it isn't "how many weeks," it's "is the team still finding new questions"
Say it like this
"This isn't really a question about picking the right number of weeks. It's a question of whether the team is still discovering something new, or just still showing up."
Why this works
This is where a strong answer separates from someone who just names a fixed timebox and calls it done.
Stage 4
Give the early signal
Say it like this
"The signal is new questions logged per week, written down as they come up, not recalled later. Two or more real new ones, keep going. Down to the same one or two, repeating, stop."
Why this works
Names the exact number the whole answer turns on, instead of a vague appeal to "team judgment."
Stage 5
Prove it with the compressed evidence
Say it like this
"LoomCheck's new-question count ran five or six a week through week three, held around three or four through week eight, then flattened to one by week nine. The spike didn't actually stop until week fourteen."
Why this works
Compresses the whole case into the one number that had already told the real story, five weeks early.
Stage 6
Name the AI-specific reasoning and the trade-off, then close
Say it like this
"The honest reason this isn't a generic project-management question is that model exploration has no fixed finish line the way a normal feature build does, you genuinely don't know what you'll find until you look. We accepted the risk of stopping a week or two too early, in exchange for never again burning five extra weeks after the real discovery had already ended."
Why this works
This is the load-bearing, AI-specific judgment: unlike a normal build, a model spike's real endpoint is a fact you discover, not a date you schedule.

Let's learn

What does it actually mean to say a team is "still exploring"?

Before LoomCheck, a quality inspector walked Warpline's weaving floor every two hours, scanning rolls by eye, and usually caught a defect only after several meters had already been woven wrong, about 40 meters of scrap per miss. LoomCheck's spike set out to test whether a loom-mounted camera and model could catch the same defects in real time, before the waste piled up.

Hand sketched icon list titled LEAD the four letters. Four rows: Link the business outcome that matters. Early signal what moves first, shown in a different color. Abuse how the metric gets gamed. Decision what you would do at each level.
The four letters, held up as one page. Early signal is the step a fixed calendar quietly replaces.

Here's the turn: LoomCheck's spike was given a loose "explore until it feels done" timebox, since real discovery, what defects a camera can even see, what lighting works, was genuinely needed. Nobody set a signal for when that discovery had actually ended. The spike kept running on habit long after it stopped teaching the team anything new.

New questions logged per week, and where the real discovery ended
6 3 0 flattens, week 9 Actually stopped, week 14 Week 1
The signal had already gone flat five weeks before the calendar did.

At its worst, a spike that should have taken nine weeks runs fourteen, quietly burning engineering time nobody budgeted for, while the actual finding, sitting there since week nine, waits to be written down.

The spike didn't run long because discovery was slow. It ran long because nobody was watching for the week discovery stopped.
The choice I would take back Warpline kept no shared, standing log of past spikes and how long real discovery had typically taken on each one. That made sense when LoomCheck was the company's first real exploration project. It stopped making sense once a second team started its own open-ended spike with no shared history to check its pace against either.

What I would leave alone: for a narrow, low-stakes probe with an obvious answer, like confirming a camera resolution is even high enough to see a thread, a fixed two-day timebox is genuinely fine. There's no real discovery curve to protect.

The lesson: a spike doesn't end when the calendar says so. It ends the week the team stops finding anything it didn't already know, and that week rarely lines up with the date anyone originally guessed.

Now here is the same thing as a story

The short version above is what you'd say defending the timeline in a steering review. Read this one for what it felt like the week a sibling team's cancelled project made everyone look twice at Priya's own.

Her name is Priya. She has run feasibility spikes at Warpline for four years, and she is the one people trust to say honestly when something isn't working yet.

Hand sketched flow diagram titled The loop a timebox should protect, third step emphasized. Five steps left to right: Form a hypothesis. Run the probe. Check the signal, shown in a different color. Extend or stop. Report the finding.
The third step is the one a fixed calendar quietly skips.

LoomCheck started with a printed sheet Priya annotated by hand at the loom, one row per week, jotting down whatever the team was still trying to figure out. Weeks one through three were full of new lines: which lighting angle showed thread breaks best, whether dye streaking looked different from tension faults, how far upstream a defect could be caught.

Hand sketched labeled parts diagram titled What counts as genuine discovery. A question mark box icon at the center labeled Still Exploring, with four labeled callouts around it: New questions per week, A surprise an expert didn't predict, Cost of one more week, Stakeholder patience left.
Four things a real discovery signal has to account for, none of which a fixed date can see.

Without a shared log of how long past spikes at Warpline usually took, individual engineers started keeping their own private spreadsheets tracking hypotheses tried and days spent, since there was no official one to check against. Nobody built it maliciously. It just quietly became the only record that existed.

Knowledge spark: why would engineers build their own private tracking spreadsheet? When there's no shared, standing history of how long exploration normally takes on a project like this, people still need some way to judge whether they're on track. Without an official signal to lean on, they build a private one, quietly, one spreadsheet at a time, and nobody else ever sees it.

A sibling team's automated color-matching spike, run the same open-ended way, hit month six with nothing to show but "we're still exploring," and got cancelled outright once leadership noticed. That failure, not LoomCheck's own numbers, was what finally got someone to ask how long LoomCheck itself had actually been running.

Hand sketched comparison titled Two clocks. Left panel, a gauge icon labeled Lagging metric, caption spike quality score, rings late. Right panel, a gauge icon labeled Leading metric, caption new questions per week, rings first.
The leading clock had already rung, five weeks before anyone thought to check the lagging one.

The real question was never how many weeks LoomCheck should have run. It was whether anyone was tracking the one number that would have said, out loud, exactly when to stop.

Hand sketched metaphor scene titled How it gets gamed. Left, a document icon labeled On Paper, caption a checkbox, hours logged as spent. Right, a person icon labeled In Reality, caption the real question, still unanswered.
A metric that only counts hours spent can look busy while the actual question sits untouched.

When LoomCheck's loose timebox was first set, someone said, "let's not rush real discovery, we'll know when we're done," and it sounded reasonable, since at the time nobody had a signal to check that feeling against.

Total spike cost, under three different stopping rules
$60k $30k $0 $8k, too early Fixed 2-week cutoff $36k Signal-based stop, wk 9 $56k Actual, ran to wk 14
The signal-based stop captures the same real discovery as what actually happened, for 20 thousand dollars less.

Rerun the same fourteen weeks with a shared question-count log tracked from week one: the flattening at week nine gets flagged the same week it happens, the finding gets written up while it's fresh, and the last five weeks of quiet, habitual work never get spent at all.

What I'd tell myself, hearing about that cancelled sibling project: the calendar was never a signal. It just felt like one, right up until someone finally checked.

LEAD, the clock that actually matteredNot a script for cutting every spike short. LEAD is what tells you the exact week discovery quietly became habit.

L
Link. The business outcome that actually matters.
Whether LoomCheck deserves real engineering investment next quarter, decided cheaply, not whether the spike ran a tidy number of weeks.
Grounds the whole answer in the decision the spike exists to inform.
E
Early signal. What moves before the outcome does.
New genuinely novel questions logged per week. It flattened to one by week nine, five weeks before the spike actually stopped.
This is the hardest step, and the one a fixed calendar quietly replaces with a guess.
A
Abuse. How this metric gets gamed.
A team under deadline pressure can pad the count with trivial rephrasing of old questions, or log vague, unfalsifiable "questions" that were never real hypotheses to begin with.
Naming the metric's own weak point is what keeps it honest instead of decorative.
D
Decision. What you'd do at each level.
Two or more real new questions a week, extend, at roughly 4 thousand dollars a week in engineer time. One or fewer, repeating, stop and write the finding regardless of the calendar date.
A metric nobody acts on is a dashboard decoration. This one has a stated action at every level.

The recap, one line per letter: link is whether LoomCheck earns real investment next quarter, early signal is new questions logged per week, abuse is padding the count with trivial or vague entries, and decision is extend at two or more real questions a week, stop the moment it flattens.

And if you want to be sure it really works, try it somewhere elseSame four letters, a small-town newsroom instead of a weaving floor. The signal stays the same shape, the way people stop watching it doesn't.

Rutger Achebe-Lindstrom edits The Fenmoor Gazette, where a spike tested whether a model could draft headline options for an editor to pick from. Mapped onto LEAD: link is whether the tool earns a permanent spot in the daily workflow. Early signal is the same idea, new question count per week, until the team stopped tracking it once the model's drafts started sounding consistently good. Abuse would have been logging vague, feel-good notes instead of real open questions. Decision should have been to keep the human check regardless of how polished the drafts felt, since a felt sense of "it's fine now" is not the same thing as a tracked signal.

Hand sketched comparison titled Checking a draft headline. Left panel, a document icon labeled Week 1, caption editor checks every drafted headline. Right panel, a question mark box icon labeled Week 9, caption editor stops checking, ships them cold.
A different flip than LoomCheck's: not a private workaround, but an editor who stopped checking once the drafts felt reliably good.
Hand sketched decision tree titled Extend the timebox, or stop. Root, did this week teach us something new. Three branches: new questions still forming leads to extend one more week, shown in a different color. Same 2 questions as last week leads to stop write the finding. No new questions at all leads to stop decide now.
The same three-way call LoomCheck should have made five weeks earlier than it did.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "count new questions per week, not days elapsed, and stop when that count goes flat," and stop.
Cost: there's no time to build a formal tracking process before the spike starts. Say so honestly, and keep even a simple shared note of new questions each week, rather than relying on individual memory of how things feel.
The exploration goes faster than expected, for real: if genuinely new questions dry up by week two instead of week nine, that's still a real result worth writing up immediately, not a sign something went wrong.

Where people run it wrong.
They timebox exploration with a fixed date chosen before anyone knew what there was to discover.
They let each person's private sense of progress stand in for a shared, trackable signal.
They keep exploring out of habit once a project already has momentum, long after the real discovery has stopped.

How to use it live. The moment someone asks how to timebox exploration, ask yourself: what's the one number that tells you discovery is still happening, and who's actually writing it down each week. Name it, and the rest of the timebox follows on its own.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip: with no shared, standing log of past spike lengths, engineers quietly built their own private spreadsheets to track progress, instead of raising it as a shared, official signal.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Vantongeren, the AI PM at Warpline Textiles, who ran LoomCheck's feasibility spike on a loose, calendar-based timebox.
3 · THE HABIT
What did the team keep doing, out of habit, once real discovery had already flattened?
Tap to flip
ANSWER
They kept exploring on the same loose schedule for five more weeks, since nobody was watching the new-question count that had already gone flat at week nine.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Raising a progress question through a shared, official process, versus quietly building a private spreadsheet instead. No shared middle ground existed once the official log never got built.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never keeping a shared, standing log of past spikes and how long real discovery typically took on each one, leaving every team to guess its own pace.
6 · THE NUMBER
Fill in the blank: LoomCheck's new-question count flattened by week ___, but the spike didn't actually stop until week ___.
Tap to flip
ANSWER
Nine, and fourteen.
7 · THE REPLAY
Same fourteen weeks, with a shared question-count log tracked from week one. What changes?
Tap to flip
ANSWER
The flattening at week nine gets flagged the same week it happens, the finding gets written up while it's fresh, and the last five weeks of habitual, unproductive work never get spent.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
The Fenmoor Gazette's headline-drafting spike. The flip is over-trust: the editor stopped checking drafted headlines once they started sounding consistently good.

Check yourself Score: 0 / 0

Multiple choice
1. What is the early signal this answer uses to decide when to stop exploring?
  • A. The total number of weeks the spike has run.
  • B. The number of genuinely new questions the team logs each week.
  • C. The number of engineers assigned to the project.
  • D. The model's accuracy on the held-out test set.
Show hint
Look at the E step in the LEAD recap.
Show answer
B. A calendar can't tell you whether a team is still discovering something new, but a flattening question count can, and does, weeks before anyone notices otherwise.
True or false
2. True or false: LoomCheck's spike stopped the same week its new-question count flattened.
  • True
  • False
Show hint
Look at the line chart, "new questions logged per week."
Show answer
False. The count flattened at week nine, but the spike kept running for five more weeks, until week fourteen, because nobody was tracking that signal.
Fill in the blank
3. Fill in the blank: a signal-based stop at week nine would have cost about ___ thousand dollars, compared to about ___ thousand for what actually happened.
Show hint
Look at the bar chart, "total spike cost, under three different stopping rules."
Show answer
36; 56. The extra 20 thousand dollars was spent entirely after the real discovery had already gone flat.
Short answer, where it wouldn't matter
4. Name a kind of spike where a fixed, calendar-based timebox would genuinely be fine, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A narrow, low-stakes probe with an obvious answer, like confirming a camera's resolution is high enough to see a thread. There's no real discovery curve there to protect.
Short answer, apply it yourself
5. Think of a time you or your team kept working on something past the point it was teaching you anything new. What would a "new questions per week" count have shown, if you'd tracked it?
Show hint
Think about the difference between staying busy and still discovering something.
Show answer
Model answer: A months-long research rabbit hole that kept circling the same two or three open questions. Tracking a weekly count would have shown the flattening well before anyone admitted it out loud.
Short answer, work the number
6. If LoomCheck's weekly extension cost had been 8 thousand dollars instead of 4 thousand, would the case for stopping at week nine be even stronger or weaker?
Show hint
Think about how the cost of each extra week changes the trade-off in the D step.
Show answer
Model answer: Stronger. A higher weekly cost makes the five wasted weeks after week nine even more expensive, which only strengthens the case for stopping the moment the signal flattens.
Before you close the answer
Why this works
Tests whether you'll timebox exploration with a guessed date, or find the one signal that tells you, honestly, the week real discovery actually ended.
Follow-up traps
"Isn't counting questions just as gameable as any other metric?" Response: yes, which is why the answer names that risk directly, padding with trivial rephrasing, and pairs the count with a plain definition of what counts as genuinely new, agreed before the spike starts.

"What if the team is embarrassed to admit discovery has stopped?" Response: that's exactly why the signal needs to be shared and tracked from week one, not left to any one person's private read on how things feel.
If pressed
The shared log this answer proposes also tracks which new questions turned into an actual design change versus which were dead ends, so a future spike can compare not just how many questions came up, but how many of them actually mattered.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more