CaseAdvancedAI Opportunity & Model Strategy / Opportunity identification for AI / #8
Describe how you would size an AI opportunity before knowing whether it is technically feasible.
GUARD · a guess wearing an estimate's clothes
Plumbline is Anvilrow Technologies' tool that reads a stack of blueprint PDFs for a construction project and drafts a full material takeoff and labor estimate, a bid draft a general contractor's estimator checks before it goes out the door. Leofric Cartwright, the product manager who owns Plumbline's feasibility and sizing calls, found out how far a clean number can travel before anyone tests it. Dorothea Kingsley, Anvilrow's VP of Growth, sized a new market at fourteen million dollars and carried it straight to the board. Ysabel Everhart, who owns Everhart Restoration Group, just wanted a bid she could trust before her busiest season started.
The direct answer
Size it as a range tied to which technical approach turns out to work, not as one clean number. Say both scenarios out loud: what the opportunity is worth if a lightweight fix closes the gap, and what it's worth if a heavier approach turns out to be needed, then name the spike that will tell you which one you're in and put a date on it. No board slide, roadmap slot, or customer promise gets built on top of the number until that spike has actually run.
Do this, in order
Size it as a range tied to feasibility, and say so plainly, not as one smoothed number.Why: this is the direct answer, and every item below just protects it.
Run the technical spike on real, untested documents before any number leaves the building.Why: Plumbline's 95 percent was a real number. It was also measured on a completely different population of documents.
Put a hard date on the spike, and hold every roadmap slot and customer promise until it passes.Why: an undated "we'll find out" lets a guess sit in a board deck indefinitely, looking exactly like a checked number.
Track how many sizing memos carry an explicit range and a named spike, every quarter, not just after the one that went wrong.Why: a quiet quarter can mean the discipline is working, or it can mean everyone's gone back to guessing.
Leave the sizing process alone for expansions that plug into the same clean document type the model has already proven itself on.Why: gating what doesn't need gating just teaches the team the whole process is theater.
When a spike comes back low, revise the number loudly, with a dated note saying what changed.Why: a quiet downward revision teaches the next sizing exercise nothing at all.
How to answer this, stage by stage
Nobody is grading whether you think Dorothea's math was good. They're grading whether you can turn "size it before you know if it works" into a specific, defensible number, or two.
1
Scope it to one product and one number
Say it like this
"Let me make this concrete. Say a company builds an AI tool that reads blueprint PDFs and drafts a construction bid. It works well on new-construction projects. Someone wants to size the opportunity of expanding it into renovation projects, where the blueprints look completely different. That's the scenario I'll run, because 'size an opportunity before you know if it's feasible' means nothing until there's an actual number and an actual unknown sitting under it."
Why this works
Keeps the answer from turning into generic market-sizing advice.
2
Say your structure out loud
Say it like this
"I'll run this as GUARD. Groups: who's actually carrying the risk of a wrong number. Unequal: whose plans get built on top of it without ever seeing the uncertainty. Ability to contest: can anyone downstream tell how solid the number really is. Reduce: the actual fix to how the number gets sized. Detect: how you'd know sizing discipline is slipping."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about the number.
3
Reframe what "sizing" actually means here
Say it like this
"This isn't really a market-sizing question. A market is fully knowable, you just have to go find the data. This is different. The honest answer to 'how big is this' depends on whether a model can do something nobody's tested yet, read messy real scans instead of clean CAD files, and that's not a research question. It's an engineering question with a real answer. We just don't have it yet."
Why this works
Separates this from generic estimation-under-uncertainty and locates the real judgment: a model capability, not a market fact, is the unknown.
4
Give the one decision
Say it like this
"Here's what I'd actually do: size it as two numbers, not one. If a lightweight fix gets the model working on the new document type, the opportunity is this much. If it needs a heavier approach, it's this much less, near term, and this much more, later. I name the two-week spike that tells us which one we're in, and I put a date on it."
Why this works
This matches the direct answer word for word. If it doesn't, the interviewer notices before you do.
5
Prove it with the compressed failure
Say it like this
"Here's what happens without it. A VP sizes an opportunity at fourteen million, carrying an old accuracy number forward onto documents nobody's tested it against. The board approves it. Engineering gets a ten-week slot. Sales promises a customer a pilot. Six weeks in, someone finally checks, and the real number is sixty-one percent, not ninety-five. Now there's a blown roadmap slot, a customer four weeks from her season with a bid tool she can't fully trust, and a board still holding a number nobody's told them is wrong."
Why this works
This is the story below, compressed to four sentences. The long version proves it actually happened this way.
6
Say what you'd watch, and what you'd leave alone
Say it like this
"I'd track how many sizing memos carry an explicit range and a named spike date, every quarter, because that number tells you if the discipline is real or if everyone's quietly gone back to guessing. And I'd leave the sizing process alone for expansions that use the same clean documents the model's already proven itself on. That's not where the risk lives."
Why this works
Shows judgment about where the real risk sits, not blanket caution slapped onto every sizing exercise.
7
Close on something checkable
Say it like this
"You'll know it's working when a number that turns out wrong gets revised loudly, with a dated note saying what changed. You'll know it's broken when the roadmap just quietly re-scopes around a smaller plan and nobody ever tells the board the original number was wrong."
Why this works
Ends on a test the interviewer could actually go verify, not a promise that it's handled.
Let's learn
Plumbline sits behind one upload button on Anvilrow's estimating dashboard, the button a general contractor's estimator clicks after dragging in a stack of blueprint PDFs. A few minutes later, a full material takeoff and a labor estimate come back: how much lumber, how much concrete, how many sheets of drywall, priced and ready for the estimator to check before the bid goes out.
Knowledge spark: what's a takeoff?
A takeoff is the exact list of amounts a project needs, counted straight off the blueprints: board feet of lumber, cubic yards of concrete, square feet of drywall. Every bid starts with one. Get the takeoff wrong and the whole bid is wrong under it.
On new-construction projects, clean CAD-native blueprints straight out of an architect's design software, Plumbline is checked hard. Against 600 real sheets, matched by hand against a certified estimator's own takeoff, it lands within 5 percent of the right quantity 95 percent of the time. That's the number on Anvilrow's homepage, and it's a real, earned number.
Renovation and retrofit projects run on a different kind of document entirely. Blueprints for an existing building get scanned, marked up with hand-drawn revision clouds where something changed mid-project, stamped by two or three different permit reviews, sometimes carrying two different scale bars from two different renovation eras on the same sheet. Nobody had ever run Plumbline against a real batch of them.
One dot on this chart had actually been checked when Dorothea built her slide. The other two were still just a guess wearing a decimal point.
Here's the turn. Dorothea Kingsley, Anvilrow's VP of Growth, wanted to size the renovation-bid opportunity for the board. She took the 95 percent number, the only accuracy figure she had, and carried it forward onto a market of general contractors whose blueprints Plumbline had never once been tested against. The mistake wasn't the market math. Renovation-focused contractors really are a big market. The mistake was letting one clean number stand in for a technical question nobody had actually answered yet.
Four things on the slide. Three of them are real math. None of them says the accuracy number had never been checked on this kind of document.
We didn't lose the number. We lost the sentence saying nobody had checked it.
At its worst, this is Sixtus Fairfax's engineering team six weeks into a ten-week roadmap slot, building screens and an upload flow for a capability the model can't yet reliably run. It's Ysabel Everhart, who owns Everhart Restoration Group, four weeks from her busiest bid season, holding a promised pilot that only works cleanly on two of the six renovation jobs she tried it on. And it's a board still carrying fourteen million dollars in projected revenue that nobody has told them is built on a number nobody checked.
The choice I would take back
Nothing in Anvilrow's process required a sizing number, once it left the building as a board slide or a customer promise, to carry a check against real, untested documents. That was fine the first two times Anvilrow expanded into a new segment, because both of those segments still used the same clean CAD blueprints Plumbline had already proven itself on. It stopped being fine the first time the new segment's documents were a genuinely different kind of thing, not just a bigger pile of the same kind.
What I would leave alone
Sizing an expansion into new states or new trades for new-construction work never needed a spike gate. The documents are the same clean CAD files Plumbline was built and tested on. Gating that too, just to look consistent, would only teach the team the whole process is theater.
The lesson: an opportunity that depends on an untested model capability isn't a market-sizing problem wearing a technical costume. It's a technical question wearing a business suit. The job isn't picking a confident number. It's finding out, on real documents, what's actually true before the number ever leaves the room.
Now here is the same thing as a story
Stage five above compresses this into four sentences. Here's the ten weeks underneath, the part a stand-up answer skips.
Leofric Cartwright can tell, from the first three sheets of a blueprint set, roughly how good a week Plumbline is about to have. Four years into owning the model's feasibility and sizing calls, he reads a set of drawings the way an old mechanic listens to an engine: not for what's wrong yet, just for what kind of thing he's looking at.
Dorothea Kingsley runs growth at Anvilrow, and she's good at her job in a specific way: she can turn the rough shape of an opportunity into a number a board will actually act on. Ten weeks before the quarter she'd usually build a growth deck for, she pulled up Plumbline's own numbers page. Ninety-five percent, on new-construction takeoffs. She multiplied that against the renovation-contractor market, a market Anvilrow's sales team had been asking about for a year, and landed on fourteen million dollars in new revenue within eighteen months, if the self-serve motion that worked for new-construction customers worked here too.
She didn't invent the ninety-five percent. She just never asked whether it applied to a document nobody had shown her yet. Nobody in the room asked either. The board approved the number on a Tuesday in April, and by Thursday, Sixtus Fairfax's engineering team had a Q3 slot: ten weeks to ship renovation-bid support.
Neither person here did anything careless. One of them just had the only copy of the assumption underneath the figure.
Dorothea, wanting a flagship account to point to, called Ysabel Everhart that same week. Ysabel owns Everhart Restoration Group, a renovation-focused general contractor two towns over, the kind of company the fourteen-million-dollar number was supposed to represent by the thousand. Dorothea promised her a working pilot in time for the fall bid season, ten weeks out. Ysabel, who'd spent years pricing renovation jobs by hand, said yes on the spot.
For six weeks, Sixtus's team built. A new upload flow for renovation projects. A revised results screen. Nothing about the model itself, because nobody had asked them to touch it. Why would they. The number said ninety-five percent.
In week four, in an ordinary Tuesday planning meeting, Cassander Woodrell, three weeks into the job, asked the question nobody else in the room had thought to say out loud: "Wait, has anyone actually run this on a real renovation set yet? Like an actual scanned one?" The room went quiet in the specific way a room goes quiet when the answer is no and everyone realizes it at the same moment.
Four of these five weeks looked completely fine, if the only thing anyone was watching was the roadmap burn-down.
Leofric ran the spike that should have run back in April. Two weeks, forty real renovation blueprint sets, two hundred forty pages, pulled from Anvilrow's own customer files. He also, on a hunch, ran fifteen sets of historic-restoration drawings, hand-drafted originals from buildings old enough that nobody had ever digitized the plans. The renovation number came back at sixty-one percent. The historic number came back at thirty-four.
The model wasn't guessing randomly. It was doing something specific, and once Leofric looked closely, obvious in hindsight: it read a hand-drawn revision cloud, the scribble an architect draws around a part of a drawing that's changed, the same way it read a real wall boundary. Both are just closed loops on a scanned page. Nothing in training had ever told it the two were different.
We didn't take the ninety-five percent away from Plumbline. We took it from a set of documents it had never once seen.
By week six, with four weeks left before Ysabel's bid season, Leofric had to tell Dorothea, Sixtus, and the board three things at once: the model couldn't reliably run renovation bids yet, not at the quality the fourteen-million number assumed. Sixtus's team had spent six of ten weeks building around a capability that wasn't actually there. And Ysabel was going to find out, one way or another, before her season started.
The decision Leofric would take back sits earlier than any of that. Anvilrow had expanded Plumbline twice before, once into a new state's building codes, once into a new trade's material list, and both times the ninety-five percent held, because both times the underlying documents were the same clean CAD files the model had always seen. Two clean wins in a row had quietly taught everyone the number just travels. Nobody had ever built a rule saying it wouldn't always.
So here's what changed. Every sizing memo that reaches the board or becomes a roadmap commitment now has to carry two numbers, not one, each tied to a named scenario, plus a dated spike that resolves which one is real. Dorothea's redo, run properly this time: if a lightweight fine-tune closes the gap on renovation documents, the opportunity is eleven million in eighteen months, close to the original. If it needs a heavier pipeline, handwriting-aware extraction plus a mandatory human check before any renovation bid ships, it's three million in eighteen months, but could reach twenty million in three years once the review cost comes down. The two-week spike, run in week two this time instead of week six, is what tells you which one you're actually building toward.
Ysabel got a real answer in week two of the redo, not week six: the heavier pipeline was needed, her bids would carry a short human check for now, and Plumbline would still save her most of a day per bid, just not the fully hands-off version she'd been promised. She kept the pilot. She just kept it with an accurate picture of what it actually was.
What I'd tell myself, sitting in that April board meeting where nobody asked if the number applied: two wins in a row isn't a pattern. It's a coincidence you haven't been charged for yet.
GUARD, checked against one clean number and a ten-week clock
This was never really about whether Dorothea's math was good. It was good. GUARD is for naming who actually pays when a number nobody's tested reaches a roadmap before the thing it's about does.
GGroups. Who's actually carrying the risk of a wrong number.
Sixtus Fairfax's engineering team, committed to a ten-week roadmap slot for a capability nobody had confirmed existed at the assumed quality bar. Dorothea herself, who made a real downstream promise to a real customer based on a number she had no way yet to know was solid. And Ysabel Everhart, who has to live with whatever Anvilrow actually built, whether or not anyone told her it was sized off an assumption.
None of these three did anything careless. Dorothea sized the best number she had. Sixtus's team built to the roadmap they were given. Naming all three before reacting to just the number is the whole first move.
UUnequal. Where the harm actually lands.
Sixtus's team and Ysabel are the ones who absorb the cost of a wrong number, and both sit several steps removed from the spreadsheet where the extrapolation happened. Neither ever saw the assumption. They only ever saw the number that came out the other end, sitting in a roadmap slot or a promised date, looking exactly as solid as a checked one.
The harm doesn't land on whoever made the guess. It lands on whoever built real plans on top of it, usually without ever seeing that it was a guess at all.
AAbility to contest. Could anyone downstream tell how solid the number really was.
Dorothea's board slide showed one figure, fourteen million dollars, eighteen months, with no confidence range and no note saying it depended on an untested capability. Sixtus had no way to know the number he was building a roadmap slot around had never been checked on the documents it actually depended on. The only real check that ever happened was an informal question from an employee three weeks into the job, not anything built into the process itself.
This is GUARD's sharpest question for this exact scenario: not is the number wrong, but could anyone who had to build on top of it have told.
Four of these five steps happened exactly as designed. The missing one, a completed spike, was never built into the path at all.
Quantity-extraction accuracy by blueprint source, Plumbline's eval sets
New-construction, testedRenovation, tested lateHistoric, tested late
Only the first bar had been checked when Dorothea's number reached the board. The other two were carried forward from it, unverified, for six more weeks.
RReduce. The actual fix, not a policy memo.
Any sizing memo that reaches the board or becomes a roadmap commitment now carries an explicit range tied to a named feasibility scenario, plus a dated technical spike that resolves which scenario is real. No external commitment, a customer promise, a fixed launch date, ships before that spike has actually run.
The alternative worth naming and rejecting: require a completed spike before any opportunity even reaches a sizing conversation at all. That fully removes the risk of a wrong number. It also means every one of the dozen-plus candidate opportunities Anvilrow ranks in an average quarter, most of which get killed in the first five minutes of that conversation, would need two weeks of engineering time first, which would grind the whole prioritization process to a stop.
Four lines. None of them apologize for not knowing yet. They just say what's actually known, and what isn't.
DDetect. How you'd know sizing discipline is slipping.
Track the share of sizing memos that reach the board or a roadmap decision carrying an explicit, dated range, every quarter. Before the fix, that share sat between 7 and 14 percent, whatever a given VP happened to include on their own. After the fix, it climbed past 90 percent within two quarters.
The failure worth naming plainly: if a wrong number only ever gets fixed quietly, the roadmap re-scoped around a smaller plan with nobody told the original figure was wrong, that's not sizing discipline. That's just a mistake nobody's allowed to name out loud.
Share of sizing memos carrying an explicit, dated range, by quarter
Share of sizing memos with an explicit, dated range, per quarter
Dorothea's slide shipped in the lowest quarter on record, 7 percent. The fix didn't just recover from that. It reversed a slide that was already heading toward zero.
The trade-off, said out loud
Routing renovation-type bids to a mandatory human check adds roughly 45 minutes and 65 dollars of reviewer time per bid, on exactly the segment where the model's accuracy doesn't clear the bar on its own yet. Anvilrow took that cost on purpose, for the segment carrying real risk, instead of either shipping wrong bids for free or reviewing every bid, new-construction included, by hand, which would have erased the instant turnaround that makes Plumbline worth paying for in the first place.
And if you want to be sure it really works, try it somewhere else
Same five letters, a claims-coding tool instead of a bid tool, and this time the untested documents are surgical operative notes instead of scanned blueprints.
Codewright, built by Marrowbank Health Analytics, reads a clinician's note and assigns the billing codes for a claim, so a practice can submit it without a coder starting from scratch. Anatoly Cavanaugh owns Codewright's coding model the way Leofric owns Plumbline's.
Same three branches Anvilrow faced with Plumbline, a different document entirely. The honest branch is still the one that survives an actual spike.
On primary-care visit notes, checked against 800 real claims and a certified coder's own assignment, Codewright gets the full code set right 92 percent of the time. Sylvester Millhouse, who runs revenue-cycle partnerships, sized the opportunity of expanding into orthopedic surgical coding at nine million dollars in new revenue within twelve months, carrying the 92 percent number forward the same way Dorothea carried Plumbline's.
Same rank, mapped onto Codewright: size it as two numbers. Surgical operative notes bundle several procedures into one note, and deciding which ones are separately billable and which are folded into another code is a judgment call nobody had tested the model against. A two-week spike against fifty real orthopedic notes found the real accuracy at 58 percent, not 92, because bundling logic doesn't behave anything like a single visit's diagnosis code. If a lightweight fix, flagging bundling questions for a human coder to check, closes enough of the gap, the opportunity is seven million in twelve months. If it needs a separate model trained specifically on surgical bundling rules, it's two million near term, climbing toward fifteen million in three years.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: never let a single point estimate leave the room without naming which feasibility scenario it assumes. A number with no range attached is a guess wearing an estimate's clothes.
Cost: no budget for a full two-week spike this quarter. Run a scoped one-day test on 10 to 15 real documents instead of zero. A small real sample beats an extrapolation from a completely different population.
The model got better, for real: say Plumbline's renovation accuracy jumps to 90 percent next quarter after a training pass on scanned documents. That's a reason to update the range and lift the mandatory-review requirement, on purpose, with a new eval run backing it up, not a reason it should have been assumed at 95 percent from day one.
Where people run it wrong.
They treat sizing as a market-research exercise only, and never connect the number to whether the model has actually been tested on the real documents it would need to read.
They let whichever number reaches the board first become the anchor everyone quietly designs around, even after a spike proves it wrong.
They wait for a blown deadline or an upset customer to reveal the gap, instead of gating any external promise behind a dated feasibility check.
How to use it live. Before answering, ask out loud: "has anyone actually run this on real renovation blueprints yet, or are we assuming it behaves like new-construction ones?" Naming that split buys you a few seconds of thinking time, and it's most of the real answer.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
Which framework fits "size an AI opportunity before you know if it's technically feasible"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real test isn't whether fourteen million is a good guess, it's whether anyone built a way to check the guess before real time and a customer's schedule got committed to it.
2 · THE PEOPLE
Who are the people this answer names?
Tap to flip
ANSWER
Leofric Cartwright, Plumbline's product manager. Dorothea Kingsley, VP of Growth, who sized the renovation opportunity. Sixtus Fairfax, the engineering lead who owned the Q3 slot. Ysabel Everhart, owner of Everhart Restoration Group, promised a pilot. Cassander Woodrell, the new engineer who asked the question.
3 · THE OLD HABIT
What did the team default to, out of habit, before there was a real process?
Tap to flip
ANSWER
Carry the last proven accuracy number forward onto a new population of documents, because the last two expansions happened to use similar documents and the pattern held, so nobody ever built a required check for when it wouldn't.
4 · THE TWO-WAY TRAP
What's the two-way trap this question is actually testing?
Tap to flip
ANSWER
Commit to one clean confident number too early, and you might burn a roadmap slot and a customer's trust on a capability that isn't there yet. Refuse to size anything until every unknown is resolved, and you can't rank opportunities against each other at all, so nothing ever ships.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Nothing in Anvilrow's process required a sizing number, once it left the building as a board slide or a customer promise, to carry a check against real, untested documents. Dorothea's fourteen million was a single number because nothing forced it to be anything else.
6 · THE NUMBER
Fill in the blank: Plumbline's accuracy holds at ___ percent on ___ new-construction sheets, but drops to ___ percent on the 40 real renovation sets Leofric finally tested.
Tap to flip
ANSWER
95 percent on 600 sheets. 61 percent on the renovation sets. That 34-point gap is the whole story: a real number, applied to the wrong population.
7 · THE REPLAY
Same ten weeks, new process. What changes?
Tap to flip
ANSWER
Dorothea's slide carries two numbers with a spike date attached. The spike runs in week two, not week six. Ysabel gets told honestly, in week two, which scenario Plumbline is actually in, with eight weeks left to plan around it instead of four.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and who plays the equivalent roles?
Tap to flip
ANSWER
Codewright, Marrowbank Health Analytics' claims-coding tool. Anatoly Cavanaugh plays Leofric's role. Sylvester Millhouse plays Dorothea's role, sizing an orthopedic surgical-coding expansion at nine million before anyone tested whether the model could handle bundled procedure codes at all.
Check yourself Score: 0 / 0
True or false, with why
1. True or false: Plumbline's 95 percent accuracy number was wrong.
True
False
Show hint
Check the "Let's learn" section's opening numbers, and the turn paragraph.
Show answer
False. 95 percent was a real, correctly measured number on 600 new-construction sheets. The mistake was applying it to renovation blueprints, a population it was never tested against.
Multiple choice
2. Why doesn't it work to require a completed technical spike before every single opportunity even reaches a sizing conversation?
A. Because spikes are always inaccurate and shouldn't be trusted.
B. Because most candidate opportunities get killed in the first ranking conversation, and requiring two weeks of engineering time before any of them could even be roughly compared would grind that whole process to a stop.
C. Because engineers refuse to run more than one spike a quarter.
D. Because a spike only ever works on new-construction documents.
Show hint
Check the Reduce step's rejected alternative, in the GUARD recap.
Show answer
B. The fix gates the moment a number leaves the building, a board slide or a customer promise, not the earlier, cheap conversation where most ideas get ranked and killed.
Fill in the blank
3. Plumbline's quantity-extraction accuracy is ___ percent on new-construction sheets, checked against 600 real sheets, and ___ percent on the 40 real renovation sets Leofric's spike tested.
Show hint
Check flashcard 6, and the bar chart in the GUARD recap.
Show answer
95 percent, and 61 percent. That gap is what a single carried-forward number hid for six weeks.
Short answer, name the rejected alternative
4. Anvilrow considered one other fix besides the conditioned-range rule. What was it, and why did they reject it?
Show hint
Check the Reduce step in the GUARD recap.
Show answer
Model answer: Require a completed technical spike before any opportunity even reaches a sizing conversation. Rejected because Anvilrow ranks a dozen or more candidate ideas a quarter, most killed in the first five minutes, and requiring two weeks of engineering time before any of them could even be roughly sized would grind that whole prioritization process to a stop.
Short answer, apply it yourself
5. Think of a product you use that involves an AI judgment call. Name one "opportunity" someone might size for it, and what real technical unknown that sizing would quietly assume away.
Show hint
Look for a place where an existing accuracy number gets carried forward onto a new kind of input it's never been tested on.
Show answer
Model answer: A receipt-scanning expense app sizing an expansion into handwritten paper receipts, carrying forward its accuracy on printed receipts. Whether the model can read handwriting at all is a completely different, untested question.
Fill in the blank, work the number
6. The share of sizing memos with an explicit range fell from 14 percent in Q2 to 7 percent in Q3. If that same decline had continued for two more quarters instead of the fix shipping in Q4, about where would Q5's share have landed, compared to the actual post-fix figure of 87 percent?
Show hint
The drop from Q2 to Q3 was roughly half. Apply that same roughly-halving step twice more.
Show answer
Roughly 2 to 3 percent, versus the actual 87 percent. The fix didn't just stop a decline. It reversed one that was already heading toward zero.
Before you close the answer
Why this works
Tests whether you can size an opportunity for a capability nobody's proven yet without either freezing, never committing to a number, or borrowing false confidence from a completely different, already-proven capability. Most candidates give one clean number. That's the exact failure this question is built to catch.
Follow-up traps
"What if you genuinely don't have two weeks before the board needs an answer?" Response: ship the range anyway, with a stated "spike not yet run" flag and a date it will land, rather than a false single number. A dated placeholder is more honest than manufactured precision.
"Isn't giving a range just a way to avoid committing to a real number?" Response: no, because the range is bounded by two named, buildable scenarios, a lightweight fix or a heavier pipeline, not vague hedging, and the spike date is itself a real, checkable commitment.
If pressed
The specific reason scanned renovation blueprints broke the model: its extraction layer treats a hand-drawn revision cloud, the scribble marking a changed area, the same way it treats a real wall boundary, because both render as a closed, hand-drawn loop on a scanned page. Nothing in training ever taught it revision clouds aren't walls, which is exactly the kind of gap a clean number can't show you, only a real document can.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.