ConceptIntermediateModel Fluency & the AI PM Role / The AI literacy baseline every PM needs / #16
Describe what a reasoning model does differently and when the extra cost is worth it.
SPARK · the rule that decides when Sinew, Keelson's contract-risk reader, actually pays for the slow careful read
Keelson builds Sinew, an AI tool that reads a contract before signature and flags the clauses that carry real risk. Stanwick Components, an industrial parts manufacturer, runs every vendor and customer contract through it. Its senior contract reviewer, Bronwyn Merrion, is the one who has to trust, or double check, whatever Sinew tells her.
The direct answer
A reasoning model doesn't answer in one pass. It works through a problem in steps, checking earlier steps against later ones, before it commits to an answer, at real extra time and cost. Never leave that switch to a static setting or a person's memory. Route to it automatically, from a cheap check of the request's own structure, like how many risk clauses in a contract actually reference each other, not from a label someone set once and forgot.
Do this, in order
Build one cheap, automatic check that reads the document's own structure, and route to reasoning mode only when it says so.Why: a static category label describes what a contract used to look like, not what this one actually contains.
Compute the trigger from the contract itself, never from a customer-set category or a reviewer's memory.Why: labels go stale the moment the underlying documents change, and nothing tells anyone when that happens.
Tune the cut line against a labeled set of real contracts, not a guess.Why: a threshold nobody checked ships confidently wrong on exactly the one contract that needed the careful read.
Watch the over-trigger cost: how often Reasoning Lane fires on contracts that turn out fine.Why: too loose and every contract waits three minutes, so reviewers start batching contracts at day's end instead of checking them as they land.
Watch the under-trigger cost: sample what Fast Lane clears and check it by hand sometimes.Why: too tight and a genuinely complex contract slips through with a false all-clear, the way Kestrion's nearly did.
Keep reasoning mode off simple lookups and formatting checks entirely.Why: a two-page NDA has nothing to chain through, so the extra deliberation just burns time and money.
How to answer this, stage by stage
Nobody is grading whether you can define a reasoning model. They are grading whether you can turn "it thinks longer" into a rule that decides, cheaply, when that's actually worth paying for.
1
Scope it to one company, one reviewer
Say it like this
"Let's ground this in one company. Stanwick Components runs its contracts through a tool called Sinew, and Bronwyn Merrion, their senior contract reviewer, is the one who actually has to trust or double check whatever Sinew tells her."
Why this works
Keeps the answer from turning into generic advice about "using AI wisely."
2
Name the method out loud
Say it like this
"I'll run this as SPARK. Situation, how review works today with no clear rule for the slow mode. Payoff, the habit I actually want. Anchor, the one design decision. Risk, what breaks if I get it wrong in either direction. Keep out, what I'll deliberately never send to reasoning mode."
Why this works
Two seconds of structure signals you have a plan, not just an opinion about expensive models.
3
Reframe what the question is actually testing
Say it like this
"This isn't really 'is the fancy model worth it.' It's whether we can tell, cheaply, before we pay for the slow read, which requests actually need it and which don't."
Why this works
Separates a real product decision from a vague "smarter model, better answers" take.
4
State precisely what a reasoning model does differently
Say it like this
"A standard model reads the contract once and answers once. A reasoning model reads it, works through the pieces that might connect, one at a time, indemnification, then the liability cap, then the insurance clause, checks those steps against each other, and only then answers. That's real extra time, close to nine times as long here, and real extra cost. On a request where nothing actually connects, all of that is wasted."
Why this works
This is the technical core of the question. Skip it and the rest of SPARK has nothing to hang on.
5
Give the anchor: the actual routing rule
Say it like this
"Before Sinew spends a second thinking, it runs a cheap scan: how many of the named risk clauses, indemnification, liability cap, insurance, termination, actually reference each other in this specific document. Two or more connected types, it goes to the slow careful read. Zero or one, it stays fast. The document decides. Not a category a customer's admin picked six months ago."
Why this works
This is the direct answer, said as a rule someone could actually build, not a description of one.
6
Prove it against the near miss
Say it like this
"Here's why this mattered. Stanwick's old setup had every contract filed as 'master supply agreement' running the fast read, because that category used to mean simple. A new vendor, Kestrion Forgeworks, sent over a much denser one, and it sailed through Fast Lane with nothing flagged. Two nights before signing, Bronwyn had a hunch and checked it by hand, and found the indemnification clause was actually capped at about forty thousand dollars by a totally separate section, against a recall that could run into millions."
Why this works
One real near miss does more work than any abstract argument about "smarter routing."
7
Close on the keep-out and the one line
Say it like this
"I'd never send a two-page NDA or a missing-signature-block check to reasoning mode, that's pure waste with nothing to chain through. So: read the document's own structure, not a stale label, and only pay for the slow careful pass when the clauses actually connect."
Why this works
Shows judgment about where the expensive mode doesn't belong, then restates the decision in one breath.
Let's learn
Every week, before Sinew existed at all, Stanwick's legal team read every incoming contract by hand. A simple mutual NDA took Bronwyn Merrion about ten minutes. A real master agreement, the kind where indemnification, a liability cap, and an insurance clause all lean on each other, took her close to three hours, cross-referencing section numbers on a legal pad as she went.
This is the work Sinew was built to take off her desk. Not the reading. The cross-referencing.
Keelson built Sinew with two speeds. Fast Lane reads a contract once and answers once, about 20 seconds, a few cents. Reasoning Lane reads it, works through the clauses that might connect to each other, checks those steps against each other, and only then answers, about three minutes, roughly ten times the cost. For a long stretch, nobody at Keelson or Stanwick could say which contracts actually got which read.
Knowledge spark: what does "cross-reference" mean in a contract?
A clause that leans on another clause by name, "subject to Section 14," "as defined in Section 3." Read alone, a clause can look completely ordinary. Read against the section it points to, the promise it makes can shrink, or disappear.
This is the actual technical difference the question is asking about. Not "smarter." Slower, on purpose, because it checks its own work.
Stanwick sends Sinew about 70 contracts a month. Roughly 56 are simple, standard NDAs and purchase orders under a set dollar value. Roughly 14 are structurally complex, master agreements with indemnification, liability caps, and insurance requirements that reference each other.
Average time to clear one contract, by routing policy
Cheap, but blind to structureSafe, but wastefulWhat the gate actually spends
Weighted across Stanwick's real mix, 56 simple contracts and 14 structurally complex ones a month, the gate averages about 52 seconds a contract, a little over a quarter of what running Reasoning Lane on everything would cost.
Then came the part that mattered. Reasoning Lane on every contract wasn't really the problem, and neither were the handful of times Fast Lane guessed wrong on something genuinely simple. Those were cheap mistakes. The real problem showed up the day a contract that badly needed the careful read got the fast one instead, and nothing on the screen said so.
The extra seconds and extra dollars were never the real cost. The real cost was a "no risk flagged" tag that meant something different depending on a setting nobody in the room could see.
What it costs at its worst: a reviewer who has learned to trust a clean tag stops doing her own cross-check. The one contract structurally complex enough to hide a real problem sails through with a false green light, and by the time anyone notices, it already has a signature on it. That's worse than never having built Sinew at all, because at least without it, Bronwyn would have read every page herself.
The choice I would take back
Keelson gave every customer a settings page: pick, per contract type, whether Sinew runs Fast or Reasoning by default. It made sense with a handful of early customers whose contracts really were simple and consistent. Stanwick's admin set "master supply agreement" to Fast Lane back when their master agreements really were short, boilerplate purchase orders in different clothes. Nobody ever revisited that label as the documents underneath it changed.
What I would leave alone: Stanwick's simple mutual NDAs and standard purchase orders, still the majority of what comes through Sinew, stay exactly as fast as they've always been. Nothing in them has a second clause to check against, so there's no reason to make Bronwyn wait three minutes for a read that a single pass already gets right.
The lesson: a setting is a promise about a document, made once. The document doesn't send an update when it stops keeping that promise.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one for why the rule had to change at all, and what nearly went out the door before it did.
Bronwyn Merrion can read a limitation-of-liability clause and tell you, in about ten seconds, whether it's been quietly gutted by a section three pages later that redefines what counts as "liability" in the first place. Nine years reviewing contracts will do that.
Before Sinew arrived, that instinct was her whole job. She read every incoming agreement herself, indemnification against liability cap against insurance, every time, because nobody else was going to.
Sinew launched at Stanwick running Reasoning Lane on everything, no toggle yet, every contract getting the slow, careful read. For six good months, that was fine. Flags came back reliable. Bronwyn started trusting a clean tag the way she'd trust a second lawyer who never got tired.
Then Keelson, fielding cost and speed complaints from bigger customers, shipped the settings page: a Fast or Reasoning default, per contract category, per account. Stanwick's operations admin, not Bronwyn, set "master supply agreement" to Fast, because ninety percent of the ones they'd signed that year really were short and standard. Nobody told legal.
The habit thinned in three small beats. First, Bronwyn noticed master agreements cleared faster than she remembered, and figured Sinew had simply gotten quicker. Second, she stopped opening the underlying document on anything Sinew cleared, the way she used to for the first six months, because it had been right every time. Third, "no risk flagged" stopped meaning "read carefully" to her at all. It just meant cleared, full stop, and she had no way to know which read had actually happened.
Same wrong contract, two different outcomes, depending on what decided the lane: a label, or the document itself.
Kestrion Forgeworks was Stanwick's first overseas precision-parts supplier, and its agreement arrived forty pages long, nearly four times Stanwick's usual template, filed under the same old category. Fast Lane. Cleared. Nothing flagged.
It was a Friday afternoon, signature due Monday. Something nagged at Bronwyn: the file felt thick for something Sinew had cleared so quickly. Out of a habit she hadn't used in months, she opened Section 9, the indemnification promise, and started checking it against everything it might quietly touch.
She spent that evening and came back Saturday morning, close to four hours in total, cross-referencing a forty-page agreement clause by clause, the exact work Sinew was supposed to have replaced, because she no longer trusted what "cleared" actually meant and had no way to check from the screen.
We didn't just lose a Saturday morning. We lost the one thing Sinew was built to give her back.
What she found: Section 9 promised Kestrion would cover Stanwick's costs if a defective part caused a recall. Section 14, read alone, looked like an ordinary limit, nobody pays more than the value of parts delivered in the last 90 days. Nobody had connected the two. Read together, the recall promise was worth at most about forty thousand dollars. A real recall on precision parts could run into millions. Section 22 required Kestrion to carry insurance, which sounded like a backstop, until you read that the policy only covered Kestrion's own mistakes, not a promise Kestrion made in a contract.
The old decision sat in a much smaller meeting, eighteen months earlier, when Keelson had five customers and was deciding how to let people control Reasoning Lane's cost as the company grew. Building an automatic classifier that read each document's own structure felt like a bigger engineering bet than the team was ready to make. A settings page felt clean, easy to demo, easy to explain in a sales call. Lysandra Novaczek, newer to the product team then, backed it. Every customer's contracts really were simple back then. The bet was reasonable, for exactly as long as that stayed true.
Run the same afternoon again, months later, after Lysandra's redesign shipped. Kestrion Forgeworks' contract lands in Sinew. Before Bronwyn even opens her inbox, the structural gate has already counted three connected risk clause types and routed it to Reasoning Lane. Three minutes after upload, the flag is sitting there: indemnification effectively capped by Section 14, verify against real exposure. Bronwyn checks it against the raw text herself, about ten minutes, confirms it, and flags it to procurement that Monday morning, days ahead of signature instead of hours before it. Her legal pad stays in the drawer.
What I'd tell myself, sitting in that first settings meeting: we built customers a knob because a knob was easy to explain in a demo. We should have built the document a voice, because the document is the only thing that actually knows what's in it.
SPARK, the five calls behind the reasoning switch
Not a script for sounding technical. SPARK is what forces you to name the one cheap check that decides, before Sinew spends a dollar thinking, whether this contract actually needs it.
SSituation. Who is this person, and how does the moment go today, with no clear rule?
Bronwyn Merrion reviews roughly 70 contracts a month for Stanwick Components through Sinew. Right now, whether a given contract gets the fast read or the careful one depends on a category label her admin set months ago, not on what's actually written in that document.
One person, one company, one real workload. Never a segment of "legal teams."
PPayoff. What habit do I want this to build?
I want every reviewer to trust a "no risk flagged" tag only when it was actually earned by a structural check, not because a category label usually means simple. The habit is the product. A calmer Monday morning is downstream of that habit, not the goal itself.
Name the thing they'll stop doing wrong. That's the payoff, not the routing rule's word count.
The whole page in one picture. A quick read for what doesn't connect. A slow, checked read for what does.
AAnchor. The one decision everything else hangs on.
Before Sinew answers, it runs a cheap structural scan: how many named risk clause types, indemnification, limitation of liability, insurance, termination, actually reference each other in this specific document. Two or more connected types routes to Reasoning Lane. Zero or one stays in Fast Lane. Fast Lane contracts also get a lightweight keyword scan as a second, cheap signal, so risky language with no formal cross-reference still has a way to escalate. Lysandra also looked at letting a reviewer click "run the careful read" herself on any contract she chose, and turned it down: that puts the judgment call right back on the one person who's already too busy to remember to click it on exactly the contract that needs it.
Concrete enough to argue with. This is the answer to the question.
The anchor, close up. Not a setting a customer picks. A count, taken fresh off the document, every single time.
RRisk. What breaks the first time it's wrong?
Set the cut line too loose and Fast Lane nearly disappears, every contract with any two connected clauses waits three minutes, and Bronwyn starts batching review to the end of the day instead of checking contracts as they land, the exact workaround the always-reasoning phase caused. Set it too tight and a contract like Kestrion's slips through with a false all-clear, the way it nearly did. Keelson accepts a real trade here: the gate deliberately runs Reasoning Lane on some contracts that turn out fine, a few wasted minutes, to avoid missing the one that isn't, because a missed cross-clause risk costs Stanwick far more than a few wasted minutes ever could. The cut line isn't a guess. It's tuned against roughly 200 real contracts Keelson had already labeled, complex ones a lawyer had actually found a cross-clause issue in, simple ones that were genuinely fine, and it's rechecked every quarter as customers bring in new vendors with contract templates the gate has never seen. That drift is the risk that outlasts launch day: a cut line tuned on today's templates quietly stops fitting next year's, the same way the old static toggle did.
Not "accuracy drops." What the reviewer does next, in either direction.
The flag has to show its work, which sections it chained through, or a reviewer is just trusting a new kind of label.
KKeep out. What I deliberately will not build into this rule.
Reasoning mode never runs on a two-page mutual NDA, a standard purchase order, or a plain formatting check, like a missing signature block or a defined term used inconsistently. None of those have a second clause to check against. Paying nine times the wait and ten times the cost for a careful read with nothing to actually chain through is pure waste.
Shows judgment instead of a wish to be careful everywhere. Ties straight back to Risk: the wrong kind of caution is its own kind of wrong answer.
The recap, one line per letter: situation is a reviewer whose "cleared" tag stopped meaning one consistent thing, payoff is trust that's actually earned by a structural check instead of a category label, anchor is the cross-reference gate itself, tuned against real labeled contracts rather than a launch-day guess, risk is either extreme costing Stanwick on a different day, and keep out draws the line at what has nothing to chain through, never at "when in doubt, run it anyway."
And if you want to be sure it really works, try it somewhere else
Same five letters, a hospital's overnight reading room instead of a legal team, and this time the missed connection isn't a liability cap. It's a nodule that grew between two scans.
Solmark builds Nightread, an AI tool that pre-reads radiology reports overnight for Kellenbeck Health, a hospital network, and flags reads that need a second look before a radiologist signs off in the morning. Thoraya Duskwright runs the overnight reading queue, roughly 90 reports a shift.
Same shaped gate, a report instead of a contract. Count what actually connects, then decide the lane.
Mapped onto SPARK: the situation is Thoraya's overnight queue, most reports routine, a few genuinely dense, a finding's size mentioned in one section, a prior scan's measurement mentioned in another. The payoff is every radiologist trusting a "no second look needed" tag only when it's earned. The anchor is the same shaped gate, count how many findings a report correlates against another section or a prior scan, route those to Reasoning Lane so the model actually checks whether the numbers agree, otherwise Fast Lane. The risk is the same two-sided cost: over-trigger and radiologists lose a night to routine studies waiting on a slow read; under-trigger and a nodule that quietly grew gets the all-clear. Keep out draws the same line: a routine single-finding chest film, or a check for a missing required field, never needs the slow pass.
Reads Fast Lane cleared, where a radiologist's own double check caught a real missed correlation, per month, Kellenbeck Health
Building, while the label decided the laneAfter the structural gate shipped
Zero missed correlations in month one, climbing to three by month five as report templates got denser. Solmark's structural gate shipped between month five and month six. None caught by luck since, because none needed to be.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, say the rule, and what it protects against.
Cost: no budget for a new classifier this quarter. Have the hospital's existing report software run the same structural count as a plain rule, and gate the existing reasoning call behind it, no new model spend.
The model got better, for real: say Nightread's single-pass accuracy doubles overnight. The gate still matters, because a better single-pass model still only reads a report once. It can still miss a correlation across two sections it never re-reads.
Where people run it wrong.
They route every long report to reasoning mode because more pages feels riskier, instead of counting what actually cross-references.
They tune the cut line once at launch and never recheck it as report templates change hospital to hospital.
They let a busy technologist override the gate for speed, which quietly recreates the exact 2am judgment call the rule was built to remove.
How to use it live. Before answering a "what does the expensive mode do, and when's it worth it" question cold, ask yourself one thing: what's the cheap signal about this specific request, checkable before you spend a dollar, that tells you whether it actually needs multiple steps. Naming that signal first is usually exactly what the question is listening for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to design one rule for when an expensive mode is worth its cost?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Runs forward from a real gap in the room instead of working backward from a failure.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Bronwyn Merrion, Stanwick Components' senior contract reviewer. Lysandra Novaczek, the AI PM at Keelson who owns Sinew's routing rule. Kestrion Forgeworks, the new vendor whose contract almost went out with a hidden risk.
3 · THE PAYOFF
What habit does the routing rule exist to build?
Tap to flip
ANSWER
Every reviewer trusts a "no risk flagged" tag only when it was actually earned by a structural check, not because a category label usually means simple.
4 · THE ANCHOR
What's the one concrete decision Lysandra owns in this answer?
Tap to flip
ANSWER
A cheap structural scan that counts how many risk clause types, indemnification, liability cap, insurance, termination, reference each other in this specific document, and routes to Reasoning Lane only above that cut line.
5 · THE OLD DECISION
What decision would Lysandra take back?
Tap to flip
ANSWER
The static per-category settings page. It made sense with five early customers whose contracts really were simple. It stopped making sense once contract templates started varying, and nobody ever revisited the label.
6 · THE NUMBER
Fill in the blank: Fast Lane reads a contract in about ___ seconds. Reasoning Lane takes about ___.
Tap to flip
ANSWER
20 seconds, and about 180 seconds, roughly nine times as long, for about ten times the cost.
7 · THE RISK, SURVIVED
What breaks if the cut line is wrong in either direction, and how does the anchor survive it?
Tap to flip
ANSWER
Too loose and reviewers start batching contracts instead of checking them as they land. Too tight and a real risk slips through with a false all-clear. It survives because the cut line is tuned against roughly 200 labeled real contracts and rechecked every quarter, not set once and forgotten.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs SPARK again on a different product. Which one, and what's the equivalent anchor?
Tap to flip
ANSWER
Nightread, Solmark's overnight radiology reading tool for Kellenbeck Health. The equivalent anchor counts correlated findings across report sections instead of cross-referenced contract clauses.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: about ___ of Stanwick's 70 monthly contracts are structurally complex enough that Sinew should route them to Reasoning Lane.
Show hint
Check the numbers under "Let's learn," right before the time-by-policy chart.
Show answer
14. About 20 percent. The other 56 are simple NDAs and standard purchase orders that never needed the slow read in the first place.
Multiple choice
2. Why did Fast Lane clear Kestrion Forgeworks' contract with nothing flagged, even though the risk in it was real?
A. Because Fast Lane doesn't check indemnification clauses at all.
B. Because the contract was filed under a category label set to Fast Lane months earlier, not because of anything in this specific document.
C. Because reasoning mode was down that week for maintenance.
D. Because reasoning mode is only used for contracts written in a language other than English.
Show hint
Look at what actually decided the lane before Lysandra's redesign, in the key point box titled "The choice I would take back."
Show answer
B. The old design let a customer-set category decide the lane, not the document's own structure. Kestrion's agreement was structurally denser than Stanwick's usual template, but it still carried the same old label.
True or false
3. True or false: the real fix here was training Bronwyn to double check every complex contract by hand from now on, just in case.
True
False
Show hint
Think about what Sinew was supposed to take off her desk in the first place.
Show answer
False. That just rebuilds the exact workload Sinew was meant to remove, and only works as long as one person keeps remembering to distrust a label. The real fix has to live in what triggers Reasoning Lane, not in a habit a reviewer has to maintain forever.
Short answer, name the reversal
4. What old decision would Lysandra take back, and why did it make sense when Keelson first made it?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Keelson's settings page let a customer set Fast or Reasoning as the default per contract category. It made sense with five early customers whose contracts really were simple and consistent. It stopped making sense once contract templates started varying customer to customer, and nobody ever went back to update the label as the documents underneath it changed.
Short answer, apply it yourself
5. Think of an AI feature you've used that seems to have a "fast" mode and a "careful" mode. What's one cheap signal about the actual request that should decide which one runs, instead of a setting someone picked once?
Show hint
Look for something you could check about the request itself in under a second, before spending on the slow mode.
Show answer
Model answer: A coding assistant's quick autocomplete versus a full multi-file refactor suggestion. The cheap signal could be how many files or functions the proposed change actually touches, not a manual toggle a developer has to remember to flip on the one change that needed it.
Short answer, work the number
6. If Stanwick's mix shifted from 20 percent structurally complex contracts to 40 percent, would the structural gate still save meaningful time over running Reasoning Lane on everything? Show the math.
Show hint
Use the same weighted average as the time-by-policy chart, just with a different split.
Show answer
Model answer: Yes, but the margin shrinks. At 40 percent complex, the weighted average becomes (0.6 × 20s) + (0.4 × 180s) = 12 + 72 = 84 seconds, still 53 percent faster than always running Reasoning Lane at 180 seconds. The gate keeps paying off, it just pays off less as more of the actual work genuinely needs the slow read.
Before you close the answer
Why this works
Tests whether you treat "should we use the expensive model" as a per-request routing decision with a cheap, checkable trigger, not a blanket policy or a shrug of "use your judgment." Most candidates can describe what a reasoning model does. Fewer can say, precisely, what decides whether to pay for it on this exact request.
Follow-up traps
"What about risk written in plain prose, with no explicit section cross-references at all?" Response: the structural count alone would miss that, which is exactly why Fast Lane contracts also get a lightweight keyword scan as a second, cheap signal, and anything that hits there escalates to Reasoning Lane too.
"Wasn't a static toggle simpler for customers to understand and trust than an automatic classifier deciding for them?" Response: simpler to explain isn't the same as accurate, and the cost of a wrong static label is silent. The automatic gate's decision is visible in the flag itself, which sections it chained through, so a reviewer can check its reasoning instead of just trusting a setting.
If pressed
A reasoning model can chain through clauses that don't actually connect and flag a risk that isn't real. Sinew's guardrail against that: every Reasoning Lane flag has to cite the literal section numbers it chained through, not just assert a conclusion, so a reviewer can verify the chain in under a minute instead of re-deriving it from scratch.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.