ConceptAdvancedModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #17
How do you keep expectations calibrated when the underlying models genuinely improve every quarter?
BOUND · pricing what a vendor's record quarter actually buys Fathomline, an AI tool that reads earnings filings and footnotes for investors
Fathomline reads a company's earnings filing, footnotes included, and drafts a confidence-tagged research brief an investor can act on. Emyr Godalming owns its model. Halbert Runcorn, who has never opened a training log, runs client success and loves a good screenshot. Cobalt Reef Capital, a hedge fund, runs its due diligence through Fathomline, and Florentyna Aspinwall, the fund's director of research, is the one who caught a wrong number half an hour before her desk would have published it.
The direct answer
Don't let a vendor's benchmark headline become your own claim. Re-run your product's own eval set, especially its hardest real cases, against every meaningful new model release before you update anything you tell yourself or a client about how much better it got. Report the number you actually measured. Some quarters that number is a real leap. Some quarters it barely moves. Both are the honest answer, and neither one is the vendor's number.
Do this, in order
Re-run your own eval set on every meaningful model release before updating any claim.Why: this is the one decision the whole answer is built to defend.
Own the real number and where it came from, not the vendor's blended score.Why: a number with a source survives a follow-up question, a feeling does not.
Break "the model got better" into the two different things it bundles before you believe either one.Why: a general benchmark and your hardest real cases are rarely testing the same skill.
Give a range across quarters instead of a fixed expected pace.Why: some releases are a big local win, some are almost nothing, and a "minor" one can beat a "record" one.
Sanity check any improvement claim against your hardest cases before it leaves the building.Why: this is the exact check that would have stopped a wrong number reaching a client.
Split "model quality" by case difficulty, not one blended dashboard line.Why: a blended number hides the one slice the people relying on it actually care about.
How to answer this, stage by stage
Nobody is grading whether you can say "the model got smarter." They are grading whether you can turn that into a number you would actually stand behind on a client call.
01
Scope it to one product, one promise, one quarter
Say it like this
"Let's ground this. Fathomline reads earnings filings for investors, and our head of client success just told a hedge fund the model 'catches every footnote inconsistency now' because the vendor posted a record quarter on their own benchmark. That's the specific claim I'm going to check, not model improvement in general."
Why this works
Pins an abstract question to one real, checkable claim instead of a lecture on model quality.
02
Reframe what the question is actually testing
Say it like this
"The real question isn't whether the model got better. It almost always does, a little. It's whether it got better at the exact thing our hardest cases need, and a general benchmark was never built to answer that."
Why this works
Separates a true fact, the model improved, from the actual claim being made, it improved here.
03
Break down what "improved" is bundling, the B step
Say it like this
"A financial-reasoning benchmark is mostly multi-step arithmetic and clean, structured questions. Our hardest cases are something else entirely, a contingent liability named once, in a footnote, in dense legal language. Those don't share a skill, so a benchmark jump doesn't promise anything about the second one."
Why this works
Gives the interviewer a testable reason for doubt, not a vague sense of caution.
04
Own the real numbers behind one specific quarter, the O step
Say it like this
"That quarter, the vendor's benchmark went from 84 to 91, seven points, a record. Our own eval set, the one built from our hardest real filings, went from 79 to 80.4. One and a half points. Same model, same quarter, two very different numbers."
Why this works
A number with its source survives a follow-up question. "It got better" does not.
05
Show the honest range across quarters, the U step
Say it like this
"The quarter before, a ten-point benchmark jump matched an eleven-point jump on our own eval, almost identical. The quarter after the record one, the vendor called it a minor update, half a point on their benchmark, and our own eval jumped six and a half points, because that release happened to target long, messy text. There's no fixed pace here. Some quarters match, some don't, in both directions."
Why this works
Turns "it depends" into something checkable against real history, not a shrug.
06
Run the sanity check that would have caught the bad claim, the N step
Say it like this
"If a seven-point jump on their benchmark only bought us one and a half points on our own hardest cases, then telling a client we 'catch everything now' isn't rounding up. It's a different, unverified number."
Why this works
This is the exact check that would have stopped the claim before it reached the client.
07
Name the direction, reject an alternative, and close on the real rule, the D step
Say it like this
"The single fact that swings this most isn't how big the benchmark jump is. It's how much overlap there is between what the release actually improved and what our hardest cases need. So here's the rule now: every real model release gets run against our own footnote eval before anyone hears a new number, inside the company or out. We looked at waiting for the vendor to publish their own task breakdown instead, and dropped it, because they never score footnote reading as its own line. We have to test it ourselves."
Why this works
Ends on an operating rule the interviewer can picture actually running, not a promise to be more careful.
Let's learn
What happens the quarter a model genuinely gets smarter, and your hardest cases don't notice?
Say we build a tool that reads a company's earnings filing, all of it, footnotes included, and drafts a research brief with the risky lines marked for a person to check. Before a tool like this, an investor's own analyst spent about three hours combing one filing by hand, hunting for the sentence buried on page forty about a lawsuit the company hasn't lost yet. With the tool running well, that same filing comes back as a four-minute draft, and the analyst spends about twenty minutes checking the handful of lines the tool itself flagged as unsure.
Five boxes. The fourth one, flagging what the model is unsure about, is the whole reason a bad quarter is survivable at all.
Then a new model version ships. The company behind the model announces a record quarter, their own benchmark, a broad test of financial reasoning, up seven points in one release. Word spreads inside Fathomline fast. Someone updates the pitch deck. Someone tells a client the tool "catches everything now."
Knowledge spark: what is a financial-reasoning benchmark, really?
A fixed set of test questions a model vendor publishes a score against, usually a mix of arithmetic, ratio problems, and clean, well-formed questions about a filing. It is built to be scored the same way for every model, which means it rarely tests messy, one-off, real-world text the way any single product actually uses it.
Here is the turn. That seven-point jump is real. It is also almost entirely about a different skill than the one the tool is actually weak at. On Fathomline's own hardest real cases, the ones built from messy, footnote-shaped text, the number moved by a point and a half. Nobody checked that before the claim went out, because checking felt like second-guessing good news.
The seven points were real. They just were not the seven points our hardest cases needed.
What it costs at its worst: a research brief goes out with a wrong number sitting in a footnote-derived line, stated in the same confident tone as every correct line around it. Nothing on the page warns anyone. A client's own analyst catches it, this time, by luck and by reading closely, right after being told the tool had just gotten dramatically better at exactly that kind of mistake.
The choice I would take back
Early on, the team let Fathomline's internal "how good is our model" number just mirror the vendor's own published score. That was a fair call while the two moved together, because back then most of what was wrong was common and broad, and a general benchmark caught it fine. It stopped being fair the moment the common mistakes got fixed and what was left were the narrow, footnote-shaped ones a general benchmark barely touches.
What I would leave alone: plain earnings-call summaries, no footnotes, no legal language, still track the vendor's benchmark closely, because that is common, well-covered text and a broad benchmark actually does test it. Gating every claim about that behind a full re-test would be caution with no real risk behind it.
The lesson: a number that moved for a real reason can still be the wrong number to hand someone. The question was never whether the model got better. It was whether it got better at the one thing we are actually asked to be right about.
Now here is the same thing as a story
Use the short version above when you are speaking. Read this one for the Wednesday in July a wrong number almost went out under Fathomline's name.
Emyr Godalming has owned Fathomline's model since before it read its first real filing. He reads an eval report the way an accountant reads a ledger, catching a decimal moving the wrong way before anyone else notices.
Halbert Runcorn runs client success, and he has never opened a training log in his life. What he is good at is a room. For most of the year, the dashboard on his second monitor did half his job for him.
January was launch, the vendor's benchmark at 71, Fathomline's own footnote eval close behind at 68. By April the vendor's benchmark had climbed to 84 and Fathomline's own eval had climbed almost exactly as far, to 79. Halbert started every pitch that spring with the same screenshot: two lines climbing together.
The habit thinned in three steps. In April, on a renewal call, he said "the model's getting sharper every quarter," true, and he had checked the rough shape of it with Emyr the week before. By June, pitching a new fund, he said "we're closing in on catching almost everything," an extrapolation now, not a fact, though it still happened to land close to true. By July, prepping a deck for Cobalt Reef Capital, a hedge fund weighing whether to renew, he typed "catches every footnote inconsistency now" straight onto a slide, and this time he never asked Emyr a single question about it.
The vendor had just posted what they called a record quarter, their financial-reasoning benchmark up seven points in one release, 84 to 91. Halbert's slide, and his confidence, rode straight off that number.
Two shapes of the same idea. Only one of them describes what Fathomline's own numbers actually did.
He did not know Fathomline's own footnote eval, the one built from the hardest real filings the tool sees, had moved from 79 to 80.4. One and a half points. The vendor's release had gotten a lot better at multi-step math and clean, structured questions. It had barely touched the narrow, footnote-shaped mess Fathomline actually struggles with.
Five separate reasons a filing still trips the tool up. None of them share a fix, and none of them are what a general benchmark was built to catch.
Then, on a Wednesday afternoon, Florentyna Aspinwall, Cobalt Reef's director of research, was reviewing a Fathomline-drafted brief before her desk published a note off it. One line named a contingent liability at a figure that did not match the filing. It read exactly as confident as every correct line around it, nothing about it looked unsure. She caught it, this time, only because she happened to check that specific paragraph against the source herself.
She called Fathomline's account team, not angry, just unsettled. Halbert's follow-up call with Cobalt Reef, the one built around "catches everything now," was booked for 5pm the same day.
Three things landing the same afternoon. Only one of them was actually a coincidence.
Emyr heard about the near miss at 1pm and spent the afternoon building the one chart the team should have had running automatically for months: the vendor's benchmark against Fathomline's own footnote eval, quarter over quarter, side by side. By 4:40 he was in front of Halbert with it.
Halbert remembered the January meeting where they had decided Fathomline's internal quality number would just track the vendor's own published score. "Nobody in that room was cutting a corner," Emyr told him. "The two numbers really did move together, back then." They stopped moving together the day the common mistakes got fixed and only the narrow ones were left, and nobody had gone back to check.
At 5pm, Halbert got on the call. He did not say "catches everything now." He told Florentyna the real numbers: the record quarter, the one and a half points it actually bought on the cases that mattered to her, and the near miss her own team had just caught. Then he told her what changed as of that afternoon, every new model release gets run against Fathomline's own footnote eval before any claim about it goes out, inside the company or to a client.
Cobalt Reef renewed two weeks later, at a lower initial commitment than Halbert had pitched in June, tied to hitting a real, tested number on the footnote eval by October, not a vendor's press release. Fathomline hit it in October, when what the vendor quietly called a minor update turned out to move the footnote eval by six and a half points in three weeks, because that release happened to target exactly the kind of long, messy text Fathomline had been stuck on.
What Emyr would tell his January self: he built that quality number to make an internal dashboard look finished. He never once asked what would happen the day someone read a promise straight off it to a client who would trust it.
BOUND, for what a record quarter actually buys the footnotes
Not a way to make a good number sound suspicious. BOUND turns "the model got smarter" into a cost and a range Halbert, or anyone else at Fathomline, could actually stand behind on a client call.
BBreak it down. What is "the model got better" actually bundling?
A vendor's general financial-reasoning benchmark is mostly multi-step arithmetic and clean, well-formed questions about a filing. Fathomline's own hardest real cases are a different animal entirely, a contingent liability named once, in a footnote, in dense, jurisdiction-specific language. Those are two different skills sharing one headline number.
Skip this split and "the model got smarter" sounds like one fact, when it is really a question about which specific skill actually moved.
OOwn the numbers. Where does each one actually come from?
In April, the vendor's benchmark rose from 71 to 84, thirteen points, and Fathomline's own footnote eval rose from 68 to 79, eleven points, close enough to trust the vendor's number as a rough proxy. In July, the vendor's "record quarter" release moved their benchmark from 84 to 91, seven points, and Fathomline's own eval moved from 79 to 80.4, one and a half points.
A number only counts as owned if you can say exactly where it came from when someone pushes on it. "It's smarter now" is a feeling wearing a number's clothes.
Vendor's benchmark versus Fathomline's own footnote eval, by quarter
Vendor's own benchmarkFathomline's own footnote eval
The two lines move together in April. They split hard in July, and by October the "minor" release has quietly closed most of the gap the "record" one opened.
UUse a range, not one number.
A routine earnings-call summary, plain English, no footnotes, sits on the part of this curve where the vendor's benchmark and Fathomline's own number still track closely. A filing with a contingent liability buried three footnotes deep sits on the part where they do not. What decides which part you are on is not the calendar, it is whether that release happened to target long, messy text.
A single number here repeats Halbert's exact mistake, sounding certain about something that genuinely depends on which release, and which kind of case, is being asked about.
The vendor's benchmark barely tests the skill Fathomline's hardest cases need. Footnote-style extraction is 15 percent of the test, not 45.
NNail the sanity check. Does the claim survive contact with the real numbers?
If seven benchmark points bought one and a half points on the cases that actually matter to a client, then "catches everything now" is not rounding up, it is a different, unverified claim. The right question was never "did the vendor's number go up." It was "did our number go up, on the cases we're actually being trusted with."
This is the exact check that would have stopped Halbert's line before it reached Cobalt Reef.
What would move this estimate the most, if it were wrong
Not the size of the vendor's headline. Whether the release actually targeted the skill Fathomline's hardest cases need is the fact that moves this estimate the most.
DDirection. Which assumption moves this most, and what's the actual decision?
Not how big the vendor's headline number is. The single biggest swing factor is how much overlap exists between what a release actually improved and what Fathomline's hardest real cases need. So the real decision is a standing rule, not a one-time promise: every new model release gets run against the footnote eval before any internal or client-facing claim goes out, and the number reported is the tested one, never the vendor's.
Naming the fact that actually swings the estimate, instead of the biggest number in it, is what separates a real estimator from a confident guesser.
Four inputs, not one. The near miss scored borderline-high on the first input alone, which is why the fourth one got added.
One alternative considered and rejected: wait for the vendor to publish a task-by-task breakdown of their own benchmark instead of maintaining an internal eval set. It lost, because vendors bundle messy footnote extraction inside a generic "document question answering" score, they never break it out on its own, so that transparency was never actually coming. The AI-specific risk sitting under all of this is quiet, a model can misstate a contingent liability figure from a dense footnote fluently, in the exact same tone as a correct line, with nothing on the page to warn a tired reader. The guardrail is the confidence tag, tied to a mandatory human check on anything flagged low-confidence, footnote-dense, or belonging to a case type that failed last quarter, before a brief ships. And the trade-off is real, not free, re-testing every release costs about a day of compute and an analyst's afternoon, which delays how fast Fathomline can publicly say "we got smarter" by a day or two, against shipping a claim the same day the vendor announces.
And if you want to be sure it really works, try it somewhere else
Same five letters, a building-permit checker instead of a filing reader, and this time the real question is not whether to trust the number, it is whether checking it is worth the delay at all.
Codewell reads a building-permit application against local code and flags likely violations for a human reviewer, so a small permits office does not have to read every page of every application by hand. Coretta Birtwhistle is Codewell's compliance lead at the Cinder Falls permits office, and she has never once let a vendor's own number stand in for her own test.
Run BOUND on it. Break it down: a general code-compliance benchmark mostly tests common, well-documented rules, setbacks, height limits, straightforward zoning. Cinder Falls' hardest cases are historic-building exceptions, hand-amended decades ago, written in language specific to one district and nowhere else. Own the numbers: the vendor's general code benchmark jumped from 82 to 91, nine points, in one release. Codewell's own eval on historic-exception clauses moved from 74 to 75.2, a bit over one point.
A different city, a different code book, the exact same shape of gap between a general benchmark and the narrow cases a permits office is actually trusted to catch.
Where Codewell's answer genuinely differs
Coretta's team never once mirrored a vendor's number, a permit error has legal weight for a property owner, so they built the habit of testing locally from day one. Their real question was different: whether it is worth delaying a public "we upgraded" announcement by two weeks to run the full historic-exception eval, against announcing at general release and running the eval in parallel, accepting a short window where the claim is unverified.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it, name the two different things "improved" bundles, own one real number, then say plainly that a range beats a single guess.
Cost: no budget or time to run a full re-test before a promise. Shrink the eval to the handful of cases that actually decide the client's renewal, and say honestly that it is a smaller sample.
The model got better, for real: say the new release's benchmark and Fathomline's own eval both jump by a lot in the same quarter. Still check by case type, because "both numbers went up" can still hide one narrow case type that did not move at all.
Where people run it wrong.
They treat the vendor's own benchmark score as if it were their own product's score.
They average across every case type instead of checking the hardest ones specifically, so a big win on the easy majority hides an untouched hard minority.
They assume the size of a benchmark jump predicts the size of the real jump, in either direction, when a "minor" release can matter more than a "record" one.
How to use it live. Before answering, ask yourself one plain question out loud: "did this release actually target the skill our hardest cases need, or just get bigger at what was already easy." Whichever it is, that answer decides whether the headline number means anything here.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What method fits deciding whether a vendor's quarterly model jump actually improved your product?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, name the direction. Built for turning a benchmark headline into a number you can actually defend.
2 · THE CAST
Who is this answer about?
Tap to flip
ANSWER
Emyr Godalming, Fathomline's model PM. Halbert Runcorn, Fathomline's head of client success. Florentyna Aspinwall, director of research at Cobalt Reef Capital, the client fund.
3 · THE BREAKDOWN
What does "the model got better" actually bundle together, and why doesn't one number cover it?
Tap to flip
ANSWER
A general financial-reasoning benchmark mostly tests multi-step math and clean structured questions. Fathomline's hardest real cases are messy footnote disclosures. Different skills, so a jump in one says little about the other.
4 · THE OWNED NUMBER
Fill in the blank: the vendor's "record quarter" release moved their own benchmark from 84 to ___. It moved Fathomline's own footnote eval from 79 to ___.
Tap to flip
ANSWER
91, up 7 points. 80.4, up 1.4 points. Same release, same quarter, two very different real numbers.
5 · THE RANGE
What decided whether a release moved Fathomline's real numbers a lot or almost nothing?
Tap to flip
ANSWER
Whether that specific release targeted the skill the hardest cases actually needed, long messy text extraction, not how big the release's own headline number was. A "record" release barely helped; a "minor" one helped a lot.
6 · THE SANITY CHECK
Why was Halbert's claim to Cobalt Reef wrong?
Tap to flip
ANSWER
He assumed a seven-point benchmark jump meant the tool now caught everything. It had bought one and a half points on the cases that actually mattered to his client, not a transformation.
7 · THE DIRECTION
Which single fact swings this estimate most, and what's the real rule now?
Tap to flip
ANSWER
How much overlap exists between what a release actually improved and what the hardest real cases need, not the size of the headline number. The rule: re-test the own eval set before any claim goes out, every time.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's the same shape?
Tap to flip
ANSWER
Codewell, a building-permit compliance checker for Cinder Falls. Same shape: a big general code-benchmark jump barely moved its own eval on historic-building exception clauses, the narrow cases a broad benchmark does not test.
Check yourself Score: 0 / 0
Fill in the blank
1. The quarter the vendor called a "minor update," their own benchmark moved by only ___ points, but Fathomline's own footnote eval moved by ___ points.
Show hint
Check the U step and the October marks on the line chart.
Show answer
0.5 points (91 to 91.5). 6.6 points (80.4 to 87). That release happened to target long, messy text extraction, exactly the skill the footnote eval needs, even though the vendor's own headline barely moved.
Multiple choice
2. Why did the vendor's seven-point "record quarter" benchmark jump only move Fathomline's own footnote eval by one and a half points?
A. The eval set was too small to detect the real improvement.
B. The benchmark's improvement was mostly in multi-step math and structured questions, a different skill than reading messy footnotes.
C. Fathomline's engineers had not finished integrating the new model yet.
D. Cobalt Reef's filings are unusually difficult compared to the industry average.
Show hint
Check the B step and the stacked bar chart of what the benchmark actually tests.
Show answer
B. Footnote-style extraction is a small slice of what the benchmark measures, so a broad jump on the benchmark does not have to touch it at all.
True or false
3. True or false: once you have seen a vendor's benchmark score and your own product's eval score move together for two quarters in a row, you can trust that relationship to hold for future releases too.
True
False
Show hint
Check the N step and what happened between April and July in the story.
Show answer
False. The two moved together while the common, broad mistakes were still being fixed. Once those were fixed, what was left was narrow enough that the general benchmark stopped predicting it, without warning.
Short answer, where it would not matter
4. Name a part of Fathomline where trusting the vendor's benchmark jump, without re-testing locally first, genuinely would not be a mistake.
Show hint
Look at "what I would leave alone" in Let's learn.
Show answer
Model answer: Plain earnings-call summaries with no footnotes or legal language. That is common, well-represented text, and the general benchmark actually does test it well, so a jump there reliably shows up in Fathomline's own numbers too.
Short answer, apply it yourself
5. Think of a tool you use that runs on some underlying model or engine that gets updated regularly. What is one specific, hard thing you rely on it for that a general "it got smarter" announcement might not actually have touched?
Show hint
Think about the gap between a broad headline update and your own narrow, specific use of the tool.
Show answer
Model answer: A grammar checker announces a big general writing-quality update. The thing you actually rely on it for might be catching a very specific mix-up in your own writing, and a broad update aimed at tone and fluency might never touch that one narrow pattern.
Short answer, work the number
6. If Cobalt Reef had insisted Fathomline hit 95 on its own footnote eval within one more quarter no matter what, what would the real quarterly numbers suggest about whether that is realistic?
Show hint
Compare the best quarter's real jump against the gap left to close.
Show answer
Model answer: The best quarter on record moved the footnote eval by 6.6 points, and that only happened because the release specifically targeted long-text extraction, not because of a deadline. Going from 88.1 to 95 needs another 6.9 points, roughly another best-case quarter, and only if the next release happens to target the same skill. A flat promise ignores that the size of the jump depends on what the release actually does, not on the calendar.
Before you close the answer
Why this works
Tests whether a candidate will trust a vendor's own good news at face value or go check what it actually means for their specific hardest cases. Most candidates stop at the announcement.
Follow-up traps
"Isn't re-testing every release just slowing you down for no reason?" Response: it costs about a day of compute and an analyst's afternoon, against shipping a claim that turns out untrue to a client who then stops trusting every number after it. That trade is worth a day.
"What if the general benchmark and your own eval always end up close anyway?" Response: then re-testing costs almost nothing extra and confirms the claim is safe to make. The cost only shows up the quarter they diverge, and you cannot know in advance which quarter that will be.
If pressed
The line that almost went out wrong actually scored borderline-high confidence on the model's own token uncertainty alone. That is why the confidence tag also checks whether that exact case type failed the eval last quarter, not token confidence by itself, before deciding what needs a mandatory human check.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.