ConceptIntermediateAI Opportunity & Model Strategy / Model selection from a PM lens / #2
Why is benchmark leaderboard position a weak input to model selection?
TRACEthe leaderboard measured a practice essay, Ashcombe's classrooms write the real ones
Why does a model's rank on a public leaderboard tell you so little about whether it will work for you? Ashcombe Learning builds MarginNote, a tool that drafts written feedback on a student's essay for a teacher to review and send. Ingrid Kallevik picked the essay-scoring model behind it the week school started, and by the second grading period, something about it had quietly stopped being true.
The direct answer
Leaderboard rank measures agreement with human graders on the benchmark's own essay set, not on your own students' writing. A model can sit at rank one and still fail badly on the exact population you serve, and your own dashboard can hide it, because people quietly adapt what they feed the model once they notice it struggles. Test any finalist on your own raw, unedited data before trusting where it sits on someone else's leaderboard.
Do this, in order
Test the model on your own raw, unedited data before trusting its leaderboard rank.Why: rank measures fit to the benchmark's essay set, not to the writing your own students actually turn in.
Check what population the benchmark's own essays came from.Why: a leaderboard corpus that skews one kind of writing can hide exactly where your students differ from it.
Recut your own quality numbers by segment, not just the overall average.Why: an average sitting at 81 percent can still mean one group is being failed badly.
Watch for people quietly changing what they feed the model once it starts struggling.Why: a cleaned-up input makes your own dashboard look fine while the real failure keeps happening underneath it.
Re-test on a held-out, untouched sample every term, not just at launch.Why: a benchmark rank is a snapshot, and your own student population and their writing keep changing.
How to answer this, stage by stage
Nobody is scoring whether you can define "leaderboard." They're scoring whether you can explain, concretely, how a rank-one model quietly fails a real classroom.
Stage 1
Scope it to one real tool
Say it like this
"Let's ground this in a real product. Ashcombe Learning's MarginNote drafts written feedback on a student's essay. I'll walk through exactly why the model behind it topping a public leaderboard wasn't the same thing as it working for Ashcombe's actual students."
Why this works
Keeps the answer from becoming an abstract lecture about benchmarks in general.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as TRACE. Timeline, when it actually started going wrong versus when anyone noticed. Recut, sliced by segment instead of one average. Assume nothing, rule out a tracking bug first. Cause candidates, three real hypotheses. Evidence test, the one check that tells them apart."
Why this works
Signals a repeatable way to diagnose a gap between a benchmark number and real quality, not a one-off guess.
Stage 3
Reframe: it isn't "is the leaderboard wrong," it's "whose writing was it ever measuring"
Say it like this
"This isn't really about whether the leaderboard is trustworthy. It's about noticing that a leaderboard measures agreement with graders on one specific set of essays, and asking whether that set looks anything like the essays your own students actually write."
Why this works
This is where a strong answer separates from a candidate who just says "benchmarks can be gamed" and stops.
Stage 4
Give the one decision
Say it like this
"Here's what I'd actually do: pull 200 raw, untouched essays straight from the classroom app, split them by whether the student is a multilingual learner, and run the model cold against both groups before trusting anything the vendor's leaderboard says."
Why this works
This is the direct answer, stated as an actual test you could run this week, not a philosophy about benchmarks.
Stage 5
Prove it with the compressed evidence
Say it like this
"MarginNote's model ranked first of 14 systems on a public essay-scoring leaderboard, 92 percent agreement with human graders. Ashcombe's district is 38 percent multilingual learners. A curriculum specialist pulled ten random transcripts and found six had been rewritten by the teacher before ever reaching the model. On 200 raw, untouched essays, native-English writing scored 84 percent helpful. Multilingual-influenced writing scored 29."
Why this works
Compresses the whole case into the one audit that separated the leaderboard's number from the real one.
Stage 6
Name the AI-specific reasoning, the trade-off, and close
Say it like this
"The honest reason this isn't a generic quality problem is that the model was never confidently wrong in a way anyone could see; it just answered fluently on writing patterns it had never been shown, and teachers, without meaning to, hid that from us by cleaning up input before we could measure it raw. We accepted a slower rollout, three extra weeks of testing on our own raw data, to catch that before it hit report cards district-wide."
Why this works
This is the load-bearing, AI-specific judgment. A normal feature doesn't have a headline benchmark number that quietly disagrees with the real population using it.
Let's learn
What happens the first time a leaderboard-topping model meets writing it was never actually tested on?
MarginNote reads a student's submitted essay and drafts written feedback for a teacher to review before sending it back. Before MarginNote, an English teacher at Ashcombe's partner schools spent about 18 minutes writing real comments on one essay, and a full class set of 28 essays took most of a weekend, roughly 8 and a half hours, every single time.
With MarginNote, a full draft of feedback appears in under a minute, and teachers spend their time reviewing and adjusting it instead of writing from scratch.
Same model. Two numbers. Only one of them was ever measured on Ashcombe's own students.
Here's the turn: the real problem was never that MarginNote's model was bad. It's that "rank one on a public leaderboard" and "works for Ashcombe's students" turned out to be two separate claims, and nobody had tested the second one before believing the first.
Feedback rated helpful and specific, by cohort, on 200 raw unedited essays
The overall average across all submissions looked like 81 percent. It never showed this gap, because most submissions had already been edited before reaching the model.
The leaderboard never lied. It just never met a student like the ones actually sitting in Ashcombe's classrooms.
At its worst, a district rolls a tool out to every school on the strength of a public rank, and the students who need the clearest feedback most quietly get the least useful comments, while every dashboard says it's working.
The choice I would take back
When MarginNote launched, Ashcombe's rollout email told teachers the feedback model was "the highest-rated system available," citing its leaderboard rank, with no caveat about what population it had actually been tested on. That made sense before anyone had measured a real gap. It stopped making sense once the gap turned out to be 55 points wide on the exact students the district serves most.
What I would leave alone: for the district's advanced-placement essay track, almost entirely native-English writers close to the benchmark's own style, the leaderboard rank held up fine, agreement stayed above 85 percent with no gap at all.
The lesson: a leaderboard number is a real number. It's just a real number about a different classroom than yours, until you've tested it on your own.
Now here is the same thing as a story
The short version above is what you'd say in the interview room. Read this one for what it felt like the six weeks between launch and the morning a random audit changed what Ingrid trusted.
Ingrid Kallevik could read a vendor's model card and tell within a minute whether the claims were dressed up or real.
The week school started, Ingrid picked the model behind MarginNote because it topped a public essay-scoring leaderboard, 92 percent agreement with human graders, ahead of 13 other systems. For the first month, teacher feedback on the tool was glowing. Comments looked polished, specific, and fast.
The habit thinned in three beats, quietly, among the teachers using it. At first, most teachers pasted a student's essay in exactly as written. Within a few weeks, teachers who taught sections with more multilingual learners started lightly tidying a confusing sentence or two first, since the raw draft sometimes produced feedback that read as harsh or off-target. By week six, some teachers were rewriting a struggling student's paragraph structure almost entirely before ever generating feedback, just to get something usable back.
The real gap was three weeks old by the time anyone went looking for it.
A district curriculum specialist, running an unrelated quarterly audit, pulled ten random essay-to-feedback transcripts across three schools. Six of the ten showed clear signs the essay text had been rewritten before it ever reached the model.
Knowledge spark: why would teachers start cleaning up essays before feeding them in?
Once a tool's output feels harsh, confusing, or beside the point on certain writing, people learn its failure shape fast, often without meaning to. They start smoothing what they feed it, because getting something usable back to a student today feels more urgent than reporting a pattern. It's a completely reasonable thing to do, and it quietly erases the evidence that anything was wrong.
Ingrid ran three cause candidates. First, benchmark contamination, the model might have memorized answers on essays close to the leaderboard's own set; ruled out, since the leaderboard's own held-out set was clean. Second, a tracking bug in the helpfulness survey; ruled out, since the survey was firing correctly on every submission. Third, distribution mismatch: the leaderboard's essay corpus was drawn almost entirely from suburban, native-English classrooms, and Ashcombe's district is 38 percent multilingual learners.
Two suspects cleared quickly. The third one held up under the actual test.
We weren't measuring whether the model could grade an essay. We were measuring whether it could grade the specific essays the leaderboard happened to include.
The evidence test: pull 200 raw, untouched essays straight from the classroom app before any teacher had touched them, split by cohort, and run MarginNote cold. Native-English writing: 84 percent rated helpful and specific by an independent teacher panel. Multilingual-influenced writing: 29 percent, with comments repeatedly described as generic or focused on surface grammar instead of the student's actual argument.
When the rollout email went out in August, someone on the team wrote "highest-rated system available" straight from the leaderboard page, and it read as completely fair, since nobody had yet pulled a single raw multilingual-learner essay to check.
Share of submitted essays that had been teacher-edited before reaching MarginNote, week by week
This line was climbing for six weeks before anyone measured it. It would have told the real story well before the audit did.
Rerun the same nine weeks with a raw, untouched sample tested from day one, split by cohort: the 55-point gap surfaces in week two, when only 40 raw multilingual-learner essays exist to test against, instead of week nine, after report cards for the first grading period had already gone home.
What I'd tell myself, hearing that curriculum specialist read out six of ten rewritten transcripts: the leaderboard was never lying. It just never met a student like the ones actually sitting in Ashcombe's classrooms, and I never asked it to.
TRACE, run backward from a healthy-looking averageNot a script for distrusting every benchmark. TRACE is what tells you exactly which population a rank was ever describing.
T
Timeline. When did it actually start, versus when anyone noticed?
Teachers began lightly editing input by week three. The gap was fully formed by week six. Nobody measured it until a routine audit in week nine.
The real problem is always older than the moment it gets noticed.
R
Recut. Slice it, don't average it.
Native-English writing: 84 percent helpful. Multilingual-influenced writing: 29 percent. The blended average, 81 percent, hid the second number completely.
An 81 percent average is often one segment doing fine and another cratering underneath it.
A
Assume nothing. Rule out instrumentation first.
Before trusting the segment gap as real, Ingrid confirmed the helpfulness survey was firing correctly on every submission, in every classroom.
A tracking bug can look exactly like a real quality gap on a dashboard.
C
Cause candidates. Three named, not everything possible.
Benchmark contamination, ruled out. A survey bug, ruled out. Distribution mismatch, the leaderboard's own essay corpus skewed native-English, confirmed.
This is the hardest step, and the one that separates a real diagnosis from a guess.
E
Evidence test. The one check that separates the top two.
200 raw, untouched essays, split by cohort, run cold. The 55-point gap between native and multilingual-influenced writing appeared immediately, and only on the raw sample.
The recap, one line per letter: timeline shows the gap forming weeks early, recut finds it by segment, assume nothing clears the survey, cause candidates narrows to distribution mismatch, and the evidence test is the raw-sample rerun that proved it.
And if you want to be sure it really works, try it somewhere elseSame five steps, a veterinary clinic instead of a classroom. What people quietly clean up before feeding the model changes, the method doesn't.
Callum Ostergaard runs operations for Sorrelkirk Veterinary Partners, which uses TriageNote, a model that drafts a same-day urgency read from an owner's written description of their pet's symptoms. Mapped onto TRACE: timeline, the model was picked off a public triage-accuracy leaderboard in March, and real complaints about missed urgent cases didn't surface until May, weeks after front-desk staff had quietly started rewording anxious, run-on owner messages into short clinical phrases before submitting them. Recut: urgency calls on staff-rewritten descriptions scored fine, 88 percent appropriate, but on raw owner messages, pulled straight from the online intake form, appropriate urgency calls dropped to 51 percent. Assume nothing: they first confirmed the intake form itself wasn't dropping symptom text, it wasn't. Cause candidates: model drift, ruled out, the model hadn't changed; a labeling error in the benchmark, ruled out; distribution mismatch, the leaderboard's training messages were clinician-written summaries, not panicked owner texts, confirmed. Evidence test: 150 raw owner messages, run cold, showed the model missing urgency specifically when messages were long, emotional, and non-clinical, exactly the shape of a real owner's message and exactly the shape the leaderboard never included.
A leaderboard tests the essay on the left. Your students hand in the one on the right.
Building a real eval set fixes this for good, not just for this one launch.
Four things a leaderboard's own test set will never have: your own students, in their own words.
A leaderboard rank is a real tiebreaker exactly once you've confirmed the populations actually match.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a leaderboard measures the benchmark's own essays, not your students, test raw and untouched data before trusting the rank," and stop.
Cost: no time to build a full held-out eval before launch. Say so honestly, and commit to pulling even 50 raw, untouched samples from your real population before going district-wide.
The model got better, for real: if a new model claims a higher leaderboard rank than your current one, that's still worth testing on your own raw data first, since a higher rank on someone else's essays says nothing about your own multilingual learners.
Where people run it wrong.
They treat leaderboard rank as a finished decision instead of a starting hypothesis to test.
They look at an overall average and never recut it by the segment that actually matters to them.
They never check whether people are quietly cleaning input before it reaches the model, which hides the real number from view.
How to use it live. The moment an interviewer mentions a model's leaderboard rank, ask yourself out loud: what population was that leaderboard's own test set drawn from, and does it look like the population I'd actually be serving? That question alone usually separates a strong answer from a name-drop.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family sits underneath this story?
Tap to flip
ANSWER
Pre-editing flip: teachers feed the model a student's real essay at first, then quietly sanitize it before submitting once they learn where it struggles.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ingrid Kallevik, product manager at Ashcombe Learning, who picked the model behind MarginNote off a public leaderboard and later diagnosed the real gap.
3 · THE HABIT
What did teachers stop doing because the model seemed to work?
Tap to flip
ANSWER
Submitting a student's essay exactly as written. By week six, some teachers rewrote a struggling student's paragraph structure before ever generating feedback.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Feeding the model a student's real, unedited writing, versus quietly cleaning it up first so the feedback that comes back looks usable.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling teachers in the rollout email that the model was "the highest-rated system available" based on its leaderboard rank, with no test yet run on Ashcombe's own students.
6 · THE NUMBER
Fill in the blank: on 200 raw, untouched essays, native-English writing scored ___ percent helpful, multilingual-influenced writing scored ___ percent.
Tap to flip
ANSWER
84 percent; 29 percent. A 55-point gap the blended 81 percent average never showed.
7 · THE REPLAY
Same nine weeks, raw testing from day one. What changes?
Tap to flip
ANSWER
The 55-point gap surfaces in week two, on just 40 raw essays, instead of week nine, after the first grading period's report cards had already gone home.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Sorrelkirk Veterinary Partners' TriageNote, a concealment-adjacent input flip: front-desk staff reword anxious owner messages into short clinical phrases before submitting them.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: MarginNote's model ranked ___ of 14 systems on the public leaderboard, with ___ percent agreement with human graders.
Show hint
Look at the direct answer's supporting stage in the walkthrough.
Show answer
First (rank one); 92 percent. A real, true number, just one measured on a population unlike Ashcombe's own classrooms.
Multiple choice
2. Why did the overall helpfulness average, 81 percent, fail to show the real problem?
A. The survey tool was broken and recorded the wrong scores.
B. MarginNote's model changed silently partway through the term.
C. Most submissions had already been teacher-edited before reaching the model, hiding the raw failure rate underneath a blended average.
D. Teachers stopped using the tool entirely by week six.
Show hint
Look at the "assume nothing" step, and what it ruled out.
Show answer
C. By the time of the audit, 60 percent of submissions were teacher-edited, which smoothed over the exact gap a raw sample later revealed.
True or false
3. True or false: ruling out the survey tracking bug means the segment gap "assume nothing" step found was definitely a real quality problem.
True
False
Show hint
Think about what ruling out instrumentation actually proves.
Show answer
True. Once the survey was confirmed to be firing correctly on every submission, the segment gap it recorded couldn't be explained away as a measurement error, which is exactly what "assume nothing" is for.
Short answer, where it wouldn't matter
4. Name a part of Ashcombe's own student population where the leaderboard rank held up fine, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The advanced-placement essay track, almost entirely native-English writers whose style is close to the benchmark's own essay corpus, where agreement stayed above 85 percent with no real gap.
Short answer, apply it yourself
5. Pick a tool you use that was chosen partly because of a public ranking or review score. What population might that ranking have actually been tested on, and how would you check if it matches yours?
Show hint
Think about who wrote the reviews or built the benchmark, versus who you actually are.
Show answer
Model answer: A grammar-checking tool with strong review scores from professional writers. Checking it against a first-generation college student's early draft, a very different population, would show whether the praise actually transfers.
Short answer, work the number
6. If the district were only 10 percent multilingual learners instead of 38 percent, would the same 55-point segment gap have been as urgent to catch before rollout?
Show hint
Think about how many real students the gap would actually touch.
Show answer
Model answer: Less urgent in scale, though still worth fixing. At 10 percent of students, the same gap affects far fewer kids per class, but the underlying lesson, test on your own raw data before trusting a rank, doesn't change with the percentage.
Before you close the answer
Why this works
Tests whether you treat a leaderboard rank as evidence about a specific test population, not as a finished verdict, and whether you'd think to check what people quietly do once a tool starts struggling.
Follow-up traps
"Couldn't you just fine-tune the model on Ashcombe's own essays instead of testing first?" Response: eventually yes, but you'd still need the raw, untouched test set first, to know the gap exists and to measure whether fine-tuning actually closed it.
"Isn't asking teachers not to edit input before submitting the real fix?" Response: it removes the symptom, not the cause. Teachers were editing because the raw output genuinely read as harsh or off-target for that writing style, so the model still needs to improve on that population, not just get fed cleaner input.
If pressed
The raw-sample re-test was stratified by cohort size before drawing, not randomly sampled from the whole pool, since multilingual-influenced essays made up only 38 percent of submissions and a plain random draw risked too few of them to trust the result.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.