ConceptIntermediateModel Fluency & the AI PM Role / Working with ML engineers and researchers / #1
How do you write a requirement for a team whose output is a probability distribution?
LEAD · why a graded eval beats a yes or no bar, tested on a hail severity model called Graupel
Graupel is Isohyet's hail severity model. Give it one field on one day during hail season and it hands back a full spread of outcomes, from no damage to total loss, not one flat guess. Lucasta Farrelly owns what the research team gets asked to build. Six weeks into a kickoff requirement she wrote herself, a Tuesday remark from the model's lead researcher, Ezekiel Marrowbone, told her the number everyone had been chasing meant nothing at all.
The direct answer
Write the requirement against a real eval set and a graded score with a cut off point, never a yes or no bar. Pick a sample of real cases that looks like the actual world, state the score the model has to clear, and say plainly which mistakes the team may make more of to make fewer of the ones that cost the most. A requirement that just says "correctly predict X" can't be met or failed by a team whose output is a spread of outcomes, not one guess.
Do this, in order
Write the requirement against a real eval set and a graded score with a cut off point, not a yes or no bar.Why: a spread of outcomes can't clear or miss a bar that only has two settings.
Build the eval set from a real, representative sample, with enough rare severe cases in it to actually grade them.Why: one dramatic storm proves nothing about how the model handles an ordinary Tuesday in June.
Pick the score that would show up in the first training run, like how much real damage the model actually catches, not the one that waits a whole season to prove anything.Why: that score caught the problem the same week. The real reserve numbers would have taken months.
Write down which mistakes the team may trade for which, before training even starts.Why: without that on paper, a team optimizes whatever moves easiest, and cheap mistakes crowd out the expensive ones.
Treat raw accuracy as a warning sign on a job this lopsided, never as proof of anything.Why: a model that says "no damage" every single time already scores 96 percent without learning a thing.
Leave a plain yes or no requirement alone anywhere the real answer actually is one or the other.Why: a check like "did we get a usable picture of this field today" has one right answer, and grading that like a distribution wastes a week nobody needed to spend.
How to answer this, stage by stage
Nobody is grading whether you can define a probability distribution. They're grading whether you can say, in the team's own words, what a requirement has to contain before it can actually be passed or failed.
1
Scope it to one product and one team
Say it like this
"Let's ground this in one real build. Graupel is Isohyet's hail severity model, it reads one field on one day and hands back a spread of outcomes, not one guess. Lucasta Farrelly owns what the research team gets asked to build. I'll answer using her, not the idea of a requirement in general."
Why this works
Naming a real system and a real owner stops "write a requirement" from turning into a generic checklist.
2
Say your structure out loud
Say it like this
"I'll run this as LEAD. Link, what the requirement actually has to track. Early signal, the score that would have caught the problem in week one instead of at the end of a season. Abuse, how a requirement like this gets written badly. Decision, what actually goes in a requirement a spread of outcomes can pass or fail."
Why this works
Two seconds of structure tells the interviewer you have a method, not a definition you memorized the night before.
3
Answer the literal question first, in one line
Say it like this
"Short version: you write it against a real eval set, a graded score with a cut off point, and a stated tradeoff, never a single yes or no bar. That's the whole shape of it."
Why this works
This is the actual answer, said plainly, before any story. Skip it and the interviewer has to go dig for it.
4
Reframe why anyone's actually asking this
Say it like this
"This isn't really 'can you write a requirement.' It's 'do you understand that a probabilistic system doesn't have one right answer to be correct about,' because a PM who writes 'correctly predict X' for a team like this has already asked for something that can't be met or missed."
Why this works
Shows the interviewer you get why this question exists, not just that you can recite an answer to it.
5
Give the decision, committed
Say it like this
"So here's what I'd actually do. I'd hand the team an eval set of real cases, a graded score with a real cut off, and a written tradeoff they're allowed to make. No single accuracy percentage standing in for all three of those."
Why this works
This is the direct answer, said out loud, with no hedge in it.
6
Read the actual requirement out loud, word for word
Say it like this
"Something like this: 'Score against these five hundred real field days. The model's stated chance of severe or worse damage has to sit within five points of what actually happened, checked across ten bands of confidence. The top end of its forecast has to land within forty dollars an acre of the real claims paid, on average across the set. You may trade some misses on small hail for fewer misses on the severe end, since a missed severe event costs about forty times more than a false alarm on a small one.'"
Why this works
This is what a real requirement actually sounds like. A candidate who can say this out loud has done the thing, not just described doing it.
7
Prove it with the real six weeks, numbers first
Say it like this
"Here's what actually happened at Isohyet. Lucasta's kickoff requirement asked for 95 percent accuracy on whether a field would see damaging hail. Damaging hail only happens on about 4 of every 100 field days, so a model that always says no already clears 96. Six weeks in, the real model sat at 95.8 percent accuracy and was catching about 3 percent of the real damaging days. Nobody had noticed, because the number everyone watched kept climbing toward the target."
Why this works
Two real numbers, four months apart, beat any amount of talk about what a good requirement looks like.
8
Name the abuse, say what you'd leave alone, then close
Say it like this
"Someone on the underwriting side later suggested just testing the model against last July's worst storm, since that was the costliest one on record. I'd turn that down, because passing on one storm proves nothing about an ordinary week. I also wouldn't force a graded eval onto Isohyet's own picture check, whether a field's photo came through clean that day has one right answer, and grading that like a distribution wastes a week nobody needed to spend. So: write the requirement against a real sample, a graded score, and a stated tradeoff, and a team like this finally has something it can actually pass or fail."
Why this works
Naming a bad idea you'd turn down, and a place you wouldn't touch, is what proves this is judgment and not a rule copied off a slide.
Let's learn
What does "correct" even mean for a model that never gives you one answer?
Isohyet builds weather risk tools for crop insurers. Its main model, Graupel, looks at one field on one day during hail season and hands back five numbers: the chance of no damage, minor damage, moderate damage, severe damage, and total loss. An insurer uses the whole spread to decide how much money to set aside for a policy, not one flat yes or no.
Knowledge spark: what is a probability distribution here?
Not one guess. A spread of possible outcomes, each with its own chance attached. Instead of "this field will see damaging hail," the model says "a 2 percent chance of total loss, a 6 percent chance of severe damage, an 11 percent chance of minor damage, and an 81 percent chance of nothing at all." All four numbers together are the real answer.
A checkbox has one setting. A spread has five, and every one of them matters to an insurer.
For most of Isohyet's four years, requirements were easy to write, because most of what the company built was a plain yes or no. Did a field's satellite picture come through today. Did a claim match a policy on file. Lucasta had written dozens of requirements shaped exactly like that, and every one of them worked, because a yes or no question actually has a yes or no answer.
Then the insurers asked for something new: a model that could tell them, ahead of the season, how much money to set aside for hail. That meant a real spread of outcomes, not a flag. Lucasta wrote the kickoff requirement the way she'd always written one. "The model should correctly predict whether a field will see damaging hail, at least 95 percent of the time."
It sounded sharp in the room. Everyone nodded. Here is the part nobody in that room knew yet: damaging hail only actually happens on about 4 of every 100 field days in hail season. A model that says "no damage" every single time, without reading a single number off the sky, already gets 96 of those days right. Lucasta had asked the team to clear a bar that doing nothing at all had already cleared.
Ezekiel Marrowbone's team spent six weeks chasing that 95 percent anyway. Every week the number climbed a little. Week two, 93.1. Week four, 95.1. Week six, 95.8.
Reported accuracy vs. real damage caught, by week
Reported accuracyRecall on real damaging events
Accuracy climbed for six straight weeks and never really moved after that. The number that actually meant something sat under 4 percent that whole time, because nobody had written a requirement that could see it.
The accuracy number climbed for six weeks straight. The share of real damaging days it actually caught never moved past 3 percent the whole time.
Here is the turn. The six weeks were never the real cost. The real cost is that no team, however good, could have written code that passed Lucasta's requirement and also did the job, because the requirement described something a coin flip had already cleared.
Knowledge spark: what is recall?
How much of the real thing the model actually catches. If 20 fields truly get damaged this week and the model flags 12 of them, recall is 60 percent. A model can miss almost every real case and still score high on plain accuracy, if the real cases are rare.
The requirement got a checkmark. The field that actually lost its crop never got a second look.
An underwriting VP heard about the stall and offered a fix that felt reasonable in the room: just run the model against last July's derecho, the worst storm on record, and see if it caught that one. Lucasta turned it down. A single storm, however dramatic, says nothing about how the model behaves across a normal season of small, ordinary hail days, which is most of what it will actually see.
Once Ezekiel's team rewrote the requirement around a graded score instead of a yes or no bar, they backtested Graupel against last season's real claims. Reserve pricing error, the gap between what an insurer set aside and what it actually paid out, fell hard.
Reserve pricing error per policy, before and after the eval spec fix
BeforeAfter
A model that mostly rides the base rate gives an insurer almost nothing to price against. The pricing error only closed once the requirement asked for something the model could actually be scored on.
One of these numbers is ready the same week. The other one waits for an entire season to finish.
The choice I would take back
Four months earlier, at kickoff, Lucasta wrote Graupel's requirement the same way she'd written every requirement before it: a plain yes or no bar with a percentage attached. That shape had worked fine for a fraud flagging feature at her last job, where a transaction really is either fraud or not. It stopped working the moment the output stopped being one answer and became a spread of them, and nobody in that first meeting knew the model well enough to catch it.
What I would leave alone: Graupel also runs a daily check on whether a field's satellite picture actually came through clean, no clouds blocking the view, no corrupted file. That really is a yes or no question, and a plain pass or fail requirement is exactly right for it. Forcing a graded eval spec onto a data check like that would cost a week nobody needed to spend.
The lesson: a requirement that sounds confident and measurable in a room isn't the same thing as a requirement a team can actually fail. If the honest answer to "can this bar be cleared by doing nothing at all" is yes, you didn't write a bar. You wrote a trap and handed it to your own project.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one for the four months it took to write a requirement Graupel's team could actually pass or fail.
Lucasta Farrelly had written requirements for six years before Graupel, four of them at Isohyet. She was good at it. Ask her to spec a fraud check or a claims matching rule and she'd have a clean, testable sentence back to you inside a day, because those systems only ever had one right answer, and she knew exactly how to pin one down.
Both of them doing their job carefully. Neither one had a reason yet to think the number meant nothing.
When the insurers asked for a model that could size up hail risk ahead of a season, Lucasta ran Graupel's kickoff meeting the way she ran every kickoff. A conference room, a whiteboard, the VP of underwriting leaning forward because he wanted a number he could take to a reinsurer. She wrote it on the board herself: "correctly predict whether a field will see damaging hail, at least 95 percent of the time." The room nodded. It felt like the sharpest requirement in the meeting.
For the first two weeks, checking in on Graupel's progress was the best part of Lucasta's Tuesdays. Ezekiel Marrowbone's update always had a climbing number in it, and a climbing number toward a stated target reads exactly like progress.
By week two, 93.1. She asked a couple of questions about the model's features, glanced at a chart, moved on to her next meeting. By week four, 95.1, close enough to see the finish line, and her questions got shorter. By week six, 95.8, and she mostly just said "great, keep going" and closed her laptop.
Nobody decided to stop asking real questions. A climbing number just made it feel like there was nothing left to ask.
Ezekiel wasn't hiding anything either. He'd been chasing that 95 percent honestly, trying new features, tuning a threshold, doing the job the requirement described. That Tuesday evening, week six, he said the thing he'd been sitting with for a few days. "We've hit 95.8 three weeks running. I want to be straight with you. A model that always said no would already score 96."
Lucasta didn't say anything for a second. Then she asked the only question that mattered. "How much of the real damage is it actually catching?"
Three percent.
We did not lose six weeks to a hard problem. We lost six weeks to a question a spread of outcomes was never going to answer.
I want to say the problem was that hail is hard to predict. It is. But that's not really the story. Lucasta never had a number in her head for what "done" would actually look like, because the requirement she wrote never described anything a model could be genuinely scored against. She had a percentage that felt like a bar, and it only worked as a bar until someone did the arithmetic on how rare damaging hail actually is.
The decision she'd take back sits in that first kickoff meeting, four months earlier, marker in hand, a room full of people wanting a firm number before lunch. She reached for the shape of requirement that had always worked for her, because nobody in that room, herself included, knew Graupel's output well enough yet to say it needed a different shape entirely.
Run the same four months again, with the requirement written properly from day one. The first two weeks go to building the eval set itself: five hundred real field days pulled from six seasons and forty counties, with enough severe and total loss cases folded in on purpose that they can actually be graded, not drowned out by ordinary days the way they are in raw history. The next three weeks go to tuning against a real target: predicted chance of severe or worse damage within five points of what actually happened, and the top of the forecast within forty dollars an acre of real claims paid. By week five, Graupel clears both. Backtested against last season, reserve pricing error lands at 95 dollars a policy instead of 340. It took five weeks that went somewhere, not six that didn't.
What I'd tell myself, back at that whiteboard: a requirement that sounds confident in a room and a requirement a team can actually fail are not the same sentence. I wrote the first one and called it finished.
LEAD, for a requirement a spread of outcomes can pass or fail
Not a way to prove Lucasta was careless. LEAD is what forces you to say what the requirement should actually track, and to catch a broken one before a season of real money moves on it.
LLink. The business outcome that actually matters.
Not whether the model's own accuracy number looks tidy. What actually matters to an insurer is how close the money they set aside for a policy lands to what they actually pay out. A binary accuracy score can hit 96 percent while telling an insurer nothing useful about that gap, because it never had to.
Graupel's whole reason to exist is reserve pricing error. Its own internal accuracy score is not that number, and treating it as a stand in for that number is where the requirement broke.
EEarly signal. The thing that moves before the outcome does.
Here the early signal is recall on the real damaging cases, checked against a representative sample the same week a training run finishes. Reserve pricing error can't be known for certain until an entire hail season plays out and real claims come in. Recall on a graded eval set is available the same day.
Recall sat under 4 percent for six weeks, visible in every single training run, while the reserve numbers stayed a mystery until someone finally asked the right question of the eval set instead of the accuracy score.
AAbuse. How a requirement like this gets written badly.
Two ways, both real, both at Isohyet. First, a PM asks for a plain accuracy percentage on a rare event, without checking what a model that changes nothing would already score. Second, once that stalls, someone offers a shortcut: test against one dramatic, hand picked case instead of a real representative sample, because it feels like a fast way to prove the model works.
"95 percent accuracy" asked for less than doing nothing. "Just catch last July's derecho" would have proven nothing about an ordinary week, which is nearly every week Graupel actually runs on.
DDecision. What actually goes in the requirement.
Three things, every time. A real eval set built from history, with enough rare severe cases folded in that they can be graded at all. A graded score with a real cut off point, tied to what the business actually cares about, calibration against real frequency, and the top of the forecast against real dollars paid. And a written tradeoff: which mistakes the team may make more of, in exchange for making fewer of the ones that cost the most.
Graupel's rewritten requirement let the team trade some misses on small hail for fewer misses on severe or total loss, because a missed severe event costs about forty times more in reserve error than a false alarm on a small one.
Three parts, every time. Leave one out and the requirement goes back to being a wish with a number attached.
The recap, one line per letter: link the requirement to what the business actually pays for, reserve accuracy, not the model's own internal score. The early signal is recall on a real, graded sample, because it's ready the same week and the business outcome isn't ready for months. Name both abuses plainly, a binary bar a coin flip already clears, and a single dramatic case standing in for a real sample. And the decision is what makes it real: an eval set, a graded score with a cut off, and a written tradeoff, all three, every time.
Two things worth saying outright, since the real judgment sits here. Isohyet considered raising the bar instead of changing its shape, asking for 99 percent accuracy rather than 95. That would have made the trap worse, not better, since it still asks for a single yes or no answer on a question that was never shaped like one. The AI specific failure worth naming by name is base rate masking: a model can score high raw accuracy simply by leaning on whichever outcome is already common, hiding the fact that it has learned almost nothing about the rare cases that actually matter. The guardrail that catches it is watching recall on the minority outcome by itself, as its own tracked number, never folded into one blended accuracy score. And the tradeoff was real and taken on purpose: letting the model miss a few more small, cheap hail events in exchange for catching more of the expensive ones costs some precision on the easy cases, and Isohyet accepted that cost because a missed catastrophic event costs roughly forty times more in reserve error than a false alarm on a minor one.
And if you want to be sure it really works, try it somewhere else
Same four letters, a flash flood routing tool instead of a hail model, and this time the rare event a binary requirement hides is a stranded truck instead of a ruined field.
Freshet runs Spillway, a model that forecasts flash flood risk on trucking routes for logistics companies, so a dispatcher can reroute a load before a road closes under water. Pandora Quraishi owns what Spillway's team gets asked to build, and hit a close cousin of Lucasta's exact problem two months after a similar kickoff.
Same shape of mistake as Graupel's. A flood closes a route on about 3 of every 100 route days, so saying no every time already scores 97.
Spillway's kickoff requirement asked the team to "correctly flag routes that will flood, at least 96 percent of the time." Route closing floods happen on about 3 of every 100 route days during flood season, so a model that never once flagged a route already cleared 97. Under that requirement, Spillway's real recall on actual closures sat near 5 percent for weeks, hidden behind a reported accuracy score that looked almost finished the whole time.
The decision Pandora would take back
Spillway's kickoff requirement was written the week before a board update, when a firm sounding percentage mattered more in the room than whether it described a real bar. That made sense for the meeting it was written for. It stopped making sense the moment training actually started and a model with zero skill could already clear it.
Once Pandora's team rewrote the requirement, closure prediction went from a percentage nobody trusted to a graded eval scored against five hundred real route days, with the team allowed to trade some misses on minor ponding for fewer misses on full closures, since a stranded truck costs roughly twenty five times more than an unnecessary reroute.
Cost per flagged shipment, reroute cost vs. stranding cost, before and after
Reroute costStranding cost
Reroute cost actually went up a little on purpose, more borderline routes got flagged early. Stranding cost, the expensive kind, fell by nine tenths, because the requirement finally let the team trade for it.
Mapped onto LEAD, the shape holds exactly. The link is a dispatcher's real routing cost, not Spillway's own internal score. The early signal is recall on a graded, representative sample of route days, ready the same week, while the true routing cost only shows up after a flood season plays out. The abuse Pandora found was the same one, a binary bar a near empty model already cleared. And her decision matched Lucasta's: a real eval set, a graded score with a cut off, and a written tradeoff, sized to what actually costs the most when it's missed.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: write it against a real eval set, a graded score with a cut off, and a stated tradeoff, never a yes or no bar.
Cost: no budget this quarter to build a proper labeled eval set. Whoever owns the model pulls a smaller, honest sample by hand, fifty real cases instead of five hundred, a slow real eval still beats a fast fake one.
The model got better, for real: say a new version of Graupel doubles its recall on severe cases out of the box. The eval set and its cut off stay in place anyway, because a stronger model can still be ungraded, and a higher starting number is not the same thing as a bar that means something.
Where people run it wrong.
They ask for a percentage without checking what a model that changes nothing would already score on that same measure.
They swap one dramatic hand picked case for a real representative sample, because it feels like fast proof instead of the slow kind.
They raise the bar instead of changing its shape, asking for 99 percent instead of 95, which only makes a binary trap harder to notice, not less of one.
How to use it live. When an interviewer hands you a question like this, buy yourself a second by asking out loud what a model that did nothing at all would already score. That question is the whole answer in miniature, and it previews everything you're about to say.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question that asks how to write a requirement for a probabilistic team?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Built for metric questions, it forces you to name what the requirement should actually track and the score that would catch a broken one early.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Lucasta Farrelly, who owns Graupel's requirements at Isohyet, and Ezekiel Marrowbone, who leads the research team chasing the accuracy bar she wrote.
3 · THE LINK
What should Graupel's requirement actually track?
Tap to flip
ANSWER
Reserve pricing error, how close an insurer's set aside money lands to what it actually pays out, not the model's own internal accuracy score.
4 · THE EARLY SIGNAL
What moved first here, and what took a whole season to catch up?
Tap to flip
ANSWER
Recall on real damaging events, visible in every training run. Reserve pricing error, the number that actually mattered to an insurer, only shows up once a full hail season plays out.
5 · THE OLD DECISION
What decision would Lucasta take back?
Tap to flip
ANSWER
Writing Graupel's kickoff requirement as a plain accuracy percentage, the same shape that worked for a deterministic fraud check at her last job, instead of a graded eval spec.
6 · THE NUMBER
Fill in the blank: damaging hail happens on about ___ of every 100 field days, so a model that always says no already scores ___ percent.
Tap to flip
ANSWER
4; 96 percent. Lucasta's 95 percent bar asked the team for less than doing nothing at all.
7 · THE REPLAY
Same four months, requirement written properly from day one, what changes?
Tap to flip
ANSWER
Two weeks build a real 500 case eval set, three more tune against a real cut off. By week five, reserve pricing error backtests at 95 dollars a policy instead of 340, five weeks that went somewhere instead of six that didn't.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the parallel problem there?
Tap to flip
ANSWER
Spillway, a flash flood routing model run by Pandora Quraishi at Freshet. There a "flag routes that will flood, 96 percent of the time" bar hid a model catching only 5 percent of real closures, until a graded eval spec fixed it.
Check yourself Score: 0 / 0
True or false
1. True or false: if Lucasta had simply raised the bar from 95 percent to 99 percent, that would have fixed the requirement.
True
False
Show hint
Check the "Swap the trigger and it still runs" line about raising the bar, in Section 4.
Show answer
False. Raising the number keeps the same shape of trap, a single yes or no bar. It would still say nothing about whether the model catches real damaging cases, and it makes the trap harder to notice, not less of one.
Multiple choice
2. Why did Graupel's model sit at 95.8 percent accuracy for three weeks running while catching almost none of the real damage?
A. Ezekiel's team stopped trying to improve it.
B. Damaging hail is rare enough that a model which always says no already scores about 96 percent, so the accuracy number was hiding near zero real skill.
C. The satellite imagery pipeline was broken that month.
D. Lucasta had lowered the target without telling anyone.
Show hint
Check the "A, abuse" step in the framework recap.
Show answer
B. This is base rate masking. A rare, lopsided label lets raw accuracy stay high while telling you almost nothing about the cases that actually matter.
Fill in the blank
3. For six weeks, recall on real damaging events was stuck near ___ percent, while reported accuracy sat near ___ percent the entire time.
Show hint
Look at the line chart in Let's learn.
Show answer
3 percent; 95.8 percent. The gap between those two numbers is the whole reason the original requirement never worked.
Short answer, name the reversal
4. What old decision would Lucasta take back, and why did it make sense the first time she made it?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Writing Graupel's kickoff requirement as a plain accuracy percentage, the same shape she'd always used. It made sense because it had worked for a fraud check at her last job, where the answer really was one thing or the other. It stopped making sense once the output became a spread of outcomes instead of one flag.
Short answer, apply it yourself
5. Think of an AI product you use whose real output is a spread of possibilities rather than one clean answer. What would a graded requirement for it actually need to contain?
Show hint
Think about a real sample, a score with a cut off, and a stated tradeoff, not one flat pass or fail.
Show answer
Model answer: A route planning app's arrival time is really a range, not one number. A real requirement would score it against a set of real past trips, require its stated range to actually contain the true arrival time on, say, 90 percent of them, and allow it to trade a wider range on unfamiliar routes for a tighter one on routes it knows well.
Short answer, work the number
6. If damaging hail happened on 1 field day in 100 instead of 4, what would a model that always says no score, and would that change whether a plain accuracy requirement is a fair bar?
Show hint
Work out the score first, then ask whether the requirement's real problem gets better or worse.
Show answer
99 percent, and no, it gets worse, not better. The rarer the real event, the higher the free score doing nothing already earns, so a plain accuracy bar becomes an even weaker test of whether the model has learned anything at all.
Before you close the answer
Why this works
Tests whether you understand that a probabilistic system's output isn't a single fact to be right or wrong about, and whether you can turn that understanding into an actual requirement someone could hand to an engineer today. Most candidates can say "use an eval, not accuracy" and stop there, without ever writing the eval spec itself.
Follow-up traps
"Isn't building a five hundred case eval set just slower than shipping and seeing what happens?" Response: it costs about two weeks up front, which is exactly what the six wasted weeks under the old requirement cost with nothing to show for it. The eval set is the reason the following weeks moved a real number.
"Doesn't letting the team trade misses on small hail for fewer misses on severe hail just lower the bar?" Response: no, it points the bar at what actually costs money. A missed severe event costs about forty times more in reserve error than a false alarm on a small one, so that tradeoff is the requirement doing its job, not softening it.
If pressed
Graupel's eval set oversamples severe and total loss cases about thirty to one against how often they actually occur, on purpose, so there's enough of them to grade at all. The calibration score itself is then computed with the real, un-oversampled frequency weighted back in, otherwise the oversampling would fake a miscalibrated result where none exists.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.