CaseAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #24

How do you measure impact when the AI feature changed the workflow entirely?

SPARK · continuous herd-health monitoring for a dairy cooperative

Byrewell reads three health signals off every cow in Thistlemoor Dairy Cooperative's four barns, all day, and flags the ones worth a second look before anyone would ever see it by eye. Rhydian Ockendon owns its product story at Oxendale Analytics. Byrewell replaced the twice-daily herd walk completely in Barn 1 and Barn 2, and Percy Cathmore, Thistlemoor's general manager, wants a number for it at the co-op's annual member meeting. The thing Byrewell replaced does not happen anymore, so there is nothing left to time it against.

The direct answer
Do not measure Byrewell against the walk it replaced. There is no walk left to measure against. Anchor the case to the outcome the walk was actually protecting: how many days pass between a cow getting sick and someone treating her. Build that number from Barn 3 and Barn 4, which kept walking for four more months after Barn 1 and Barn 2 switched, on purpose, so a real comparison group still exists. Report the honest cost that comes with it too: more staff hours spent checking cows who turn out fine, not a free win.
Do this, in order
  1. Anchor the case to days from illness onset to treatment, never to the walk's own minutes or headcount.Why: it is the one number that means the same thing whether a person or a sensor is doing the watching.
  2. Hold Barn 3 and Barn 4 on the old system for the full comparison window before switching them too.Why: without a barn still being walked somewhere, there is no honest "before" left anywhere in the herd.
  3. Report the false-positive follow-up rate right next to the detection-lag gain, never the raw flagged-cow count alone.Why: thirty-four flags a day reads like an emergency. It isn't, and the count by itself can't say so.
  4. Refuse to simulate what Barn 1 "would have" caught if the walk had kept going.Why: that builds a fake comparison out of the same kind of model reasoning being tested.
  5. Say plainly that Byrewell costs more staff time in the switched barns, not less.Why: an honest trade-off Percy can weigh beats a saving that isn't real.
  6. Leave the dry-cow and heifer walk exactly as it is.Why: nothing changed for them yet, so there is no measurement problem to go solve there.

How to answer this, stage by stage

Nobody is grading whether you can say "measure the outcome, not the process." They're grading whether you can build a real comparison when the thing everyone would normally compare against has quietly stopped existing.

1
Put a real barn and a real name on it before anything else
Say it like this
"Let's ground this in one case. Byrewell is Oxendale Analytics' herd-health system, running across all four of Thistlemoor Dairy Cooperative's barns. Rhydian Ockendon owns its product story. Percy Cathmore, Thistlemoor's general manager, needs a real number for the co-op's annual meeting."
Why this works
Keeps every claim after this checkable against one barn, one person, not a category of dairy farms.
2
Say the shape of the answer before the first fact
Say it like this
"I'm going to run this as SPARK. Say what's actually gone and can't be compared to anymore, name the habit worth building instead, give the real anchor, show it surviving its own bad day, then say what I refuse to fake."
Why this works
Two seconds of structure tells the interviewer a plan is already running, not five thoughts landing in whatever order they occur.
3
Name exactly what disappeared, out loud
Say it like this
"Before Byrewell, checking a cow meant a person walking past her twice a day. That walk is gone in Barn 1 and Barn 2. There's no 'minutes per check' left to compare Byrewell against, because the thing it replaced doesn't happen there anymore."
Why this works
This is the reframe the whole answer turns on. Skip it and the rest sounds like dodging the question instead of actually answering it.
4
Give the anchor, the real thing you'd measure
Say it like this
"Here's what I'd actually build. Measure days from when a cow's real problem started to when someone treats her, not minutes per check, using vet records and her own milk-yield curve to date the real onset. And keep Barn 3 and Barn 4 on the old walk for four more months after Barn 1 and Barn 2 switch, on purpose, so there's still a real 'before' to check the 'after' against."
Why this works
This is the direct answer, and it's concrete enough that Percy could hold Rhydian to it.
5
Walk it through the day it looked like it broke
Say it like this
"Here's what almost undid the whole case. A heat wave in month three flagged a hundred and eighteen cows in a single day, more than three times normal. If alert count were the score, that reads like the model panicking. It wasn't. Techs checked every one, and only three actually needed care. The case never rested on alert count, so that day didn't touch it."
Why this works
A risk you only describe is a warning. A risk you show surviving is a real design decision.
6
Refuse the shortcut, out loud
Say it like this
"Percy's team actually asked for this: run a model over Barn 3 and 4's pattern and use it to guess what Barn 1 'would have' caught if the walk had kept going. I said no. That's building the very comparison I'm supposed to be testing out of the same kind of guesswork. I'd rather hand over real Barn 3 and 4 numbers than a made-up Barn 1."
Why this works
Naming the shortcut you're refusing is what makes "the walk is gone, here's what I did instead" sound like discipline, not an excuse.
7
Land the one sentence, and stop
Say it like this
"So: anchor to days-to-treatment, not minutes-per-check. Keep two barns walking on purpose while two switch. Report the real cost, more staff hours chasing flags that turn out fine, right next to the real gain. That's the case."
Why this works
Restates the direct answer plainly, so Percy leaves with the decision, not just the story.

Let's learn

Byrewell is Oxendale Analytics' health-monitoring system for dairy cows: an ear tag that reads three signals on every cow, all day, plus a camera at the parlor exit that scores how she walks, so a sick cow gets found without anyone walking the barn to look for her.

Before Byrewell, catching a sick cow meant a person seeing it. Bevan Norwood and two techs walked all four of Thistlemoor Dairy Cooperative's barns twice a day, 5:30 in the morning and 4:30 in the afternoon, about twenty minutes a barn, watching how 1,400 cows stood, ate, and moved. On top of the walk, about 15 percent of the herd, roughly 210 cows, got a hand temperature check each day, on a rotation. Even done well, by a team that knew the herd cold, the average sick cow wasn't caught until day 3.2 of whatever was actually wrong with her. Early illness doesn't show on the outside yet.

Hand sketched icon list titled Before Byrewell, checking meant walking the barn. Three rows: a document icon, two walks a day, clipboard checklist. A gauge icon, about 15 percent hand-checked for temp. A question mark icon, day 1 and day 2 of sickness went unseen.
Even a good team walking a barn twice a day is checking from the outside. Sickness doesn't show there first.
Knowledge spark: what is rumination time? The minutes a cow spends chewing her cud each day, lying down, relaxed. It drops fast when something is wrong, often a full day before she looks sick to a person walking past. Byrewell watches it the way a nurse watches a pulse.

Byrewell doesn't wait for the outside to change. It reads rumination minutes, activity, and ear-skin temperature off every cow's tag, all the time, plus a gait score every time she passes through the parlor, and flags her the moment two of those three signals drift together and stay that way for four hours straight.

Hand sketched left to right flow diagram titled How Byrewell flags a cow. Five boxes connected by wobbly arrows: Ear tag reads 3 signals, Rolling score per cow, 2 of 3 cross the line, this box emphasized in red-orange, Held for 4 hours straight, Cow lands on the list.
Five steps, and the middle one is the whole judgment call: one odd signal is normal. Two that agree and stay agreeing is not.

In Barn 1 and Barn 2, where it has been running, the average sick cow gets caught by day 1.2. About thirty-four cows a day land on the flagged list, out of roughly seven hundred cows in those two barns. And the flags don't behave like the old checks did. About 58 percent of them turn out fine when a tech walks out to look, against something close to zero false alarms on the old visual walk, which only ever flagged a cow a trained eye could actually see was off.

Here's the turn. Catching illness two days earlier isn't the hard part to prove. The hard part is that nobody walks Barn 1 or Barn 2 anymore, so there's no "minutes per check" left to measure Byrewell against. The thing it replaced doesn't happen there. A number built the old way, whatever that would even mean now, can't tell Percy Cathmore whether Byrewell is worth what it costs.

There is nothing left in Barn 1 to time. There is still plenty left in Barn 1 to count.
Average days from illness onset to treatment, during the four-month window Barn 3 and 4 stayed on the old walk
4 days 2 days 0 3.1 days Barn 3 & 4, still walked 1.2 days Barn 1 & 2, Byrewell
Barns still walkedBarns running Byrewell
Barn 3 and 4's 3.1 days sits close to Thistlemoor's own years-long baseline of 3.2, which is what makes this comparison group trustworthy, not skewed.
Byrewell's false-positive follow-up rate, week by week, as the threshold settled in
80% 40% 0% settled near 58% W1 W2 W3 W4 W5 W6 W7 W8 W9 W10
Weekly false-positive follow-up rateSettled, week 9 to 10
Week 1 sent a tech to check three flagged cows out of every four for nothing. By week ten that had dropped to a little over half, still real, but no longer the whole story.

At its worst, this gets reported as a shrug: "we think it's working, hard to say by how much exactly," right at the meeting where the co-op decides whether to finish switching Barn 3 and Barn 4, or let Byrewell's contract quietly lapse and go back to walking.

The decision that mattered Oxendale's own rollout playbook says ship every barn a customer buys, all at once, as fast as the hardware allows. Barn 3 and Barn 4 only stayed on the old system because the install crew ran short of ear tags on delivery day, an accident, not a plan. Nobody had ever written "hold a comparison group back on purpose" into the playbook, because for two years, nobody had needed one yet.

What I would leave alone: the dry cows and the pre-fresh heifers. They aren't tagged yet, Byrewell's calibration is built off a lactating cow's own baseline, and the old twice-daily walk still covers them exactly as it always did. Nothing changed for them, so there's no measurement problem to go solve there.

The lesson: a workflow that disappears doesn't take its evidence with it, if you plan for that on day one. It only does if you wait until someone asks, and find out the accident that would have saved you already closed four months ago.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why the scariest number Byrewell ever produced, a hundred and eighteen flagged cows in one day, never actually mattered.

The clipboard hung on a nail by Barn 1's door for eleven years before anyone took it down. Bevan Norwood used it the way you use a hand you trust, without looking at it much, ticking a box next to any cow that seemed a little off, twice a day, since before Rhydian Ockendon had ever set foot on the farm.

Rhydian joined Oxendale Analytics to build Byrewell, and joined Thistlemoor's rollout the week the first ear tags went on in Barn 1. The early weeks were loud in the way new tools always are: overrides, questions, techs walking out to a flag and finding a perfectly fine cow chewing her cud in the sun. But by week six, the flagged list had started catching things days before anyone would have seen them by eye, a cow going a little quiet two mornings before her temperature would have said anything at all. Bevan trusted it fastest of anyone on the crew, because Bevan kept being the one who was right that the flag was worth walking out for.

Two months in, a new hand on the crew, still learning which end of a cow to approach first, asked Bevan a question mid-walk in Barn 3, which was still on the old system. "Why are we still doing this by hand here, if the tags work?" Bevan didn't have a sharp answer. Later that same week, the new hand asked Rhydian something harder, standing by the tablet mounted at Barn 1's parlor exit. "If nobody's walking Barn 1 anymore, how do we even know Byrewell's catching more than Bevan would have?" Rhydian didn't have an answer either. Not a real one. "We think so" isn't a number.

That was the whole problem, and it hadn't shown up as a crisis, a bad flag, an angry vet bill, anything with a name. It showed up as a question a new hand asked, because nobody had told them not to. Bevan's twenty-minute walk through Barn 1 didn't happen anymore. There was nothing left in that barn to time, count, or compare Byrewell against, and pretending otherwise, building some kind of "time saved" slide out of nothing, would have been making up a number to answer a question nobody could actually check.

What Rhydian actually had, by luck rather than plan: Barn 3 and Barn 4 hadn't switched yet. Oxendale's install crew had only shipped enough ear tags for half the herd on delivery day, so Bevan kept walking Barn 3 and Barn 4 the old way, for no reason anyone had chosen on purpose. Rhydian saw what that accident was actually worth: a real, live "before," sitting right next door to the "after," for as long as it lasted. Rhydian got the vet and both barns' techs logging the same thing on both sides, the day a cow's problem actually started, dated off her own milk-yield curve, and the day someone treated her. Not minutes. Not flags. Days.

Then, in month three, a heat wave sat over the valley for four straight days. Byrewell flagged 118 cows in Barn 1 and Barn 2 on the worst of it, more than three times the usual thirty-four. For about six hours, before any of them got checked, that number alone would have read like the whole system falling apart.

Hand sketched decision tree titled Does the case survive the heat-wave day. Root box reads Heat wave day, 118 cows flagged at once, branching into two outcomes: if alert count were the score, leading to reads like the model crying wolf, and checked against the outcome anchor, leading to 3 of 118 needed care, case holds.
The same 118 flagged cows read as either a crisis or a normal Tuesday, depending entirely on what number the case was actually built on.

Techs worked the whole list by afternoon. Three cows needed anything at all. The rest were just hot, the same as every cow in Barn 3 and Barn 4 that day, walked or not.

The number that would have sunk the whole case was never the number the case was built on.

Four months after Barn 1 and 2 switched, Barn 3 and 4 switched too, hardware finally caught up, and the comparison closed for good. By then Rhydian had it: a cow in Byrewell's barns got treated by day 1.2 on average. A cow in the barns still being walked got treated by day 3.1, close enough to Thistlemoor's own years-long baseline of 3.2 that nobody could say the comparison barns were behaving strangely. Mastitis that turned severe, needing more than routine treatment: 4.5 percent of cases in Byrewell's barns, 9.8 percent in the barns still being walked.

The decision Rhydian would take back traced to Oxendale's own rollout playbook, written two years before Thistlemoor ever signed. It said: ship every barn a customer buys, all at once, as fast as the hardware allows. That made sense when the whole job was proving Byrewell worked at all, and speed mattered more than anything else. It stopped making sense the day a customer needed to prove Byrewell was worth keeping, and the fastest possible rollout had already spent the one thing that question needed: something left unswitched to check against.

Run that playbook meeting again, with one line added: hold back a real comparison group on purpose, for every new customer, not just the ones who get lucky with a hardware delay. Same heat wave hits Thistlemoor in month three. Same 118 cows flagged in a day. This time, when Percy Cathmore asks for the number, Rhydian isn't explaining a lucky accident. There's a real "before" sitting in Barn 3 and Barn 4 because the playbook put it there.

One design proves a workflow worked by how fast it disappeared. The other proves it by what's still standing next to it, on purpose, long enough to check.

What Rhydian would tell their past self, back when that rollout playbook first got written: speed answers "does it work." It never answers "was it worth it," and those two questions need completely different things left standing when the dust settles.

Hand sketched full page metaphor scene titled Two rulers, only one of them still works. Left panel, a gauge icon labeled MINUTES PER CHECK, caption the walk that stopped happening. Right panel, a document icon labeled DAYS TO TREATMENT, caption the outcome the walk was for.
The whole answer to this question, in one picture. One ruler broke the day the walk stopped. The other one never depended on the walk existing at all.

SPARK, for a workflow that left no ruler behind

Not a way to dress up "measure the outcome" in five letters. SPARK forces the actual comparison you'd build, and makes you prove it survives the one day it looked like it had failed.

SSituation. What happens today, without this design.
Before Byrewell, checking a cow meant a person walking past her, twice a day, in every barn. Once Byrewell runs a barn, that walk stops completely. There's no "minutes per check" or "percent of herd walked" left to measure it against, because the thing those numbers described doesn't happen there anymore.
Say what's actually gone before naming a fix, or the anchor sounds like a nice number instead of the only number left standing.
PPayoff. The habit worth building.
Not "better reporting." A team that can prove Byrewell's real impact using a number that means the same thing whether a person or a sensor is doing the watching, so nobody ever has to fake a "before" just because the old workflow is gone.
Name the habit, not the mood. A number you can rebuild for the next barn, the next customer, is a habit. A one-off explanation isn't.
AAnchor. The actual thing you'd measure.
Days from when a cow's real problem started, dated by vet records and her own milk-yield curve, not by when someone noticed, to when she gets treated. Build it from Barn 3 and Barn 4, which stayed on the old walk for four more months after Barn 1 and Barn 2 switched, on purpose, so a live comparison group exists the whole time.
This is the concrete answer to the question. Everything else in the framework exists to protect it.
Hand sketched labeled parts diagram titled The comparison the rollout kept, almost by accident. A central box labeled Thistlemoor, 4 barns, with four labeled callouts around it: Barn 1 and 2, Byrewell live. Barn 3 and 4, still walked. Same outcome logged both sides. Four-month comparison window.
Not a bigger sensor. A different shape of rollout: two barns changed, two barns held still on purpose, and one outcome measured the same way on both sides.
RRisk. What breaks the first time it's tested.
A heat wave in month three flags 118 cows in a single day, more than three times normal. Read as an alert count, that looks like the model panicking. The anchor never counted alerts, so a mass false-alarm day, checked and mostly cleared within hours, doesn't touch the real case.
Design the anchor against this specific risk, not a generic "someone might doubt the sensor."
KKeep out. What doesn't ship on day one.
No simulated version of what Barn 1's walk "would have" caught. No blending Barn 1 and 2's early numbers into Barn 3 and 4's, before the window closes, just to get one clean figure sooner. No calling the case finished the day Barn 3 and 4 switch too, since that's exactly when the comparison group disappears for good.
Naming the shortcuts you're refusing is what makes "the old workflow is gone" sound like discipline instead of an excuse.
Hand sketched icon list titled What Rhydian refused to build. Three rows: a question mark icon, a simulated Barn 1 walk, backdated. A funnel icon, one bad day standing in for the case. A scale icon, the real Barn 3 and 4 log instead.
Three shortcuts that would each, quietly, make this quarter's number look cleaner and next quarter's harder to trust.

Three things worth stating directly, since this is where the real judgment sits. The alternative Percy's team pushed for, and Rhydian turned down, was a backdated simulation of what Barn 1's old walk "would have" caught while Byrewell was running there. It lost, because it would use the same kind of model reasoning being tested to invent the very comparison meant to check it, a number nobody could independently verify. The AI-specific failure worth naming is distribution shift under heat stress: a hot, humid day drops rumination and activity herd-wide, and Byrewell's per-cow baseline didn't yet know to expect that, so it read a whole barn's normal heat response as 118 separate emergencies. The guardrail already in the design is a person: every flag gets checked by a tech before any cow is treated, so a mass false-alarm day gets caught by a human within hours, not shipped as a diagnosis. And the trade-off is real and stated plainly: Byrewell costs Thistlemoor about two more hours of tech time a day in Barn 1 and Barn 2 than the old walk ever did, because catching illness two days earlier means checking a longer list, most of it false, rather than trusting a five-minute glance the way the old walk did. That's accepted on purpose, because two hours of checking beats losing a cow to mastitis nobody saw coming for three more days.

And if you want to be sure it really works, try it somewhere else

Same five letters, a building's mechanical room instead of a cattle barn, and this time the thing that stopped happening was a technician's quarterly walk-through with a clipboard of his own.

Hand sketched comparison diagram titled Same wait, a different old workflow. Left panel, a barn icon labeled Thistlemoor Dairy, caption a barn walk, twice a day, by hand. Right panel, a gauge icon labeled Sandhurst Property Group, caption a mechanical room, checked once a quarter.
Same shape of gap. A dairy barn waits on a walk that stopped. A building waits on an inspection that stopped, for the exact same reason.

Ductgauge is Emberlain Systems' continuous monitor for commercial building HVAC systems: temperature drift, vibration signature, refrigerant pressure, and filter differential pressure, read all day off equipment in Sandhurst Property Group's 40 buildings. Yestin Garrowby owns its product story. Ductgauge replaced the quarterly manual inspection completely in 22 of Sandhurst's buildings, phase one, simply because that's as many sensor kits as the install crew had ready on delivery day. The other 18 buildings, phase two, kept their quarterly inspection for four more months. Rosanwyn Sillitoe, Sandhurst's ops director, wants to know at the portfolio review whether Ductgauge is worth expanding to the rest of the buildings, or worth renewing at all.

The decision Emberlain's team would take back Emberlain's own rollout playbook, same as Oxendale's, never reserved a comparison group on purpose. Phase two only existed because of a hardware delivery limit, not a measurement plan. Yestin's team is now writing "hold back a comparison group for the first rollout window" into the standard playbook, so the next customer doesn't need to get lucky with a shipping delay to have one.

Same rank, different lever, mapped straight onto SPARK: the situation is the same shape, a quarterly inspection that simply stops happening once Ductgauge covers a building, leaving no "minutes per visit" behind. The payoff is the same trust: Rosanwyn accepts a real number instead of a guess, because it's built the same way whether a technician or a sensor is watching. The anchor is the same idea, aimed at a different outcome: unplanned major equipment failures avoided, not inspection minutes, measured against the 18 buildings still on the old quarterly schedule during the same four months. The risk is the same shape too: a cold snap in month two triggered more filter-pressure alerts than usual across phase-one buildings, and if alert count were the score, that reads like Ductgauge crying wolf the same way the heat wave did at Thistlemoor. And what stays out is the same discipline: no simulating what a quarterly inspection would have found in a phase-one building, no calling the case closed once phase two switches over too.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: anchor to the outcome the old process protected, hold back a real comparison group, refuse to fake the missing "before."
Cost: no budget to hold anything back this quarter. Fall back to an external record instead, warranty claims, insurer inspection logs, anything that exists independent of either workflow, rather than simulating the one that's gone.
The model got better, for real: say Byrewell's detection lag drops from 1.2 days to 0.6. The anchor doesn't change. A faster catch just makes the anchor number bigger, it doesn't excuse skipping the comparison group that makes the number believable.

Where people run it wrong.
They compare the new alert count straight to the old check count, two units that were never the same thing.
They let staff's memory of "the old way barely caught anything" stand in for a real logged comparison.
They let one loud day, all flags or none, decide the story before the real comparison window closes.

How to use it live. Ask, before naming any number: "what did the old process actually protect, not what did it do." That question alone tells you what to measure once the old process itself is gone.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits this question, and what's its job?
Tap to flip
ANSWER
SPARK: design the actual measurement you'd hand someone before you build it, situation, payoff, anchor, risk, keep out. Fits here because there's no clean before/after left, so designing the measurement itself is the real answer.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Rhydian Ockendon, who owns Byrewell's product story at Oxendale Analytics, and has to prove impact for a workflow that no longer exists to compare against.
3 · THE SITUATION
What's actually gone, and can't be compared to anymore?
Tap to flip
ANSWER
The twice-daily herd walk in Barn 1 and Barn 2. There's no "minutes per check" left, because nobody checks that way there anymore.
4 · THE ANCHOR
What's the concrete anchor, the actual thing Rhydian measures?
Tap to flip
ANSWER
Days from illness onset to treatment, not workflow minutes, built from Barn 3 and Barn 4, held on the old walk for four more months on purpose so a real comparison group exists.
5 · THE REJECTED ALTERNATIVE
What did Percy's team ask for instead, and why did Rhydian refuse?
Tap to flip
ANSWER
A simulated, backdated version of what Barn 1's walk "would have" caught. Refused because it would fabricate the comparison out of the same kind of model logic being tested.
6 · THE NUMBER
Fill in the blank: Byrewell barns caught illness by day ___ on average; barns still walking caught it by day ___, over a ___-month window.
Tap to flip
ANSWER
1.2 versus 3.1, over a four-month comparison window.
7 · THE REPLAY
Run the original rollout meeting again, with staggering built in as policy. What changes?
Tap to flip
ANSWER
Every future customer gets a real comparison group by design, not by the luck of how many sensor tags happened to ship on install day.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent anchor?
Tap to flip
ANSWER
Ductgauge, Emberlain Systems' continuous HVAC monitor for Sandhurst Property Group's buildings. The equivalent anchor is unplanned equipment failures avoided, measured against buildings still on the old quarterly inspection.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these numbers means the same thing before and after Byrewell arrived, making it the safe one to anchor the whole case to?
  • A. Minutes spent walking the barn
  • B. Days from illness onset to treatment
  • C. Number of cows flagged per day
  • D. Percent of the herd walked each morning
Show hint
Check the A step in the framework recap.
Show answer
B. It's measured the same way whether a person or a sensor does the watching. The other three only exist in one of the two workflows, not both.
Fill in the blank
2. Byrewell flags a cow when ___ of its ___ tracked signals cross the line and stay crossed for ___ hours straight.
Show hint
Look at the flow diagram in "Let's learn."
Show answer
2 of 3 signals, held for 4 hours. One odd reading alone isn't enough. Two signals agreeing and staying that way is what makes a flag.
True or false
3. True or false: once Barn 3 and Barn 4 switched to Byrewell too, in month four, Rhydian still had an ongoing live comparison group to check new numbers against.
  • True
  • False
Show hint
Check the K step, keep out, in the framework recap.
Show answer
False. After month four, all four barns run Byrewell. The case has to stand on what got measured during the four-month window, not on a comparison that keeps renewing itself.
Short answer, name the rejected alternative
4. What did Percy's team ask Rhydian to build instead of relying on Barn 3 and Barn 4's real numbers, and why did Rhydian refuse?
Show hint
Look at stage 6 of the walkthrough.
Show answer
Model answer: A simulated, backdated version of what Barn 1 would have caught if the walk had kept running. Refused because it would build the very comparison being tested out of the same kind of model guesswork Byrewell itself uses, a number nobody could independently check.
Short answer, apply it yourself
5. Think of a tool at your own job, or one you use, that replaced a scheduled routine with something continuous. What discrete unit disappeared, and what real outcome could you measure instead?
Show hint
Look for the thing that used to happen on a schedule, and ask what it was actually protecting.
Show answer
Model answer: A team that used to run a weekly manual security review switched to a tool that scans continuously. "Reviews per week" disappeared. What could still be measured: days between a real vulnerability appearing and it getting patched, checked against a few systems kept on manual review for a while as the comparison group.
Short answer, work the number
6. Byrewell's false-positive follow-up rate started at 74 percent in week 1 and settled at 58 percent by week 10. At the usual 34 flags a day, roughly how many fewer wasted checks a day did settling at 58 percent save, compared to staying at 74 percent?
Show hint
Multiply 34 by each percentage, then find the difference.
Show answer
About 5 fewer wasted checks a day. 74 percent of 34 is about 25; 58 percent of 34 is about 20. That's roughly 30 minutes of tech time back every day, at six minutes a check, just from the model settling in.
Before you close the answer
Why this works
Tests whether you'll invent a comparison to make the answer easier, or design a real one when the workflow that used to provide it is gone. Most candidates either fake a "before" number or give up and say it's too hard to measure.
Follow-up traps
"Why not just compare Barn 1's numbers now to Barn 1's own numbers from before Byrewell arrived?" Response: too much else moved in that same stretch of time, the season, the herd's age structure, whatever else changed at Thistlemoor, so any gap could be Byrewell or could be none of it. A live comparison group, not a comparison to your own past, is what makes the number defensible.

"Isn't 118 flagged cows in one day proof the model has a real problem?" Response: only if alert count were the metric. It isn't. Only 3 of the 118 needed care, and the case never rested on how many cows got flagged on any single day.
If pressed
The fix that actually shipped after that heat-wave day: on any day the barn's humidity index crosses a set point, Byrewell widens its per-cow threshold band and requires two separate confirmations six hours apart before a flag counts, instead of one. That's what took the next heat wave's flagged count from 118 down to 29.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more