How do you measure impact when the AI feature changed the workflow entirely?
Byrewell reads three health signals off every cow in Thistlemoor Dairy Cooperative's four barns, all day, and flags the ones worth a second look before anyone would ever see it by eye. Rhydian Ockendon owns its product story at Oxendale Analytics. Byrewell replaced the twice-daily herd walk completely in Barn 1 and Barn 2, and Percy Cathmore, Thistlemoor's general manager, wants a number for it at the co-op's annual member meeting. The thing Byrewell replaced does not happen anymore, so there is nothing left to time it against.
- Anchor the case to days from illness onset to treatment, never to the walk's own minutes or headcount.Why: it is the one number that means the same thing whether a person or a sensor is doing the watching.
- Hold Barn 3 and Barn 4 on the old system for the full comparison window before switching them too.Why: without a barn still being walked somewhere, there is no honest "before" left anywhere in the herd.
- Report the false-positive follow-up rate right next to the detection-lag gain, never the raw flagged-cow count alone.Why: thirty-four flags a day reads like an emergency. It isn't, and the count by itself can't say so.
- Refuse to simulate what Barn 1 "would have" caught if the walk had kept going.Why: that builds a fake comparison out of the same kind of model reasoning being tested.
- Say plainly that Byrewell costs more staff time in the switched barns, not less.Why: an honest trade-off Percy can weigh beats a saving that isn't real.
- Leave the dry-cow and heifer walk exactly as it is.Why: nothing changed for them yet, so there is no measurement problem to go solve there.
How to answer this, stage by stage
Nobody is grading whether you can say "measure the outcome, not the process." They're grading whether you can build a real comparison when the thing everyone would normally compare against has quietly stopped existing.
Let's learn
Byrewell is Oxendale Analytics' health-monitoring system for dairy cows: an ear tag that reads three signals on every cow, all day, plus a camera at the parlor exit that scores how she walks, so a sick cow gets found without anyone walking the barn to look for her.
Before Byrewell, catching a sick cow meant a person seeing it. Bevan Norwood and two techs walked all four of Thistlemoor Dairy Cooperative's barns twice a day, 5:30 in the morning and 4:30 in the afternoon, about twenty minutes a barn, watching how 1,400 cows stood, ate, and moved. On top of the walk, about 15 percent of the herd, roughly 210 cows, got a hand temperature check each day, on a rotation. Even done well, by a team that knew the herd cold, the average sick cow wasn't caught until day 3.2 of whatever was actually wrong with her. Early illness doesn't show on the outside yet.
Byrewell doesn't wait for the outside to change. It reads rumination minutes, activity, and ear-skin temperature off every cow's tag, all the time, plus a gait score every time she passes through the parlor, and flags her the moment two of those three signals drift together and stay that way for four hours straight.
In Barn 1 and Barn 2, where it has been running, the average sick cow gets caught by day 1.2. About thirty-four cows a day land on the flagged list, out of roughly seven hundred cows in those two barns. And the flags don't behave like the old checks did. About 58 percent of them turn out fine when a tech walks out to look, against something close to zero false alarms on the old visual walk, which only ever flagged a cow a trained eye could actually see was off.
Here's the turn. Catching illness two days earlier isn't the hard part to prove. The hard part is that nobody walks Barn 1 or Barn 2 anymore, so there's no "minutes per check" left to measure Byrewell against. The thing it replaced doesn't happen there. A number built the old way, whatever that would even mean now, can't tell Percy Cathmore whether Byrewell is worth what it costs.
At its worst, this gets reported as a shrug: "we think it's working, hard to say by how much exactly," right at the meeting where the co-op decides whether to finish switching Barn 3 and Barn 4, or let Byrewell's contract quietly lapse and go back to walking.
What I would leave alone: the dry cows and the pre-fresh heifers. They aren't tagged yet, Byrewell's calibration is built off a lactating cow's own baseline, and the old twice-daily walk still covers them exactly as it always did. Nothing changed for them, so there's no measurement problem to go solve there.
The lesson: a workflow that disappears doesn't take its evidence with it, if you plan for that on day one. It only does if you wait until someone asks, and find out the accident that would have saved you already closed four months ago.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why the scariest number Byrewell ever produced, a hundred and eighteen flagged cows in one day, never actually mattered.
The clipboard hung on a nail by Barn 1's door for eleven years before anyone took it down. Bevan Norwood used it the way you use a hand you trust, without looking at it much, ticking a box next to any cow that seemed a little off, twice a day, since before Rhydian Ockendon had ever set foot on the farm.
Rhydian joined Oxendale Analytics to build Byrewell, and joined Thistlemoor's rollout the week the first ear tags went on in Barn 1. The early weeks were loud in the way new tools always are: overrides, questions, techs walking out to a flag and finding a perfectly fine cow chewing her cud in the sun. But by week six, the flagged list had started catching things days before anyone would have seen them by eye, a cow going a little quiet two mornings before her temperature would have said anything at all. Bevan trusted it fastest of anyone on the crew, because Bevan kept being the one who was right that the flag was worth walking out for.
Two months in, a new hand on the crew, still learning which end of a cow to approach first, asked Bevan a question mid-walk in Barn 3, which was still on the old system. "Why are we still doing this by hand here, if the tags work?" Bevan didn't have a sharp answer. Later that same week, the new hand asked Rhydian something harder, standing by the tablet mounted at Barn 1's parlor exit. "If nobody's walking Barn 1 anymore, how do we even know Byrewell's catching more than Bevan would have?" Rhydian didn't have an answer either. Not a real one. "We think so" isn't a number.
That was the whole problem, and it hadn't shown up as a crisis, a bad flag, an angry vet bill, anything with a name. It showed up as a question a new hand asked, because nobody had told them not to. Bevan's twenty-minute walk through Barn 1 didn't happen anymore. There was nothing left in that barn to time, count, or compare Byrewell against, and pretending otherwise, building some kind of "time saved" slide out of nothing, would have been making up a number to answer a question nobody could actually check.
What Rhydian actually had, by luck rather than plan: Barn 3 and Barn 4 hadn't switched yet. Oxendale's install crew had only shipped enough ear tags for half the herd on delivery day, so Bevan kept walking Barn 3 and Barn 4 the old way, for no reason anyone had chosen on purpose. Rhydian saw what that accident was actually worth: a real, live "before," sitting right next door to the "after," for as long as it lasted. Rhydian got the vet and both barns' techs logging the same thing on both sides, the day a cow's problem actually started, dated off her own milk-yield curve, and the day someone treated her. Not minutes. Not flags. Days.
Then, in month three, a heat wave sat over the valley for four straight days. Byrewell flagged 118 cows in Barn 1 and Barn 2 on the worst of it, more than three times the usual thirty-four. For about six hours, before any of them got checked, that number alone would have read like the whole system falling apart.
Techs worked the whole list by afternoon. Three cows needed anything at all. The rest were just hot, the same as every cow in Barn 3 and Barn 4 that day, walked or not.
Four months after Barn 1 and 2 switched, Barn 3 and 4 switched too, hardware finally caught up, and the comparison closed for good. By then Rhydian had it: a cow in Byrewell's barns got treated by day 1.2 on average. A cow in the barns still being walked got treated by day 3.1, close enough to Thistlemoor's own years-long baseline of 3.2 that nobody could say the comparison barns were behaving strangely. Mastitis that turned severe, needing more than routine treatment: 4.5 percent of cases in Byrewell's barns, 9.8 percent in the barns still being walked.
The decision Rhydian would take back traced to Oxendale's own rollout playbook, written two years before Thistlemoor ever signed. It said: ship every barn a customer buys, all at once, as fast as the hardware allows. That made sense when the whole job was proving Byrewell worked at all, and speed mattered more than anything else. It stopped making sense the day a customer needed to prove Byrewell was worth keeping, and the fastest possible rollout had already spent the one thing that question needed: something left unswitched to check against.
Run that playbook meeting again, with one line added: hold back a real comparison group on purpose, for every new customer, not just the ones who get lucky with a hardware delay. Same heat wave hits Thistlemoor in month three. Same 118 cows flagged in a day. This time, when Percy Cathmore asks for the number, Rhydian isn't explaining a lucky accident. There's a real "before" sitting in Barn 3 and Barn 4 because the playbook put it there.
One design proves a workflow worked by how fast it disappeared. The other proves it by what's still standing next to it, on purpose, long enough to check.
What Rhydian would tell their past self, back when that rollout playbook first got written: speed answers "does it work." It never answers "was it worth it," and those two questions need completely different things left standing when the dust settles.
SPARK, for a workflow that left no ruler behind
Not a way to dress up "measure the outcome" in five letters. SPARK forces the actual comparison you'd build, and makes you prove it survives the one day it looked like it had failed.
Three things worth stating directly, since this is where the real judgment sits. The alternative Percy's team pushed for, and Rhydian turned down, was a backdated simulation of what Barn 1's old walk "would have" caught while Byrewell was running there. It lost, because it would use the same kind of model reasoning being tested to invent the very comparison meant to check it, a number nobody could independently verify. The AI-specific failure worth naming is distribution shift under heat stress: a hot, humid day drops rumination and activity herd-wide, and Byrewell's per-cow baseline didn't yet know to expect that, so it read a whole barn's normal heat response as 118 separate emergencies. The guardrail already in the design is a person: every flag gets checked by a tech before any cow is treated, so a mass false-alarm day gets caught by a human within hours, not shipped as a diagnosis. And the trade-off is real and stated plainly: Byrewell costs Thistlemoor about two more hours of tech time a day in Barn 1 and Barn 2 than the old walk ever did, because catching illness two days earlier means checking a longer list, most of it false, rather than trusting a five-minute glance the way the old walk did. That's accepted on purpose, because two hours of checking beats losing a cow to mastitis nobody saw coming for three more days.
And if you want to be sure it really works, try it somewhere else
Same five letters, a building's mechanical room instead of a cattle barn, and this time the thing that stopped happening was a technician's quarterly walk-through with a clipboard of his own.
Ductgauge is Emberlain Systems' continuous monitor for commercial building HVAC systems: temperature drift, vibration signature, refrigerant pressure, and filter differential pressure, read all day off equipment in Sandhurst Property Group's 40 buildings. Yestin Garrowby owns its product story. Ductgauge replaced the quarterly manual inspection completely in 22 of Sandhurst's buildings, phase one, simply because that's as many sensor kits as the install crew had ready on delivery day. The other 18 buildings, phase two, kept their quarterly inspection for four more months. Rosanwyn Sillitoe, Sandhurst's ops director, wants to know at the portfolio review whether Ductgauge is worth expanding to the rest of the buildings, or worth renewing at all.
Same rank, different lever, mapped straight onto SPARK: the situation is the same shape, a quarterly inspection that simply stops happening once Ductgauge covers a building, leaving no "minutes per visit" behind. The payoff is the same trust: Rosanwyn accepts a real number instead of a guess, because it's built the same way whether a technician or a sensor is watching. The anchor is the same idea, aimed at a different outcome: unplanned major equipment failures avoided, not inspection minutes, measured against the 18 buildings still on the old quarterly schedule during the same four months. The risk is the same shape too: a cold snap in month two triggered more filter-pressure alerts than usual across phase-one buildings, and if alert count were the score, that reads like Ductgauge crying wolf the same way the heat wave did at Thistlemoor. And what stays out is the same discipline: no simulating what a quarterly inspection would have found in a phase-one building, no calling the case closed once phase two switches over too.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: anchor to the outcome the old process protected, hold back a real comparison group, refuse to fake the missing "before."
Cost: no budget to hold anything back this quarter. Fall back to an external record instead, warranty claims, insurer inspection logs, anything that exists independent of either workflow, rather than simulating the one that's gone.
The model got better, for real: say Byrewell's detection lag drops from 1.2 days to 0.6. The anchor doesn't change. A faster catch just makes the anchor number bigger, it doesn't excuse skipping the comparison group that makes the number believable.
Where people run it wrong.
They compare the new alert count straight to the old check count, two units that were never the same thing.
They let staff's memory of "the old way barely caught anything" stand in for a real logged comparison.
They let one loud day, all flags or none, decide the story before the real comparison window closes.
How to use it live. Ask, before naming any number: "what did the old process actually protect, not what did it do." That question alone tells you what to measure once the old process itself is gone.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't 118 flagged cows in one day proof the model has a real problem?" Response: only if alert count were the metric. It isn't. Only 3 of the 118 needed care, and the case never rested on how many cows got flagged on any single day.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.