Artifact critiqueAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #13
Critique an ROI model that assumes every AI-completed task would otherwise have been done by a human.
BOUND · camera-based PPE compliance monitoring on construction sites
Corbelwatch watches every active camera on a Fennborough Construction Group job site for a missing hard hat, a missing hi-vis vest, missing eye protection, or a harness that isn't clipped in at height. Malachi Ekwueme owns what its ROI actually claims. For two years the slide said the same thing: Corbelwatch avoids nine million dollars a year in EHS coordinator labor, one dollar for every flag it raises. Then Graciela Basurto, Fennborough's CFO, went looking for where that nine million had actually shown up in the safety budget, and couldn't find it.
The direct answer
Stop treating every Corbelwatch flag as a catch a person would have made. Ground the model in Fennborough's own shadow-audit data: EHS coordinators independently caught only 22% to 38% of what the cameras flagged, because rounds are periodic and a real slice of what gets flagged, occluded harness clips, four-second lapses, isn't something one person standing in one spot could ever have seen. Rebuild the ROI on that real range, about $2.0 million to $3.4 million a year instead of the assumed $9 million, and Corbelwatch still returns 230% to 470%, an honest number instead of a fantasy 1,400%.
Do this, in order
Rebuild the ROI on the real, measured substitution rate, 22% to 38%, not the assumed 100%.Why: the fully-assumed model overstates avoided labor by three to four times, and that's the entire flaw the 1,400% figure was built on.
Split every flag by whether a human could ever have caught it, not just whether one happened to.Why: some flags come from fusing two camera angles through scaffold framing, a detection no single vantage point could make, no matter how often a round ran.
Value the human-uncatchable flags as injury avoidance, not as labor substitution.Why: they still carry real safety value, just not the value the naive model claimed for them, zeroing them out would undercount just as badly the other way.
Use a range on the substitution rate, not one point number.Why: 22% at a scaffold-heavy site and 38% at an open slab site are both real, a single blended figure would be wrong for whichever site actually matters.
Revisit the substitution rate every time Corbelwatch's confidence threshold changes.Why: loosening the threshold to catch more borderline cases is good for safety, but it mechanically shrinks the real substitution share, so the gap between naive and honest keeps growing.
Check the corrected number against something outside the ROI math itself.Why: comparing it to the average cost of one avoided injury claim proves the case for running Corbelwatch survives even without the labor-substitution argument at all.
How to answer this, stage by stage
Nobody is grading whether you can say "not every task would have been done by hand." They're grading whether you can turn that into a real number, and whether the reason you give for the discount is about the model, not just about people being slow.
1
Anchor it to one real feature and one real number owner
Say it like this
"Let's ground this in one feature. Corbelwatch is Fennborough Construction Group's camera system, it watches every active job site for hard hats, hi-vis vests, eye protection, and harness tie-off at height. Malachi Ekwueme owns what its ROI actually claims, and that's the number I want to rebuild."
Why this works
A generic "not every task transfers" answer stays a slogan. One real feature keeps every number checkable.
2
Name the naive assumption out loud before touching any math
Say it like this
"The ROI on the slide assumes every single flagged violation is a catch an EHS coordinator would otherwise have made by hand. That's the whole model. One number, a hundred percent substitution, and I don't believe it."
Why this works
Names the actual flaw plainly before any figure appears, which is literally the question being asked.
3
Break the equation into its real terms
Say it like this
"The naive value is flagged violations, times a hundred percent, times the cost of a human catch. The honest value is flagged violations, times the real substitution rate, times that same cost, plus whatever the flags a person could never have caught are worth as injury avoidance, a separate number entirely."
Why this works
This is BOUND's break-it-down step, said before any arithmetic, so a wrong number can't hide inside vague language.
4
Give the decision, in one breath
Say it like this
"Here's what I'd do. Rebuild the ROI on the real substitution rate, 22 to 38 percent, from an actual shadow audit, not the assumed hundred. That takes the labor-value side from $9 million down to somewhere between $2.0 and $3.4 million a year, and Corbelwatch still returns 230 to 470 percent. Real, and still good."
Why this works
This is the direct answer, said plainly before any story about how the naive assumption got made.
5
Own each number, and say where it came from
Say it like this
"180,000 flags a year, from Corbelwatch's own logs. $50 a catch, that's thirty minutes of a fully loaded EHS coordinator's time to spot it, walk over, correct it, and log it. $600,000 a year to run Corbelwatch, cameras, compute, alert triage, quarterly retraining, on-call. And 22 to 38 percent substitution, from a two-week shadow audit we ran at two sites, one scaffold-heavy, one open slab, coordinators doing their normal rounds while the cameras logged everything in parallel, unfiltered."
Why this works
Every figure traces to something a listener could go check themselves, not a guess dressed up as data.
6
Tie the discount to the model, not just to people being slow
Say it like this
"It's not only that rounds happen every couple hours instead of continuously. A real slice of what Corbelwatch flags, a harness clip only visible by fusing two camera angles through scaffold framing, a hard hat missing for four seconds while someone reaches into a bin, isn't something one person standing in one spot could ever have seen, no matter how often they walked past. Those flags were never a labor-cost question. They're a coverage question, and it gets worse, not better, as we loosen the confidence threshold to catch more of them."
Why this works
This is the actual AI-specific judgment the question is testing: the discount comes from what the model's own detection footprint can see, not a generic "people are slower" excuse.
7
Sanity check against something outside the ROI math itself
Say it like this
"Even the low end, $2.0 million a year against a $600,000 run cost, only needs to prevent about sixteen PPE-related injury claims a year to pay for itself on injury cost alone, at Fennborough's own average claim of $38,000. We log more near-misses than that most quarters. So this was never a reason to slow down Corbelwatch. It was a reason to stop quoting a number nobody could defend."
Why this works
Proves the corrected number survives a check that doesn't just reuse its own arithmetic.
8
Name the assumption that would move it most, then close
Say it like this
"If one assumption here is going to move the picture the most, it's the substitution rate itself, and it isn't staying still: every time we loosen Corbelwatch's confidence threshold to catch more borderline cases, that rate drops further, because more of what gets flagged falls into the class no human vantage point could reach. So: real substitution rate, not the assumed hundred, split labor value from injury-avoidance value, and revisit the rate every time the threshold changes."
Why this works
Closing on the thing to actually watch is what makes this sound rehearsed, not like a story that trailed off.
Let's learn
For two years, Malachi Ekwueme's quarterly slide said the same thing: Corbelwatch avoids nine million dollars a year in EHS coordinator labor. Nobody had thought to ask what fraction of that labor a coordinator would actually have spent.
Corbelwatch is Fennborough Construction Group's camera system. It watches every active job site, all shift, for a missing hard hat, a missing hi-vis vest, missing eye protection, or a harness that isn't clipped in above six feet.
Before Corbelwatch, Fennborough ran four EHS coordinators across fourteen sites, each site getting three to four timed walking rounds a shift. Nobody had ever measured exactly what fraction of real violations those rounds actually caught. The best independent guess, from a safety consultant's estimate two years back: about three in ten.
Before Corbelwatch, catching a real violation meant a coordinator's round happening to line up with the moment it was visible. Most of the time, it didn't.
With Corbelwatch running, every camera stays on for the whole shift, at every site. It doesn't blink between rounds. Fennborough's cameras log about 180,000 flagged violations a year, at roughly a cent and a half each to run through the model.
Knowledge spark: what actually counts as a Corbelwatch flag?
A frame, or a short run of frames, where the model's confidence that a required piece of PPE is missing clears its threshold. Some flags are obvious, a hard hat missing for ten minutes in plain view. Others come from fusing two camera angles through scaffold framing to resolve a harness clip no single angle could confirm alone.
Cameras don't choose when to look. A round does, four times a shift, down one path.
Here's the turn. The extra flags aren't the problem, more visibility was the whole point of building Corbelwatch. The problem is what the ROI slide did with them next: it took all 180,000 and multiplied by the value of a human catch, $50 each, as if a walking round would eventually have found every single one. It wouldn't have. Rounds happen every couple hours. A harness clip fused from two camera angles through scaffold framing was never going to be caught by someone standing in one spot.
The extra flags weren't the problem. What the model did with all of them was.
Naive labor value vs. the honest, measured value
Naive labor value assumedCorrected labor value, with range
The naive model counted all $9.0 million as avoided labor. About $6.3 million of that was never a catch a person could have made, no matter how many rounds ran. The honest range, $2.0 million to $3.4 million, is what's left, plus a separate, smaller injury-avoidance number for the rest.
What it costs at its worst: this spring, while planning two new site openings, Fennborough's regional leadership nearly held EHS coordinator headcount flat instead of adding two more, reasoning that Corbelwatch had already absorbed most of a coordinator's job. Graciela Basurto, reviewing the safety budget line by line before signing off, noticed something that didn't add up: if Corbelwatch had genuinely avoided nine million dollars in coordinator labor, overtime pay should have fallen by a lot more than the $180,000 it actually had. She paused the headcount decision and asked Malachi a plain question: "Whose number is this? Where did it actually show up?"
Two real sites, two real catch rates. The naive model's 100% never sat anywhere near either one.
The choice that mattered
Corbelwatch's ROI sheet set its substitution assumption to 100% during the pilot at Vesper Point, one site, sixty workers, one coordinator who was practically on that site full time. A shadow measurement that spring actually put pilot-era substitution close to 92%, nearly true at that scale. Nobody wrote down a reason to come back and check it as Fennborough scaled to fourteen sites and four coordinators doing timed rounds instead of one coordinator standing there all day.
What I would leave alone: the small set of flags for a hard hat missing at ground level for ten minutes or more, in plain, unobstructed view. Those really would almost always get caught by a round eventually, they aren't time-window sensitive or occlusion-dependent, so there's no reason to discount them the way the scaffold-level flags need discounting.
The lesson: a substitution rate isn't a constant you set once. It's a coverage claim, and it needs to be measured again every time the thing doing the coverage, four coordinators on rounds, or a camera's confidence threshold, actually changes.
Now here is the same thing as a story
The short version is above, for the room. Read this one when you want to feel why a number that assumes a person was always going to be there is the whole mistake.
Malachi Ekwueme has owned Corbelwatch's numbers since the pilot, back when it watched one site, Vesper Point, sixty workers, one EHS coordinator who was on that site nearly every working hour. The ROI math from that spring was simple: value avoided, divided by what it cost to build. $9,000 in avoided coordinator overtime that first quarter, divided by a $6,000 pilot build. 150%. Genuinely close to true, at that scale, with that coordinator standing right there most of the day.
For a year, nothing about that math needed to change. Fennborough had one site running Corbelwatch. The pilot coordinator really did catch almost everything the cameras also caught, because she really was almost always present. The number held up because the assumption behind it, that a person would have been there anyway, happened to be nearly true.
It stopped being true in three ordinary steps, and none of them looked like a mistake. Fennborough rolled Corbelwatch out to a second site, then a fifth, then all fourteen, and staffed EHS coverage the only sane way a construction company staffs it: four coordinators, sharing fourteen sites, each site getting three or four timed rounds instead of one person standing there the whole day. Flagged violations climbed from a few hundred a month to 180,000 a year. Every quarter, someone updated the flag count on the slide. Nobody ever updated the assumption sitting underneath it: that each of those flags was still a catch a person would otherwise have made, the way the Vesper Point coordinator once nearly did.
Some of what Corbelwatch flags was never a question of how often a round ran. It's a detection no single person, standing anywhere, could have made alone.
Then came an ordinary Tuesday in a budget meeting, not a crisis.
Graciela Basurto was signing off on two new site openings, and the plan in front of her held EHS coordinator headcount flat instead of adding the two coordinators those sites would normally need. The logic on the page was Corbelwatch's own ROI slide: nine million dollars in avoided labor implied coordinators had plenty of slack to absorb more sites. Graciela does this for a living, though, and she does one thing before signing anything: she checks whether a claimed saving shows up anywhere real. Coordinator overtime had dropped, but only by $180,000. Not the kind of drop nine million dollars in avoided labor should have left behind.
"Whose number is this," she asked Malachi, in the flat tone that isn't an accusation yet. "This one's been on the slide for two years. Has anyone ever actually measured it?"
Nobody had lied about the number. Nobody had ever measured it either.
Malachi didn't have a real answer, and that was the actual problem, not the size of the number itself.
Four points on one line. Three of them came from an actual audit. One of them, the one on the slide for two years, never did.
What Malachi built over the next two weeks was the audit that should have run the day Corbelwatch left its pilot site. He and Emre Yilmaz, Fennborough's site EHS lead, ran two weeks of shadow coverage at two sites chosen on purpose: one scaffold-heavy multi-story build, one open single-level slab job, the two ends of how hard a site is to see across. Coordinators ran their normal rounds. Cameras logged everything in parallel, unfiltered. Then Malachi compared the two lists.
At the open slab site, coordinators independently caught 38% of what the cameras also flagged. At the scaffold-heavy site, harder sightlines, more occlusion, they caught 22%. Neither number was close to a hundred. Neither was a coincidence either: it lined up almost exactly with the old three-in-ten guess for how much of the real violation rate a walking round had ever caught, back before Corbelwatch existed at all. A person can only be in one place at a time, whether or not there's a camera watching too.
The decision Malachi traced it back to sat in a single pricing meeting the week the pilot launched. Someone asked whether the ROI's substitution assumption should be a number that gets re-measured, or a fixed hundred percent baked into the formula. A hundred was faster to ship, and at Vesper Point's scale, with one coordinator on site nearly full time, it was close enough to true that nobody pushed back. Nobody wrote a date to come check it again.
The whole answer to this question in one picture. A flag was never one thing. It was always three, and only one of the three was ever a labor number.
Run that pricing meeting again, with one line added: revisit the substitution assumption every time the site count or the model's confidence threshold changes, not just once at launch. Same rollout to fourteen sites, same budget meeting, same sharp question from Graciela. This time Malachi's answer is immediate: "22 to 38 percent, measured two months ago, here's the audit." The headcount plan doesn't get cut. The two new sites get their coordinators.
One design let "the ROI" mean whatever number a pilot site happened to produce. The other lets it mean whatever the current coverage actually is.
What Malachi would tell himself, back in that first pricing meeting: a hundred percent substitution wasn't a lie. It was just a measurement of one site, with one coordinator standing there nearly full time, that got treated like a law of nature instead of a fact that would stop being true the moment the company grew.
BOUND, or the range a launch-day pilot never earned
Not a story wearing a framework's clothes. This is a substitution-rate estimate that got frozen at a hundred percent when it was almost true once, and BOUND is what turns "not every task transfers" into a number a CFO can actually check.
BBreak it down. What's the actual equation?
The naive value is flagged violations times a hundred percent substitution times the cost of a human catch. The honest value is flagged violations times the real, measured substitution rate times that same cost, plus a separate, smaller number for the flags a person could never have caught, valued as injury avoidance instead of labor avoidance. Those are two different kinds of value and the naive model collapsed them into one.
Saying the two terms apart before any figure appears stops "not every task transfers" from staying a vague hedge.
OOwn the numbers. Where did each one come from?
180,000 flags a year, from Corbelwatch's own logs. $50 a catch, thirty minutes of a fully loaded EHS coordinator's time to spot, correct, and log one violation. $600,000 a year to run Corbelwatch, cameras, inference, alert triage, quarterly retraining, incident response. 22% to 38% substitution, from a two-week shadow audit at two sites chosen for their sightline extremes. This is also where a rejected alternative sits: Malachi considered valuing the human-uncatchable flags at zero, simplest to compute, and turned it down, because they still carry real safety value, just priced as injury avoidance, not labor substitution, zeroing them out would have undercounted just as badly as the naive model overcounted.
Owning a number means saying where it came from and what got turned down instead, not just stating a figure.
UUse a range, not one number.
Flag count and run cost are solid enough to treat as point figures, both come straight from Corbelwatch's own logs and Fennborough's own invoices. Substitution rate is the genuinely uncertain one: 22% at the scaffold-heavy site, 38% at the open slab site. Malachi also turned down applying one flat 30% company-wide, since that would understate risk at exactly the sites where getting it wrong matters most. All in, the honest labor value is $2.0 million to $3.4 million a year, not one clean $2.7 million.
The whole case for measuring instead of assuming lives inside that range.
What moves the corrected ROI most, if the assumption is wrong
Substitution rate, widest swingRun costCost per catch and flag volume
Moving the substitution rate across its real 22 to 38 percent range swings the ROI by 240 points, far more than a realistic move in run cost, cost per catch, or flag volume. It's the assumption actually worth watching.
NNail the sanity check. Does the number survive being compared to something real?
Even the low end of the corrected range, $2.0 million a year, only needs to prevent about sixteen PPE-related injury claims a year to cover Corbelwatch's $600,000 run cost on injury-avoidance value alone, at Fennborough's own average claim cost of $38,000. Fennborough logs more near-misses than that most single quarters. That's the check that matters: this was never a reason to slow Corbelwatch down. The number that should have worried Malachi two years earlier was a different one, a hundred percent, quietly sitting on a slide that kept getting more impressive as the flag count climbed.
The hardest step, and the one most answers skip. A number that looks calm can still be built on an assumption nobody ever checked.
DDirection. Which single assumption would move the answer most?
By raw dollars, substitution rate swings it hardest, moving the ROI 240 points across its real range, more than run cost, cost per catch, or flag volume combined. And it isn't a static number: as Corbelwatch's confidence threshold gets tuned looser to catch more borderline, partially occluded cases, a genuinely good move for safety, the flagged count grows, but a bigger share of that growth falls into the class no human vantage point could ever reach. Every safety win on the threshold quietly widens the naive model's error.
Naming the assumption that's both the biggest swing and actively moving, not just the biggest line item, is what a good estimator does that a bad one skips.
Three things worth stating directly, since this is where the real judgment sits. The trade-off being accepted: loosening Corbelwatch's confidence threshold buys real safety coverage, catching borderline cases a tighter threshold would miss, but it buys that recall by pulling in exactly the class of flags a labor-substitution number can't explain, so the honest substitution share keeps shrinking even as the flag count climbs. The AI-specific failure mode worth naming separately is false-positive drift: a rolled-up sleeve or a shadow across a hard hat's brim can read as a violation under low winter light the way it wouldn't in summer glare, and the guardrail is a precision audit against a golden set stratified by season and site type, not one blended precision number, so a seasonal false-positive creep doesn't hide inside an average that still looks fine. And the alternative genuinely considered and rejected: freezing Corbelwatch's threshold at its pilot-era setting to keep the substitution math simple, turned down because a frozen threshold misses exactly the borderline, high-consequence cases, an unclipped harness at height, that the looser setting exists to catch.
And if you want to be sure it really works, try it somewhere else
Same five letters, a trucking fleet instead of a job site, and this time the thing no one vantage point could see wasn't a harness clip. It was a driver's eyes.
Wakewatch is Thistlegate Freight Co's in-cab driver-monitoring system. It watches for eyes-closed events, head-nods, and phone glances across the fleet, all shift, on every truck. Thistlegate runs it across 620 trucks, logging about 46,000 flagged fatigue and distraction events a year.
Same BOUND, a different fleet, a different thing no single human vantage point could have caught in time.
Radim Tavarez owns Wakewatch's numbers, and Thistlegate's original ROI sheet had the same shape as Corbelwatch's: every flagged event valued as a catch a dispatcher would otherwise have made by reviewing footage. The reality is starker here than at Fennborough. A dispatcher reviewing live feeds can realistically watch a handful of trucks at a time out of 620, and most fatigue events last under three seconds, well under what a spot-check schedule could ever intersect. Radim's own shadow measurement put dispatcher substitution at just 8% to 15%, far lower than Corbelwatch's 22% to 38%, because a truck cab has no equivalent of a walking round at all.
The decision Radim would take back
Wakewatch's ROI sheet, like Corbelwatch's, priced every flag as a dispatcher catch, set during a pilot on forty trucks where one dedicated safety reviewer really did watch most feeds live. Nobody revisited the assumption when the fleet scaled to 620 trucks and reviewer time stayed flat.
Same rank, different lever: Thistlegate's fix isn't a bigger dispatch team. It's the same substitution habit, run on Wakewatch's own numbers, with the lower 8 to 15 percent range driving a much smaller labor-value claim and a much larger share of the flags valued instead as crash-avoidance value, priced against Thistlegate's own claims history.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: don't count every flag as a human catch, measure the real substitution rate, and value the rest as injury or incident avoidance instead of zero.
Cost: there's no time this quarter to run a real shadow audit. Ship the cheap version first, a documented estimate built from last year's near-miss logs, revisited next quarter with real audit data.
The model got better, for real: say Corbelwatch's false-positive rate got cut in half. The corrected ROI improves, but the substitution problem doesn't go away, a more accurate model still isn't the same thing as a model whose every flag was a guaranteed human catch.
Where people run it wrong.
They assume the substitution rate is a hundred percent because "the AI did the work," without ever measuring what a person would have caught in parallel.
They swing too far the other way and price the human-uncatchable flags at zero, when they still carry real, separately priceable safety value.
They set the substitution rate once and never revisit it as the model's threshold, or the scale of the rollout, changes.
How to use it live. Ask one clarifying question before any arithmetic: "Has anyone actually measured what a human process would have caught in parallel, or is this substitution rate just assumed?" That question alone usually tells you whether you're about to defend a real number or a guess wearing one.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: show the arithmetic, own the assumptions. Built for estimation and ROI-modeling questions like this one, not a habit-flip story.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Malachi Ekwueme, who owns Corbelwatch's ROI at Fennborough Construction Group, and inherited a substitution assumption set during a single-site pilot where it was nearly true.
3 · THE BLIND SPOT
What did the naive ROI model assume that never got re-checked?
Tap to flip
ANSWER
That every flagged violation was a catch an EHS coordinator would otherwise have made, a hundred percent substitution, set when one coordinator was nearly always on the one pilot site, never revisited across fourteen sites and four coordinators.
4 · THE EQUATION
What's the naive value, and what's the honest value?
Tap to flip
ANSWER
Naive: flags times 100% times cost per catch. Honest: flags times the real 22 to 38 percent substitution rate times cost per catch, plus a separate, smaller injury-avoidance value for the rest.
5 · THE OLD DECISION
What decision would Malachi take back?
Tap to flip
ANSWER
Baking a fixed 100% substitution rate into the ROI formula at pilot launch, with no plan to re-measure it, decided when the assumption was almost true at one site.
6 · THE NUMBER
Fill in the blank: Corbelwatch flags about ___ violations a year. The naive ROI was ___%. The corrected range is ___% to ___%.
Tap to flip
ANSWER
180,000 flags a year. Naive ROI: 1,400%. Corrected: 230% to 470%, still healthy, just honest.
7 · THE REPLAY
Same budget meeting, new habit, what changes?
Tap to flip
ANSWER
Malachi already has the measured range when Graciela asks. The two new sites keep their planned EHS coordinators instead of the headcount plan getting cut on a number nobody had checked.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the different substitution rate?
Tap to flip
ANSWER
Wakewatch, Thistlegate Freight Co's in-cab driver-monitoring system. Its measured substitution rate is 8% to 15%, lower than Corbelwatch's, because a truck cab has no equivalent of a walking round at all.
Check yourself Score: 0 / 0
True or false
1. True or false: the shadow audit found that Fennborough's EHS coordinators would eventually have caught close to 100% of what Corbelwatch flags, just more slowly.
True
False
Show hint
Check the substitution range from the two-site shadow audit in the direct answer and stage 5.
Show answer
False. Coordinators independently caught only 22% to 38% of what Corbelwatch flagged, not because they were slow, but because a real slice of the flags could never have been caught by one person in one spot, no matter how many rounds ran.
Multiple choice
2. Why can't a human vantage point ever catch some of the violations Corbelwatch flags, no matter how often the rounds run?
A. Coordinators aren't trained on the newest PPE rules.
B. Some flags come from fusing two camera angles through scaffold framing, or resolving a lapse of just a few seconds, a detection no single glance could make.
C. The flagged images are too low-resolution for a person to interpret.
D. Coordinators are instructed to ignore minor violations.
Show hint
Look at stage 6 of the walkthrough, "tie the discount to the model."
Show answer
B. The model's own confidence threshold pulls in detections, multi-camera fusion, sub-visibility-window lapses, that were structurally never a labor-cost question, they're a coverage question only a camera can answer.
Fill in the blank
3. Corbelwatch flags about ___ violations a year, and Fennborough's naive ROI sheet valued every one of them at $___, assuming 100 percent of them would otherwise have been a human catch.
Show hint
Look at stage 5's owned numbers, or the "Let's learn" section.
Show answer
180,000 violations a year, at $50 each. That put the naive labor value at $9.0 million, against a $600,000 run cost, a naive ROI of 1,400%.
Short answer, name the rejected alternative
4. What alternative did Malachi consider instead of building the real substitution-rate range, and why was it rejected?
Show hint
Look at the O step in the framework recap.
Show answer
Model answer: Valuing every flag a human couldn't have caught at zero, the simplest option. Rejected because those flags still carry real safety value, just as injury avoidance instead of labor substitution, zeroing them out would undercount just as badly as the naive model overcounted.
Short answer, apply it yourself
5. Think of an AI product you use, or would pitch, that gets credited with "doing what a person used to do." What's one reason a person wouldn't actually have done all of it?
Show hint
Think about coverage, not effort, would a person even have been present for every case, not just whether they'd have been slower.
Show answer
Model answer: A spam filter gets credited with "catching what a moderator would have caught," but a moderator only ever reviewed a small sample of flagged posts, most spam would simply have gone unreviewed, not caught slowly, so crediting the filter with 100% substitution overstates the labor it actually replaced.
Short answer, work the number
6. If Corbelwatch's confidence threshold loosened tomorrow and flagged volume rose 20%, with almost all the new flags being borderline occlusion cases no human could have caught, what happens to the real substitution rate, and does that mean Corbelwatch got worse?
Show hint
Check the D step's point about the threshold and the trade-off paragraph right after it.
Show answer
The substitution rate falls further, since a bigger share of the growth is human-uncatchable. That's not Corbelwatch getting worse, it's catching more real risk. It just means an even smaller share of the growth should ever be counted as avoided labor, the rest belongs in the injury-avoidance number instead.
Before you close the answer
Why this works
Tests whether you'll treat "not every task transfers to a human" as real, measured arithmetic tied to what the model can actually see, rather than a soft caveat you gesture at and move past.
Follow-up traps
"Isn't $2.7 million still just a guess, since the audit only covered two sites?" Response: it's a bounded, checkable estimate, not a guess dressed as certainty. The two sites were chosen deliberately to span the sightline-complexity range, scaffold-heavy and open slab, and the range gets revisited, not treated as final.
"Why not just count every flag a person couldn't have caught as worth zero, safest assumption?" Response: that undercounts just as badly in the other direction. Those flags still carry real, checkable value, they should be priced as injury avoidance, not deleted from the model entirely.
If pressed
Corbelwatch's harness tie-off detector runs at a 0.61 confidence threshold, calibrated to hold precision above 90% on a golden set stratified by site type and season. Loosen it further and false positives start contaminating the substitution-rate audit itself, since a flag that isn't real can't be "caught" by anyone, human or camera.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.