CalculationAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #16

What business impact does an AI feature have on churn, and how would you measure it?

LEAD · AI-drafted performance reviews for a retail chain's managers

Candor drafts and summarizes performance reviews inside Fenhollow's people-management platform. Hartsbridge Retail Group, a 140-store home-goods chain, rolled it out to every store manager because Candor was the reason Hartsbridge picked Fenhollow's higher tier over a cheaper competitor two years ago. Chiamaka Adebayo owns Candor's product story at Fenhollow. Franklin Vandenburg runs People at Hartsbridge, and he's the one who has to decide, at renewal, whether Candor is still worth the eighteen percent it adds to the bill.

The direct answer
Track the weekly share of reviews where a manager throws out Candor's draft and writes the review from a blank page instead, not adoption or completion. That reversion rate climbs for months before a renewal decision does, while completion stays perfect the whole time. Tie it to a real dollar, the Candor add-on revenue at risk, and act once the rate crosses 30 percent and holds, instead of waiting for the renewal conversation to tell you.
Do this, in order
  1. Track the weekly rewrite rate, not adoption or completion.Why: completion is mandatory paperwork that stays near 100 percent no matter what. Only the share of drafts a manager throws out tells you whether Candor still earns the price Hartsbridge pays for it.
  2. Tie the rewrite rate to a real dollar, the add-on revenue at risk, before finance asks.Why: a leading indicator nobody has priced is a chart nobody acts on. Pricing it is what turns a metric into a decision.
  3. Set real thresholds and act on them, don't wait for the number to become a renewal conversation.Why: under 15 percent is healthy, 15 to 30 earns a sample audit, over 30 sustained earns a Customer Success call. Each threshold buys weeks the annual renewal date never gives you.
  4. Never let "the feature is switched on" count as "the feature is being used."Why: that's the easiest way this exact metric gets gamed, and it hides a worsening account behind a number that can only tick up.
  5. Only draft when there's enough source material to draft from.Why: a draft built on three thin notes is a guess wearing a paragraph's clothing. The guardrail against a made-up detail is refusing to draft past that point, not catching it after the fact.
  6. Keep the draft-to-final history, for every review, on purpose.Why: without it nobody can compute the rewrite rate at all, which is exactly the gap that let Hartsbridge's number climb for eight months unnoticed.

How to answer this, stage by stage

Nobody is grading whether you can define churn. They're grading whether you can name a number that would have looked fine right up until the morning the renewal call went badly, and say what you'd have watched instead.

1
Anchor it in one account before naming any metric
Say it like this
"Let me ground this in one case. Candor is Fenhollow's AI feature that drafts and summarizes performance reviews. Hartsbridge Retail Group runs it across a hundred and forty stores, and Franklin Vandenburg is the one who has to decide, at renewal, whether it's still worth the eighteen percent premium they're paying for it."
Why this works
Keeps the answer checkable against a real account instead of a hypothetical dashboard.
2
Name the framework and its shape out loud
Say it like this
"I'll run this as LEAD. Link it to the real dollar outcome, find the early signal that moves before that dollar does, name how the metric gets gamed, then say what I'd actually do at each level of it."
Why this works
Signals a plan before the interviewer has to wonder whether one is coming.
3
Reject the obvious metric before proposing the real one
Say it like this
"The tempting metric is adoption, the percent of managers who ever open Candor's draft. It's wrong, because it sits near 95 percent from week one whether the manager keeps a single word or throws the whole thing out. It can't move, so it can't warn you."
Why this works
Shows the reasoning behind the metric choice, not just the metric, which is what actually survives a follow-up question.
4
Give the real link and the real early signal, in one breath
Say it like this
"Here's my link: Candor's own add-on revenue, twenty-seven thousand dollars a year on this account, and what it says about the base contract underneath it. Here's the early signal: the weekly share of reviews where a manager writes from a blank page instead of starting from Candor's draft. On Hartsbridge's account that went from 8 percent to 41 percent over ten months, while on-time completion sat at 97 percent the entire time."
Why this works
This is the direct answer, said in one breath, with the number that proves it isn't a hunch.
5
Name how the metric gets gamed, before someone else finds it for you
Say it like this
"Two ways this number lies to you. One, an early report only looked at the twelve beta customers who chose Candor first, our own champions, who were always going to renew. Two, the retention dashboard counted any account with Candor switched on as 'Candor-active,' even when almost nobody there was keeping the drafts. Both make the number look better than the account actually is."
Why this works
Naming the gaming mode unprompted is the single move that separates a candidate who's used a metric from one who's only heard of it.
6
Say what you'd actually do at each threshold
Say it like this
"Under 15 percent, leave it alone, that's what we sold. Between 15 and 30, pull a sample of the discarded drafts and check them against the manager's own notes for made-up specifics. Past 30 percent, sustained for three weeks, Customer Success gets looped in before the renewal conversation starts, not after."
Why this works
A metric with no attached decision is a chart nobody acts on. This is what makes it a real answer instead of dashboard decoration.
7
Close on the decision, in one breath
Say it like this
"So: don't watch whether the review got submitted, watch whether Candor's own words survived into it. That number moves months before a renewal does, and it's the one Hartsbridge's account was trying to tell us the whole time."
Why this works
Restates the direct answer plainly, so the interviewer leaves with the decision, not just the story behind it.

Let's learn

Candor reads a manager's notes on an employee, logged over a review period, and writes a first draft of that employee's performance review, plus a short summary of themes across the manager's whole team. The manager can keep it, edit it, or ignore it and start over.

Hand sketched numbered icon list titled Before Candor, every review started blank. Three rows: a document icon, 140 managers wrote every review from a blank page. A gauge icon, a full write-up took 45 to 60 minutes per employee. A person icon, busy managers often shrank it to two generic lines.
Before Candor, a Hartsbridge store manager spent 45 to 60 minutes writing each employee's review from nothing.

Before Candor, a Hartsbridge store manager wrote every review of every employee from a blank page, 45 to 60 minutes each. With around nine employees per store and 140 stores, reviews land on a rolling schedule tied to each employee's hire date rather than one big cycle, so roughly 25 to 30 reviews get written across the whole account every week, all year.

With Candor turned on, a manager opens a review and finds a draft already sitting there, built from the notes they logged that period. In the first four weeks after rollout, managers kept the bulk of that draft 92 percent of the time, only 8 percent of reviews got written from scratch instead. Reviews that used to take an hour were taking twelve minutes.

Here's the turn. Ten months later, that 8 percent had become 41 percent. Almost half of every week's reviews were being written as if Candor didn't exist. And the number everyone at Fenhollow was actually watching, the share of reviews submitted on time, never moved. It sat at 97 percent in week one and 97 percent in week forty-four. Completion is mandatory. A manager writes the review whether they trust the draft or not. It just never told anyone which one was happening.

The leading edge, charted: Hartsbridge's weekly reversion rate against the completion rate that hid it
100% 50% 0% completion: 97%, flat watch line: 15% act line: 30% renewal call Wk 1 Wk 14 Wk 31 Wk 44
Weekly reversion rate, blank-page reviewsOn-time completion, unmoved the whole time
The reversion rate crossed the watch line at week 14 and the act line at week 31, thirteen weeks before anyone had a renewal conversation with Franklin. Completion never left 97 percent, which is exactly why it never raised a hand.
Knowledge spark: what's a hallucination, here? Candor sometimes wrote a specific detail, a project name, a task, a number, that never appeared anywhere in the manager's own logged notes. It sounded exactly as confident as the true sentences next to it. The model wasn't lying on purpose. It was filling a gap the way a good guesser fills a gap, and a manager who caught it once stopped trusting the rest of the paragraph too.

What that costs at its worst is not just a worse review. It's a false sense that everything is fine, dressed up as data. When Candor was still in private beta, Fenhollow's first internal report on it looked at the twelve customers who signed up earliest, all of them already Fenhollow's biggest champions. Every one of them renewed. The report said "100 percent retention on Candor accounts, versus 91 percent baseline," which was true and told nobody anything, because those twelve were never going to leave regardless of what Candor's drafts were like.

Hand sketched labeled parts diagram titled What the early Candor report actually counted. A center box with a question mark labeled Candor, active, with four callouts around it: switched on, yes. Draft opened, sometimes. Draft kept, rarely. Counted as retained, yes.
Hartsbridge itself got counted as "Candor-active, retained" for three straight quarters, right through the stretch where its reversion rate climbed past 30 percent, because the dashboard only checked whether the toggle was on.

That's the second, quieter way the number gets gamed. "Candor-active" meant an admin had switched the feature on for the account, nothing about whether the store managers were actually keeping the drafts. An account can only turn a feature on, rarely off, so a metric built that way can only ever look healthy or flat, never honestly bad, even while the real thing it's supposed to stand for is falling apart underneath it.

Candor's add-on costs Hartsbridge $27,000 a year, on top of a $150,000 base platform contract. At renewal, once Franklin actually looked past the completion number, he asked to drop the add-on and keep the base platform. That's $27,000 gone on one account, and a warning sign on the base contract too, since Candor was the reason Hartsbridge chose the pricier tier over a competitor two years earlier.

The lagging outcome, charted: what happened to the add-on revenue Hartsbridge was paying for Candor
$160K $80K 0 $150,000 Base platform (renewed) $27,000 Candor, budgeted (at signing) $0 Candor, actual renewal (dropped)
HeldExpected at contract signingActual, at renewal
The base contract survived. The one line item wired directly to whether managers still trusted Candor's own words didn't.
Hand sketched flow diagram titled What Fenhollow never saved. Four boxes left to right: Candor drafts, manager edits, final save, history none kept. The fourth box is outlined in red-orange.
When a manager edited or replaced Candor's draft, the final save simply overwrote it. Nothing recorded what the draft had said in the first place.
The choice Chiamaka would take back When Fenhollow's engineers first built the review editor, they let a manager's edits overwrite Candor's draft in place, no separate record kept of the two versions. With a few hundred users, that was a sane way to keep the database small. Nobody could have told you, later, how much of any given "final" review had actually survived from Candor's own words, and for a long time nobody needed to. Once thousands of reviews a year were quietly answering the question of whether an account still trusted the tool, that same decision meant nobody could see it happening.

What I would leave alone: Candor also lets a manager paste in a messy pile of one-on-one notes and get back a tidy paragraph, purely for their own reference, never submitted as an official review. That path doesn't need this instrumentation. Nobody's trust in it decides whether Hartsbridge renews. It's a convenience, not the thing the contract is actually paying for.

The lesson: a metric that can only improve or hold flat is not measuring health, it's measuring participation. If the number you're watching would look the same whether people trust the tool or are quietly working around it, you haven't found the metric yet, you've found the paperwork.

Now here is the same thing as a story

The short version is above. Read this one when you want to feel why eleven hours of Franklin's own time, not a bad model, is what almost cost Fenhollow the account.

Franklin Vandenburg built Hartsbridge's whole review process himself, six years ago, out of a shared spreadsheet nobody trusted and a filing cabinet nobody opened. By the time Candor arrived, 140 store managers ran reviews the way he'd taught them to, on time, on a schedule, without him having to chase anyone.

The first months with Candor were good ones. Managers who used to grumble about review season started mentioning, almost in passing, that it wasn't so bad anymore. Franklin watched the completion dashboard every Monday morning, the way he always had. It sat at 97, 98 percent, week after week, exactly what he expected from a process he'd built to be reliable.

He stopped opening a sample of the actual reviews to read them. He'd done that for years out of habit, a dozen or so a month, just to get a feel for quality. Once Candor was writing most of the first draft, reading them felt like reading his own instructions back to himself. So he checked the one number instead, and the one number kept saying everything was fine.

The thing that cracked it wasn't a complaint. On a regional call, a store manager mentioned, almost as an aside, that she'd stopped using "the AI thing" a while back, easier to just write it herself. Nobody else on the call reacted. Franklin barely reacted either, at first.

But it stuck with him. A few days later he pulled up her store's last three reviews and read them properly for the first time in months. Then he pulled a few more stores. Then, over the next three weeks, with the renewal call approaching, he read ninety-one reviews by hand, about eleven hours he didn't actually have, because he had no number that could tell him which stores to worry about and which ones were fine. He was checking everything because he could no longer trust that checking nothing was safe.

Hand sketched comparison diagram titled One clock barely moved, the other ran the whole time. Left panel, a gauge icon with a needle only slightly past zero, labeled reviews completed on time, caption 97 percent, week 1 through week 44. Right panel, a gauge icon with a needle far around the dial, labeled managers rewriting from scratch, caption 8 percent to 41 percent, same 44 weeks.
Franklin had one gauge on his desk. The one that was actually moving wasn't on it.
We did not lose Franklin's trust to a bad model. We lost it to a dashboard that only ever had one gauge on it.

What he found in those ninety-one reviews wasn't a disaster. It was a pattern. Roughly one in six of the rewritten ones had a specific reason behind them, a claim in Candor's draft that wasn't anywhere in the manager's own notes, an achievement invented out of a gap the model filled the way a confident guesser fills a gap. The rest were softer: a tone that read like a corporate handbook, not like a store manager talking about someone she actually supervised. Once a manager caught the tool inventing one detail, or just sounding wrong once, they didn't go back and check draft by draft after that. They just stopped starting from it.

The decision that let this hide traced back to a short conversation at Fenhollow, more than a year earlier, when the review editor was first being built. An engineer asked whether the system should keep both versions of a review, Candor's original draft and whatever the manager finally saved, or just keep the final one and let the draft disappear. Keeping both meant a bigger, messier schema for a feature that, at the time, had a few hundred users total. The room chose simple. Nobody in it was deciding, on purpose, that Fenhollow would spend a year unable to answer whether Candor was still worth what it charged for it. They were deciding to ship a smaller feature faster, which was the right call, for exactly as long as it stayed true.

Run the same forty-four weeks again, with the history kept and the thresholds in place instead. The reversion rate crosses 15 percent in week fourteen, same as before. This time it's caught the same week, not the same year. Chiamaka's team pulls a sample, finds the hallucinated specifics and the tone mismatch, and ships a grounding fix by week eighteen: Candor now only drafts against material it can point back to, and it says less when a manager has logged less. By week twenty-four the rate is back under 15 percent. Franklin never opens a spreadsheet of ninety-one reviews. He spends zero hours on it. The renewal call runs eleven minutes, and it's about next year's store count, not about whether Candor still works.

Hand sketched decision tree titled What the team does with the rate. Root box, this account's weekly reversion rate, branching into three leaves. Under 15 percent leads to leave it alone. 15 to 30 percent leads to sample audit the drafts. Over 30 percent for 3 weeks leads to flag to Customer Success, this leaf outlined in red-orange.
The same three numbers, acted on in week fourteen instead of discovered in week forty-four, is the entire difference between the two versions of this story.

One design waits for a person to notice something feels off, on a regional call, months after it started. The other design notices for you, and hands Chiamaka's team a reason to act while there's still a renewal left to save.

What Chiamaka would tell herself, back in that short meeting about the schema: it wasn't really a storage decision. It was a decision that, for a year, nobody would be able to prove whether Candor was still earning what Hartsbridge paid for it, and nobody in the room meant to make that call on purpose.

LEAD: telling a dashboard from a warning

Not a way to dress up "watch engagement" in four letters. LEAD forces you to name the real dollar, find the number that moves before it does, and say out loud how that number gets faked.

LLink. The business outcome that actually matters, not the model's own score.
Candor's own add-on revenue on the Hartsbridge account, $27,000 a year, and what it signals about the $150,000 base contract it was sold alongside. Not Candor's draft-quality score. Not adoption. The dollar Fenhollow actually keeps or loses.
Name the real dollar before anything else, or every metric that follows is measuring something adjacent to the thing that pays the bills.
EEarly signal. The thing that moves weeks before the outcome does.
The weekly reversion rate, the share of reviews where a manager writes from a blank page instead of starting from Candor's draft. It went from 8 percent to 41 percent across forty-four weeks, while completion sat at 97 percent the entire time and told nobody anything.
This is the hardest step and the whole reason LEAD exists. A number that would have looked perfectly healthy right up until the renewal call is worse than no number at all.
AAbuse. How this metric gets gamed, by the team or the account.
Once by cohort: an early report only measured the twelve champion accounts who adopted first, all of whom renewed regardless. Once by definition: "Candor-active" meant switched on, not used, so a worsening account still counted as a retained one for three quarters running.
Naming how a metric gets gamed before anyone else finds the gap is what makes it a real number instead of a number you happen to like.
DDecision. What actually changes at each threshold.
Under 15 percent, leave it alone. From 15 to 30, pull a sample of discarded drafts and check them against the manager's real notes. Past 30 percent, sustained three weeks, Customer Success gets looped in before the renewal conversation, not after. If many accounts cross 30 at once, that's a model or prompt problem, not an account problem, and it goes back to the eval set before it ships again.
A metric with no attached action is a chart nobody uses. This is what turns the E step into something you actually do something about.

And if you want to be sure it really works, try it somewhere else

Same four letters, a regional internet provider instead of a retail chain, and this time the thing that breaks trust isn't an invented achievement, it's an outdated promise.

Firstline is Corvallo's AI feature for helpdesk software. It drafts an agent's reply to an incoming support ticket, pulling from the customer's account history and the company's own policy pages. Baymarsh Broadband, a regional internet provider, runs it across its whole support team. Meskerem Haile runs Support there, and Firstline was part of why Baymarsh signed Corvallo's higher tier in the first place.

Hand sketched flow diagram titled The same signal, a different desk. Four boxes left to right: ticket arrives, Firstline drafts, agent's choice, weekly rewrite rate. The third box, agent's choice, is outlined in teal.
Same shape as Candor's story, a different desk, and a different reason agents stop trusting the draft.
The decision Baymarsh would take back Firstline's grounding source, the document it drafts replies against, was Baymarsh's policy page as it stood the day Firstline launched. When Baymarsh quietly changed its refund window six months later, nobody re-pointed Firstline at the new version. Agents started catching replies that confidently promised a refund window that no longer existed, a policy hallucination, not an invented fact but a true one that had gone stale.

L, link: Firstline's own add-on revenue on the Baymarsh account, and the base helpdesk contract it was sold to justify. E, early signal: the weekly full-rewrite rate, tickets where an agent deletes Firstline's draft entirely and writes a new reply from nothing. A, abuse: "sent using Firstline" got counted even when an agent kept only the opening line and rewrote everything after it, a looser definition of adoption that let real reversion hide inside a metric that looked like usage. D, decision: under 10 percent, leave it alone; 10 to 25, sample the rewrites for policy claims that don't match the current policy page; past 25, sustained, flag the account and check whether Firstline's grounding source has gone stale, since that's a different fix than retraining the model.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to it: "Don't watch whether a ticket got answered, watch whether the agent kept Firstline's own words in the reply. That number moves before a renewal does."
Cost: no budget this quarter for a policy-audit pipeline. Ship the cheap version: a support lead spot-checks ten rewritten replies a week by hand, not a full tool.
The model got better, for real: say Firstline's grounding gets fixed and the stale-policy problem goes away entirely. The full-rewrite rate should fall. If it doesn't within a few weeks, trust was never only about accuracy, and it's worth asking what else agents stopped believing along the way.

Where people run it wrong.
They watch average reply time instead of rewrite rate, and reply time can get faster even as trust craters, because a fast rewrite from scratch still beats a slow one.
They let "sent using Firstline" cover a one-line edit, so real reversion hides inside a number that looks like adoption.
They wait for a CSAT survey to say something's wrong, and CSAT is slow and sparse, and by the time it moves the account's decision is basically already made.

How to use it live. Ask which number would have looked perfectly healthy right up until the morning it broke, before you name a metric. That buys real thinking time, and it reframes the whole question before you have to guess at a number.

Two things worth stating directly, since this is where the real judgment sits. The alternative Fenhollow considered and rejected for the E step was the simplest possible one, the percent of managers who ever opened Candor's draft at all. It lost, because open rate sat near 95 percent from week one regardless of what happened after, everyone was curious once, so it carried no predictive power. The AI-specific failure worth naming twice, once per product, is a model stating something false with full confidence, an invented achievement at Hartsbridge, a stale policy promise at Baymarsh, and the guardrail in both cases is the same shape: sample real output against real source material on a schedule, don't wait for a customer to catch it first. The trade-off Fenhollow accepted on purpose: Candor now drafts less often than it used to, refusing to write anything when a manager's notes are too thin to draft from safely, which means fewer employees get the time-saving draft in exchange for every draft that does get generated being one a manager can actually trust.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
LEAD: link a metric to the real dollar outcome, then find the signal that moves before that outcome does.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Franklin Vandenburg, Head of People at Hartsbridge Retail Group, who built the company's whole review process from a spreadsheet six years ago.
3 · THE LINK
What's the L step here?
Tap to flip
ANSWER
Candor's $27,000-a-year add-on revenue on the Hartsbridge account, and what it signals about the $150,000 base contract sold alongside it.
4 · THE EARLY SIGNAL
What's the E step here?
Tap to flip
ANSWER
The weekly share of reviews where a manager writes from a blank page instead of starting from Candor's draft. It went from 8 percent to 41 percent over 44 weeks, while completion held at 97 percent the whole time.
5 · THE OLD DECISION
What decision would Chiamaka take back?
Tap to flip
ANSWER
Letting a manager's edits overwrite Candor's original draft in place, with no saved history. It made sense at v1 scale, with a few hundred users and a simpler schema.
6 · THE NUMBER
Fill in the blank: the reversion rate went from ___ percent to ___ percent over ___ weeks, while completion held at ___ percent.
Tap to flip
ANSWER
8 percent to 41 percent over 44 weeks, while completion held at 97 percent.
7 · THE REPLAY
Same 44 weeks, new dashboard, what changes?
Tap to flip
ANSWER
Chiamaka's team catches the crossing at week 14 instead of week 44. Franklin spends zero hours manually auditing reviews before the renewal call, instead of the eleven hours he actually spent reading ninety-one of them by hand.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Firstline, Corvallo's ticket-reply drafting feature for Baymarsh Broadband. The early signal is the weekly full-rewrite rate on agent replies, caused by a stale policy source rather than an invented fact.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Franklin's dashboard look fine for ten months while Candor's real trouble was building underneath it?
  • A. Candor's underlying model accuracy was actually fine the whole time.
  • B. He was tracking whether reviews got submitted on time, a number that stays near 100 percent no matter how much managers trust the draft.
  • C. Hartsbridge hadn't turned Candor on for most of its stores.
  • D. Fenhollow raised Candor's price mid-contract.
Show hint
Ask what completion actually measures, and whether it's mandatory regardless of trust.
Show answer
B. The model wasn't the problem for most of that time. Completion is a paperwork number, and paperwork gets filed whether or not the manager still believes the draft.
Fill in the blank
2. The weekly reversion rate on Hartsbridge's account went from ___ percent to ___ percent over about ___ weeks, while completion stayed at ___ percent the whole time.
Show hint
It's stated in the "here's the turn" paragraph in Let's learn, and again in the line chart.
Show answer
8 percent to 41 percent over about 44 weeks. Completion stayed at 97 percent. The gap between those two numbers is the entire teaching point of this answer.
True or false
3. True or false: an account where Candor is switched on for every store should always count as a "Candor-active, retained" account in a churn-impact report.
  • True
  • False
Show hint
Ask what "switched on" actually tells you about whether anyone is using it.
Show answer
False. Being switched on says nothing about whether managers keep the drafts. Hartsbridge was counted this way for three quarters while its reversion rate climbed past 30 percent, which is exactly how the metric hid a worsening account.
Short answer, name the old decision
4. What old decision would Chiamaka take back, and why did it make sense when Candor first shipped?
Show hint
Look at the key point box right after the second chart, titled "The choice Chiamaka would take back."
Show answer
Model answer: Letting a manager's edits overwrite Candor's draft in place, keeping no separate record of the two versions. It made sense at v1 scale with a few hundred users and a simpler database, and it stopped making sense once thousands of reviews a year needed to answer whether an account still trusted the tool.
Short answer, apply it yourself
5. Think of a tool at your own job, or a service you use, with a mandatory step baked into it, something you'd do with or without the tool actually helping. Name the number that mandatory step would keep looking healthy on, and a different number, closer to real behavior, that would move first if people quietly stopped trusting the tool.
Show hint
Ask which number is required paperwork, and which one only moves if trust actually changes.
Show answer
Model answer: A code-review tool that requires an "approved" review before merging. Approval rate stays near 100 percent forever, because it's mandatory. The number that would move first is how long a reviewer actually spends on the diff before approving, which can quietly shrink to seconds long before anyone admits the reviews aren't real anymore.
Short answer, work the number
6. If Hartsbridge's reversion rate had crossed 30 percent at week 20 instead of week 31, and the intervention rule requires three sustained weeks above 30 percent before Customer Success gets looped in, in what week would the flag have fired, and would that still have landed before the week-44 renewal call?
Show hint
Check the D step in the framework recap, and add the three sustained weeks to week 20.
Show answer
Around week 23. That's still about 21 weeks of lead time before the week-44 renewal call, plenty of room to act before the number ever became a conversation with Franklin.
Before you close the answer
Why this works
Tests whether you'll reach for a metric a model's output actually has to earn, or one that just proves a mandatory process ran. Most candidates measure whether the AI got used. The stronger answer measures whether anyone still believes what it produces.
Follow-up traps
"Isn't a 41 percent rewrite rate just proof the model is bad, so why not just retrain it right away?" Response: retraining is the fix once a sample audit confirms what's actually wrong. Acting on the aggregate rate alone, without sampling the discarded drafts first, risks fixing tone when the real problem was a made-up fact, or the reverse.

"What if a manager rewrites the draft just because they're a strong writer who prefers their own voice, not because they distrust it?" Response: that's exactly why the 15-to-30 tier samples drafts before anyone acts on them. A rising rate gets checked against real notes before it's treated as a trust problem, never assumed to be one on sight.
If pressed
Candor's grounding fix wasn't just a prompt tweak. It now only generates a draft when a manager has logged at least three dated observations for that employee. Below that, the field stays blank and the manager writes it unassisted, which is the actual guardrail behind the sample audits, not a promise, a refusal to guess.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more