ConceptAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #7

Explain the problem with measuring acceptance rate of AI suggestions.

The direct answer
Acceptance rate only proves a marketer clicked accept or left the top suggestion alone and hit send. It does not prove the suggestion was any good. Split it by how much the send mattered, and compare suggestions people edited before sending against ones they accepted untouched. If the untouched ones perform worse, the rate is climbing because people trust it, not because the model is actually getting better.
Do this, in order
  1. Stop reading a rising acceptance rate as proof the suggestions got better.Why: a click only means the marketer did not stop the send. It says nothing about whether the line was strong.
  2. Split acceptance rate by how much the send mattered, not one blended number.Why: a routine newsletter and a big seasonal campaign get folded into the same average, and the average hides the expensive send getting worse.
  3. Check whether "accepted" quietly changed its own meaning before you trust the trend at all.Why: a redesign that auto-fills the top suggestion can turn doing nothing into an accepted choice, with nobody deciding that on purpose.
  4. Compare suggestions a marketer edited before sending against ones they accepted untouched, holding the send's stakes steady.Why: this is the one check that tells a genuinely good suggestion apart from one that is just easy to rubber-stamp.
  5. Watch whether the suggestion model is quietly drifting toward safe, generic lines because those get clicked more.Why: a model trained on clicks will chase the click, not the open, unless something else is grading it.
  6. Keep a second number next to acceptance rate: real performance on a held-out set of past sends.Why: acceptance without an outcome check is a popularity count, not a quality bar.

How to answer this, stage by stage

Nobody is grading whether you know that acceptance rate can be gamed. They are grading whether you can find out, in front of them, live, using a method instead of a hunch. Seven moves get you there.

1
Scope it to one product and one number
Say it like this
"Let's ground this. Everquill is Copperline's tool that suggests subject lines and ad copy right inside the campaign builder. Verity Ozanne runs product for it, and the number everyone was cheering was acceptance rate, the share of suggestions marketers used instead of writing their own."
Why this works
Grounds the diagnosis in a real product and a real number before naming a single suspect.
2
Say the method out loud before naming a cause
Say it like this
"Here's how I'd frame it. A metric that's climbing is not automatically a metric that's working. I rule out the boring explanations first, then narrow down to what's actually happening, instead of jumping straight to 'the model must be getting better.'"
Why this works
States TRACE as a rule you can reuse, not a reaction you had to this one story.
3
Lay the timeline down and name the gap
Say it like this
"Over two quarters, acceptance rate went from 41 percent to 74 percent. Nobody could point to one launch that caused it, it was a slow climb. But open rate on those same campaigns barely moved, 22.4 percent to 21.6 percent. A number climbing while the thing it's supposed to predict sits flat, that gap is where I start."
Why this works
This is the T step, and the gap itself is the whole reason to keep digging instead of celebrating.
4
Recut the number before trusting it
Say it like this
"I don't stop at the blended average. I split acceptance by how much the send mattered. Routine newsletter lines went from 52 percent accepted to 88. Black Friday, cart abandonment, win back lines went from 34 percent to 71, almost the same climb. That's the part that should worry you, because a weak line on a routine newsletter costs almost nothing, and a weak line on Black Friday costs real money at scale. Nobody was scrutinizing the expensive sends any harder than the cheap ones."
Why this works
This is the R step, the strongest single move in this framework, because the blended number would never have shown it.
5
Rule out instrumentation before behavior
Say it like this
"Before I chase a story about marketers or a story about the model, I check whether 'accepted' still means what it meant in January. Turns out a March redesign auto-fills the top suggestion into the subject line box now. If nobody touches it and hits send, that counts as accepted. So part of this climb might just be a new default, not new trust."
Why this works
This is the A step. A tracking change can look exactly like a behavior change on a dashboard, and you would chase the wrong story without checking.
6
Name the real suspects, and the one check that tells them apart
Say it like this
"Three things could be happening at once, and they are not the same problem. One, people click accept without really reading the line, because nothing has gone wrong yet. Two, the new default is inflating the number on its own, whether or not anyone is paying attention. Three, the model itself learned that safe, generic lines get accepted more often, so it ranks those higher, even if a bolder line would open better. The check that tells them apart: pull every high stakes send from last quarter and compare open rate for the ones sent exactly as suggested against the ones a marketer edited first."
Why this works
This is the C and E steps together, three named causes and the one test that actually separates them instead of guessing.
7
Close on the diagnosis, defended in one line
Say it like this
"So here's the answer: acceptance rate went up because it got easier to trust without checking, not because the suggestions got better where it counted. Untouched sends underperformed edited ones by about five points of open rate, on exactly the campaigns where that gap is expensive. Acceptance rate is not wrong to track. It is wrong to trust alone."
Why this works
Closes on the actual diagnosis, something a reader can apply to any climbing metric, not just this one.

Let's learn

What happens when a number keeps climbing and everyone quietly assumes that means the thing underneath it is getting better?

Say a marketing tool suggests the subject line and the ad copy for a campaign, right inside the screen where a marketer builds it. She can accept the suggestion, edit it, or write her own. Before a tool like this, a marketer wrote every subject line herself, testing two or three ideas by feel, about eight minutes of thinking per line. With the suggestion tool running, three ready lines show up the moment she opens the campaign, and most weeks she just reads them and clicks one, cutting that eight minutes to under one.

Blended acceptance rate versus blended open rate, six months
100% 50% 0% 41% 74% 22.4% 21.6% Jan Jun Acceptance rate Open rate
Acceptance rate nearly doubled in six months. Open rate on those same sends did not move. That gap, not the climb itself, is where the real question starts.

Here is the turn. A climbing acceptance number got read as the model getting sharper. It was not the model. It was marketers reading the number and trusting it more, while the suggestions themselves, on the sends that mattered most, quietly got less scrutiny, not more accuracy.

We did not build a model that got worse at writing subject lines. We built a habit that stopped checking the sends where checking actually paid for itself.

Splitting by stakes makes it worse, not better. Routine newsletter suggestions went from 52 percent accepted to 88 percent, which is fine, a weak newsletter line barely costs anything. But Black Friday, cart abandonment, and win back suggestions climbed almost the same amount, 34 percent to 71 percent, on sends where a weak line costs real revenue. And the actual result on those expensive sends tells a different story than the acceptance line does.

Open rate by campaign stakes, January versus June
25% 0% 21.9% 22.1% 24.1% 18.7% Routine newsletter High stakes campaigns
Routine sends held steady. High stakes sends, the ones with the most money behind them, dropped over five points of open rate in the same six months acceptance rate climbed just as fast for both groups.
Knowledge spark: why would a model learn to suggest something worse? A suggestion model often trains on what people click and keep. If a bland, safe line gets accepted more often than a bold one, the model learns bland wins, even if the bold line, on the rare send where someone actually used it, got more opens. It is chasing the click, not the result.

At its worst, this costs the exact opposite of what the metric promised. The team ships more campaigns with less human judgment applied to the ones with real money behind them, while a dashboard says everything is going great.

The decision that mattered "Accepted" was let to mean two different things, a suggestion someone chose on purpose and a suggestion nobody touched before sending, and both got counted the same way. Split those apart before you trust the number at all.

The choice I would take back is not the auto-fill redesign itself. It is that when that redesign shipped, nobody split the "accepted" event in two, one for a suggestion a person actually picked, one for a suggestion that just sat there until send. That was a fine call for a five minute decision about a UI tweak. It stopped being fine the moment the team started using that same number as proof the model was working.

What I would leave alone: acceptance rate on routine, low stakes sends. A weekly newsletter line that underperforms by a point costs almost nothing, and marketers already treat those suggestions the right way, quickly, without much second guessing. Chasing that number down to the segment level would be effort spent where the cost of being wrong is small.

The lesson: a metric that only measures whether someone let something through will always look healthier than the thing it is supposed to be checking. It rewards silence, not judgment, and those are not the same thing.

Now here is the same thing as a story

The short version sits above. Read on if you want to feel how ordinary the Monday looked from Verity's side, right up until the morning she pulled the thread.

Every Monday morning, Verity opens the same dashboard tile first, before anything else on her list. She has run product for Everquill for three years, since before it could suggest anything longer than a single line, and she knows the difference between a metric that is really moving and one that is just bouncing around.

Everquill launched inside Copperline's campaign builder two years back. For the first several months, watching it work felt like watching something click into place. A marketer would open a new campaign, and three subject lines and two ad copy blocks would already be sitting there, waiting. Most people used them as a starting point at first. They read all three, took the one they liked best, and changed a word or two to sound like their own brand.

By the second quarter, that habit had started to thin. First it was just the routine sends, the weekly newsletter, the receipt confirmations, where a marketer would glance at the top suggestion and click it without reading the other two. Then it spread to the seasonal campaigns. By the time Black Friday planning started that second year, most marketers using Everquill were doing the same thing on every send: glance, click, move on.

The trigger, when it came, was almost nothing. In March, Copperline shipped a small redesign. The top suggestion now sat pre filled in the subject line box the second a marketer opened the campaign, instead of waiting in a side panel until someone clicked to see it. Nobody flagged it as a big change. It saved a click.

Verity did not notice anything for weeks. The acceptance number kept climbing, and every week it climbed, it looked like more proof the model was getting good. Forty one percent in January. Fifty eight in March. Seventy four by June.

Then, on an ordinary Monday, she did something she almost skipped: she pulled up open rates next to acceptance, split by campaign type, instead of just eyeballing the acceptance graph on its own.

Black Friday sends: 71 percent acceptance, open rate 18.7 percent, down from 24.1 the year before. Routine newsletter: 88 percent acceptance, open rate basically unchanged. Same climbing line on the acceptance side. Completely different story underneath it.

Verity spent that week not looking at the model at all. She looked at the March redesign notes first, because a jump that clean deserved a boring explanation before an interesting one. She found it inside an hour: the auto fill change had quietly redefined what counted as accepted. A marketer who did nothing at all, who never even opened the suggestion panel, was now credited with choosing it.

That was not the whole story either. Even after correcting for the auto fill, high stakes acceptance had climbed almost as fast as routine acceptance, and the two should never have moved together. So she ran the one check that would actually separate the real suspects: every high stakes send from the past quarter, sorted into two piles, sent exactly as suggested, and sent after a marketer changed at least a few words.

The untouched pile opened at 18.2 percent. The edited pile opened at 23.6 percent. Same suggestion engine, same day, same size of audience. The only difference was whether a person had actually looked hard enough to change something.

The decision Verity would take back was made a year earlier, in a short meeting about the March redesign, when someone asked whether auto filling the top suggestion needed its own tracking, separate from a suggestion someone picked on purpose. The answer in the room was no, it is the same event, do not overbuild it. That was the right call for a five minute decision about a UI tweak. It was the wrong call for a number the whole team was about to lean on as proof of quality.

The replay: with the split metric back in place, edited versus untouched broken out by campaign stakes, Verity's team caught the next drift inside three weeks instead of two quarters. A ranking update that August had quietly started favoring lines with the word "today" in them, because those got accepted fastest, and by week three the edited pile's open rate on high stakes sends had already pulled two points ahead of the untouched pile again. They rolled the ranking change back before it ever touched a real Black Friday.

The old design gave Verity one number and let her believe it. The new one gives her two numbers that have to agree before she believes either.

What I would tell myself, looking back at that ordinary Monday: a number that only measures whether someone let something through will always look healthier than the thing it is supposed to be checking.

TRACE, spelled out in five checks

This is a question about whether a metric is lying to you, not a story with a two setting switch, so TRACE fits, not a framework built around a habit that snaps.

T
Timeline. When did it start, and what shipped near that date, including things that looked like improvements.
Acceptance climbed from 41 to 74 percent over two quarters, while open rate barely moved. A March redesign, filed as an improvement, sits right at the bend in that line.
R
Recut. Slice by segment, cohort, or stakes.
Split by how much a send mattered: routine acceptance and high stakes acceptance climbed at nearly the same rate, even though a weak high stakes line costs far more.
A
Assume nothing. Rule out instrumentation before behavior.
The March redesign silently changed what "accepted" meant, crediting an auto filled default as if it were a real choice.
C
Cause candidates. Three named, not a list of everything possible.
Rubber stamp acceptance, default inertia from the auto fill, and a suggestion model quietly learning that safe, generic lines get clicked more often.
E
Evidence test. The one check that separates the top suspects.
Compare open rate on high stakes sends accepted untouched against ones a marketer edited first: 18.2 percent versus 23.6 percent.

One thing worth saying out loud here, since this is exactly where an AI product question earns its name. The alternative most people reach for by reflex is a confirmation step, a small popup asking "are you sure" before a high stakes suggestion gets used untouched. That got ruled out on purpose. Friction on the ninety percent of sends where the suggestion really is fine just trains people to click through popups without reading them, the same rubber stamp problem wearing a different button. The real bar for Everquill's ranking model is not "never suggest a bland line." It is a calibrated one: on a held out set of past high stakes campaigns, the model's top suggestion has to beat a plain human written control on open rate at least six times out of ten, checked every quarter, not trusted just because acceptance looks good this month. That trade is real too. Splitting the metric and running the quarterly check costs a data analyst's time and a slower release cycle for ranking updates. It buys back the ability to catch a blandness drift in three weeks instead of two quarters.

And if you want to be sure it really works, try it somewhere else

Same five checks, a completely different product and a much higher cost of being wrong, so the method proves itself instead of repeating a story you happened to prepare.

Corrievale Health sells Lanternwood, a chat triage assistant vet techs use overnight, when a pet owner messages in about a symptom and someone has to decide fast whether it can wait until morning or needs an emergency visit tonight.

T, timeline. Acceptance of Lanternwood's suggested replies climbed from 46 percent to 79 percent over four months. Nothing dramatic shipped, except a July release that added a one tap "send as suggested" button to the top of the chat window.
R, recut. Split by how serious the symptom sounded on Corrievale's own scale, acceptance for mild cases and acceptance for serious cases climbed at almost the same rate, even though a wrong call on a serious case can cost an animal its life.
A, assume nothing. Malaika Brambila's team checked the July release notes first. The one tap button had not changed what counted as accepted, that part held up clean. What had changed was staffing: two techs left in June, and the overnight shift ran short handed the rest of the summer.
C, cause candidates. Techs waving serious cases through faster because they were stretched thin, at exactly the moment they had less time to double check anything. A model quietly favoring "monitor at home" replies over "come in now" ones, because owners followed the calmer advice without pushing back, which read as a good outcome in the data even when it was not. Or a real improvement, since Lanternwood's underlying triage model had genuinely gotten better that same quarter.
E, evidence test. Corrievale pulled every serious case chat from the short handed months and checked which ones later needed an emergency visit anyway, comparing suggestions sent untouched against ones a tech had changed. Untouched serious case replies led to a same night emergency visit later on 6 percent of chats. Edited ones, on the same kind of case, led to one on 2 percent.

Same shape, worse stakes At Everquill, the hidden cost was a subject line nobody opened. At Lanternwood, it is a sick animal that needed a same night visit and got told to wait instead. The check does not change: compare suggestions sent exactly as written against ones a person actually changed.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule: acceptance rate proves a click happened, not that the suggestion was good, so split it by stakes and check edited versus untouched before trusting it.
Cost: pulling every high stakes send by hand takes a data analyst a week. Do not skip the check, sample fifty of them instead of all of them. The method still works on a sample, it just needs enough volume to trust the split.
The model got better, for real: say the underlying model genuinely improved that quarter. That still is not proof acceptance rate can be trusted on its own. You would expect edited and untouched suggestions to perform about the same if the model really got better everywhere, not just for the ones nobody double checked.

Where people run it wrong.
They treat a rising acceptance number as the whole story, without ever splitting it by how much a bad call would cost.
They assume the number's definition never changed, and go hunting for a behavior story before ruling out a tracking one.
They fix the acceptance number itself, a nag, a popup, a reminder, instead of checking whether the number ever meant what people thought it meant.

How to use it live. Say the rule before naming a single cause: "A metric climbing is not the same as the thing it stands for getting better, I check that gap before I explain it." That buys the room to actually work through the five checks instead of guessing the first plausible story out loud.

Flashcards, tap to flip

1 · THE FRAMEWORK
What framework fits diagnosing why a climbing metric might be lying to you?
Tap to flip
ANSWER
TRACE: rule out the boring explanations first, timeline, recut, assume nothing, then narrow to the real cause with a named evidence test.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Verity Ozanne, product manager for Everquill, Copperline's AI subject line and ad copy suggestion tool, three years into running it.
3 · THE HABIT
What did marketers stop doing because it worked?
Tap to flip
ANSWER
They stopped reading all three suggested lines and picking one on purpose. They started letting the top suggestion sit in the box and hitting send.
4 · THE GAP
What is the gap between the two numbers in this story?
Tap to flip
ANSWER
Acceptance rate climbed from 41 percent to 74 percent over two quarters. Open rate on those same campaigns barely moved, 22.4 percent to 21.6 percent.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting an auto filled default suggestion count as accepted the same way a suggestion someone actively chose did, instead of splitting the metric in two.
6 · THE NUMBER
Fill in the blank: high stakes suggestions sent exactly as written opened at ___ percent. Ones a marketer edited first opened at ___ percent.
Tap to flip
ANSWER
18.2 percent untouched, 23.6 percent edited. Same engine, same day, only the editing differed.
7 · THE EVIDENCE TEST
What is the one check that separates a genuinely better suggestion from one that is just easy to rubber stamp?
Tap to flip
ANSWER
Compare open rate on suggestions sent exactly as written against ones a marketer edited first, holding campaign stakes steady. If untouched loses, the acceptance number was measuring trust, not quality.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs this same check on a different product. Which product, and what did the evidence test find?
Tap to flip
ANSWER
Lanternwood, Corrievale Health's vet triage chat assistant. Untouched serious case replies led to a same night emergency visit later 6 percent of the time, versus 2 percent for edited ones.

Check yourself Score: 0 / 0

True or false
1. True or false: a rising acceptance rate on its own is good evidence that an AI suggestion tool is getting better.
  • True
  • False
Show hint
Think about what a click actually proves, and what else could cause it.
Show answer
False. It only shows people are letting suggestions through more often, which can also happen from a UI change or from people trusting the tool too much to double check it.
Multiple choice
2. Why did splitting acceptance rate by campaign stakes matter more than looking at the blended number?
  • A. Blended numbers are always wrong and should never be reported.
  • B. High stakes and routine campaigns climbed at nearly the same rate, hiding that high stakes suggestions needed more scrutiny, not the same amount.
  • C. Everquill only tracks blended numbers by default and cannot be changed.
  • D. Routine campaigns do not get AI suggestions at all.
Show hint
Look at the two acceptance curves in stage 4 of the walkthrough.
Show answer
B. A blended average bundles a cheap mistake and an expensive one into the same line. The split is what showed the expensive campaigns were not getting extra scrutiny.
Fill in the blank
3. Copperline's ___ release quietly changed what counted as accepted, by auto filling the top suggestion into the subject line box.
Show hint
Check the A step in the story and the framework recap.
Show answer
March. The redesign shipped in March, and acceptance rate had already started climbing by the time anyone thought to check the definition.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory about the March redesign, not a dial anyone could just turn back up.
Show answer
Model answer: The decision not to split the "accepted" event into a real choice and an untouched default when the auto fill redesign shipped. It made sense because it was a five minute call about a small UI tweak, not one anyone expected to become the team's main quality metric a year later.
Short answer, apply it yourself
5. Think of a tool you use where "usage went up" gets treated as "it's working." What is one number you would check next to it before believing that?
Show hint
Look for a number that could be flattered by a default, not just by real behavior.
Show answer
Model answer: A to do app where tasks marked complete keeps rising. Check that against a harder number, like tasks that actually got a follow up or a real outcome, since marking something done and it actually being done are not the same thing.
Short answer, the number question
6. If the untouched versus edited gap had been 1 point of open rate instead of 5, would the same conclusion still hold?
Show hint
Think about how big a gap needs to be before it stops looking like noise.
Show answer
Not as strongly. A 1 point gap on a metric that naturally wobbles a bit is closer to noise, so you would want a bigger sample or a longer window before calling it proof of rubber stamping. A 5 point gap on the same sample size is a real signal.
Before you close the answer
Why this works
Tests whether you will treat a friendly looking metric as proof, or go find out what it is actually measuring before you trust it. Most candidates stop at "acceptance went up, great."
Follow-up traps
"Isn't checking edited versus untouched suggestions just proving people who care more get better results, nothing to do with the model?" Response: partly true, and that is exactly the point. Acceptance rate cannot tell those two things apart, which is why it needs a second number sitting next to it.

"What if splitting the metric by stakes just adds a dashboard nobody reads?" Response: it only earns its place if someone acts on the split, so pair it with the quarterly held out check and a real rollback decision, like the August ranking change that got pulled before it shipped.
If pressed
The held out eval Verity's team built samples fifty high stakes sends a quarter, checks the model's top suggestion against a plain human written control blind, and requires the model to beat the control on open rate at least six times out of ten before that quarter's ranking update ships to everyone.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more