Explain the problem with measuring acceptance rate of AI suggestions.
- Stop reading a rising acceptance rate as proof the suggestions got better.Why: a click only means the marketer did not stop the send. It says nothing about whether the line was strong.
- Split acceptance rate by how much the send mattered, not one blended number.Why: a routine newsletter and a big seasonal campaign get folded into the same average, and the average hides the expensive send getting worse.
- Check whether "accepted" quietly changed its own meaning before you trust the trend at all.Why: a redesign that auto-fills the top suggestion can turn doing nothing into an accepted choice, with nobody deciding that on purpose.
- Compare suggestions a marketer edited before sending against ones they accepted untouched, holding the send's stakes steady.Why: this is the one check that tells a genuinely good suggestion apart from one that is just easy to rubber-stamp.
- Watch whether the suggestion model is quietly drifting toward safe, generic lines because those get clicked more.Why: a model trained on clicks will chase the click, not the open, unless something else is grading it.
- Keep a second number next to acceptance rate: real performance on a held-out set of past sends.Why: acceptance without an outcome check is a popularity count, not a quality bar.
How to answer this, stage by stage
Nobody is grading whether you know that acceptance rate can be gamed. They are grading whether you can find out, in front of them, live, using a method instead of a hunch. Seven moves get you there.
Let's learn
What happens when a number keeps climbing and everyone quietly assumes that means the thing underneath it is getting better?
Say a marketing tool suggests the subject line and the ad copy for a campaign, right inside the screen where a marketer builds it. She can accept the suggestion, edit it, or write her own. Before a tool like this, a marketer wrote every subject line herself, testing two or three ideas by feel, about eight minutes of thinking per line. With the suggestion tool running, three ready lines show up the moment she opens the campaign, and most weeks she just reads them and clicks one, cutting that eight minutes to under one.
Here is the turn. A climbing acceptance number got read as the model getting sharper. It was not the model. It was marketers reading the number and trusting it more, while the suggestions themselves, on the sends that mattered most, quietly got less scrutiny, not more accuracy.
Splitting by stakes makes it worse, not better. Routine newsletter suggestions went from 52 percent accepted to 88 percent, which is fine, a weak newsletter line barely costs anything. But Black Friday, cart abandonment, and win back suggestions climbed almost the same amount, 34 percent to 71 percent, on sends where a weak line costs real revenue. And the actual result on those expensive sends tells a different story than the acceptance line does.
At its worst, this costs the exact opposite of what the metric promised. The team ships more campaigns with less human judgment applied to the ones with real money behind them, while a dashboard says everything is going great.
The choice I would take back is not the auto-fill redesign itself. It is that when that redesign shipped, nobody split the "accepted" event in two, one for a suggestion a person actually picked, one for a suggestion that just sat there until send. That was a fine call for a five minute decision about a UI tweak. It stopped being fine the moment the team started using that same number as proof the model was working.
What I would leave alone: acceptance rate on routine, low stakes sends. A weekly newsletter line that underperforms by a point costs almost nothing, and marketers already treat those suggestions the right way, quickly, without much second guessing. Chasing that number down to the segment level would be effort spent where the cost of being wrong is small.
The lesson: a metric that only measures whether someone let something through will always look healthier than the thing it is supposed to be checking. It rewards silence, not judgment, and those are not the same thing.
Now here is the same thing as a story
The short version sits above. Read on if you want to feel how ordinary the Monday looked from Verity's side, right up until the morning she pulled the thread.
Every Monday morning, Verity opens the same dashboard tile first, before anything else on her list. She has run product for Everquill for three years, since before it could suggest anything longer than a single line, and she knows the difference between a metric that is really moving and one that is just bouncing around.
Everquill launched inside Copperline's campaign builder two years back. For the first several months, watching it work felt like watching something click into place. A marketer would open a new campaign, and three subject lines and two ad copy blocks would already be sitting there, waiting. Most people used them as a starting point at first. They read all three, took the one they liked best, and changed a word or two to sound like their own brand.
By the second quarter, that habit had started to thin. First it was just the routine sends, the weekly newsletter, the receipt confirmations, where a marketer would glance at the top suggestion and click it without reading the other two. Then it spread to the seasonal campaigns. By the time Black Friday planning started that second year, most marketers using Everquill were doing the same thing on every send: glance, click, move on.
The trigger, when it came, was almost nothing. In March, Copperline shipped a small redesign. The top suggestion now sat pre filled in the subject line box the second a marketer opened the campaign, instead of waiting in a side panel until someone clicked to see it. Nobody flagged it as a big change. It saved a click.
Verity did not notice anything for weeks. The acceptance number kept climbing, and every week it climbed, it looked like more proof the model was getting good. Forty one percent in January. Fifty eight in March. Seventy four by June.
Then, on an ordinary Monday, she did something she almost skipped: she pulled up open rates next to acceptance, split by campaign type, instead of just eyeballing the acceptance graph on its own.
Black Friday sends: 71 percent acceptance, open rate 18.7 percent, down from 24.1 the year before. Routine newsletter: 88 percent acceptance, open rate basically unchanged. Same climbing line on the acceptance side. Completely different story underneath it.
Verity spent that week not looking at the model at all. She looked at the March redesign notes first, because a jump that clean deserved a boring explanation before an interesting one. She found it inside an hour: the auto fill change had quietly redefined what counted as accepted. A marketer who did nothing at all, who never even opened the suggestion panel, was now credited with choosing it.
That was not the whole story either. Even after correcting for the auto fill, high stakes acceptance had climbed almost as fast as routine acceptance, and the two should never have moved together. So she ran the one check that would actually separate the real suspects: every high stakes send from the past quarter, sorted into two piles, sent exactly as suggested, and sent after a marketer changed at least a few words.
The untouched pile opened at 18.2 percent. The edited pile opened at 23.6 percent. Same suggestion engine, same day, same size of audience. The only difference was whether a person had actually looked hard enough to change something.
The decision Verity would take back was made a year earlier, in a short meeting about the March redesign, when someone asked whether auto filling the top suggestion needed its own tracking, separate from a suggestion someone picked on purpose. The answer in the room was no, it is the same event, do not overbuild it. That was the right call for a five minute decision about a UI tweak. It was the wrong call for a number the whole team was about to lean on as proof of quality.
The replay: with the split metric back in place, edited versus untouched broken out by campaign stakes, Verity's team caught the next drift inside three weeks instead of two quarters. A ranking update that August had quietly started favoring lines with the word "today" in them, because those got accepted fastest, and by week three the edited pile's open rate on high stakes sends had already pulled two points ahead of the untouched pile again. They rolled the ranking change back before it ever touched a real Black Friday.
The old design gave Verity one number and let her believe it. The new one gives her two numbers that have to agree before she believes either.
What I would tell myself, looking back at that ordinary Monday: a number that only measures whether someone let something through will always look healthier than the thing it is supposed to be checking.
TRACE, spelled out in five checks
This is a question about whether a metric is lying to you, not a story with a two setting switch, so TRACE fits, not a framework built around a habit that snaps.
One thing worth saying out loud here, since this is exactly where an AI product question earns its name. The alternative most people reach for by reflex is a confirmation step, a small popup asking "are you sure" before a high stakes suggestion gets used untouched. That got ruled out on purpose. Friction on the ninety percent of sends where the suggestion really is fine just trains people to click through popups without reading them, the same rubber stamp problem wearing a different button. The real bar for Everquill's ranking model is not "never suggest a bland line." It is a calibrated one: on a held out set of past high stakes campaigns, the model's top suggestion has to beat a plain human written control on open rate at least six times out of ten, checked every quarter, not trusted just because acceptance looks good this month. That trade is real too. Splitting the metric and running the quarterly check costs a data analyst's time and a slower release cycle for ranking updates. It buys back the ability to catch a blandness drift in three weeks instead of two quarters.
And if you want to be sure it really works, try it somewhere else
Same five checks, a completely different product and a much higher cost of being wrong, so the method proves itself instead of repeating a story you happened to prepare.
Corrievale Health sells Lanternwood, a chat triage assistant vet techs use overnight, when a pet owner messages in about a symptom and someone has to decide fast whether it can wait until morning or needs an emergency visit tonight.
T, timeline. Acceptance of Lanternwood's suggested replies climbed from 46 percent to 79 percent over four months. Nothing dramatic shipped, except a July release that added a one tap "send as suggested" button to the top of the chat window.
R, recut. Split by how serious the symptom sounded on Corrievale's own scale, acceptance for mild cases and acceptance for serious cases climbed at almost the same rate, even though a wrong call on a serious case can cost an animal its life.
A, assume nothing. Malaika Brambila's team checked the July release notes first. The one tap button had not changed what counted as accepted, that part held up clean. What had changed was staffing: two techs left in June, and the overnight shift ran short handed the rest of the summer.
C, cause candidates. Techs waving serious cases through faster because they were stretched thin, at exactly the moment they had less time to double check anything. A model quietly favoring "monitor at home" replies over "come in now" ones, because owners followed the calmer advice without pushing back, which read as a good outcome in the data even when it was not. Or a real improvement, since Lanternwood's underlying triage model had genuinely gotten better that same quarter.
E, evidence test. Corrievale pulled every serious case chat from the short handed months and checked which ones later needed an emergency visit anyway, comparing suggestions sent untouched against ones a tech had changed. Untouched serious case replies led to a same night emergency visit later on 6 percent of chats. Edited ones, on the same kind of case, led to one on 2 percent.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule: acceptance rate proves a click happened, not that the suggestion was good, so split it by stakes and check edited versus untouched before trusting it.
Cost: pulling every high stakes send by hand takes a data analyst a week. Do not skip the check, sample fifty of them instead of all of them. The method still works on a sample, it just needs enough volume to trust the split.
The model got better, for real: say the underlying model genuinely improved that quarter. That still is not proof acceptance rate can be trusted on its own. You would expect edited and untouched suggestions to perform about the same if the model really got better everywhere, not just for the ones nobody double checked.
Where people run it wrong.
They treat a rising acceptance number as the whole story, without ever splitting it by how much a bad call would cost.
They assume the number's definition never changed, and go hunting for a behavior story before ruling out a tracking one.
They fix the acceptance number itself, a nag, a popup, a reminder, instead of checking whether the number ever meant what people thought it meant.
How to use it live. Say the rule before naming a single cause: "A metric climbing is not the same as the thing it stands for getting better, I check that gap before I explain it." That buys the room to actually work through the five checks instead of guessing the first plausible story out loud.
Flashcards, tap to flip
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if splitting the metric by stakes just adds a dashboard nobody reads?" Response: it only earns its place if someone acts on the split, so pair it with the quarterly held out check and a real rollback decision, like the August ranking change that got pulled before it shipped.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?