CaseAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #9
How would you detect that users are working around your AI feature rather than with it?
In one breath
Don't watch whether people open your AI feature. Watch how much of what it generates actually survives to the thing they submit, section by section, not as one score for the whole document. Build a similarity check between the model's draft and the final output, per section, and flag the section that keeps coming back near-empty of the model's own words, even while every other part of the document, and the person's daily use of the tool, looks completely healthy.
Ranked, not chronological
Build a per-section similarity check between what the model generated and what actually gets submitted.Why: it is the only signal that survives a person who still opens your tool every single day.
Stop trusting daily use, submit counts, or completion rate as your workaround signal.Why: all three stayed perfectly healthy the whole time Odessa was quietly rewriting the one section that mattered.
Give every generated section its own "keep" or "regenerate this only" control.Why: it removes the reason leaving the app was ever the easiest option.
Calibrate the flag against a labeled set of confirmed workarounds, with a probability band, not one hard cutoff.Why: a single hard percentage punishes someone who is honestly personalizing, not routing around the feature.
Leave a small dip in bullet-level overlap alone.Why: a person swapping in their own real numbers makes the resume more true, not less trusted, and shouldn't trip the same flag.
Route every flagged section back to the model team as a labeled example, not just a dashboard number.Why: that closes the exact blind spot that let the habit form unseen in the first place.
How to answer this, stage by stage
Nobody is grading whether you can name "engagement" as a metric. They are grading whether you can find a signal that survives a person who never stops opening the app. Eight moves get you there.
1
Scope it to one real product, one real person
Say it like this
"Let's ground this. Wordbridge is Northstem Careers' tool. It turns a job posting and someone's raw work history into a tailored resume and cover letter. Serafina Vos runs product for it, and Odessa Fairweather, eight weeks into her job search, is one of the people typing into it every morning."
Why this works
Grounds a detection question in a real product and a real person before naming a single signal.
2
Reframe the question before naming a signal
Say it like this
"Here's how I'd frame it. This isn't really 'is the feature getting used.' It's 'is what ships still the model's work, or did someone quietly build their own version of this feature right next to it, and just keep clicking through mine on the way.' Only the second question is worth answering."
Why this works
Separates usage from substitution, which is the actual thing this question is testing.
3
Name why the obvious metric would miss it
Say it like this
"If I just watched daily active use, or how many people hit submit, I'd see nothing wrong. Odessa opened Wordbridge every single morning of her search, for the entire eight weeks. A workaround doesn't look like someone leaving. It looks like someone still there, using less of what you actually built than your dashboard thinks she is."
Why this works
Names the alternative most people reach for, usage or completion rate, and says plainly why it fails here.
4
Give the one decision, the actual detection method
Say it like this
"So here's what I'd build. Score each generated section on its own, bullets, opening line, closing line, comparing the model's draft against what actually gets submitted. Not one score for the whole document. When one section keeps coming back edited past a real cut-off, checked against known workaround examples, while the rest stays close, that section is the one being worked around."
Why this works
This is the concrete, defensible thing you'd build. It matches the direct answer word for word.
5
Prove it with the failure it would have caught
Say it like this
"With Odessa, her bullets stayed close to what Wordbridge wrote, around 88 percent overlap, steady for weeks. Her opening line fell to 9 percent by week three, because she was rewriting it by hand in a separate document and pasting it back in before she submitted. A single score for the whole document would have averaged those two numbers together and told me everything was fine."
Why this works
Makes the abstract signal concrete with the real numbers behind this exact answer.
6
Say what you'd do with the flag, not just watch it
Say it like this
"Once I know it's the opening line specifically, I'd give Odessa a 'regenerate just this' button, so fixing it never means leaving the app. And every flagged section becomes a labeled example for the model team, a real case of what isn't landing, instead of sitting in a Google Doc none of us will ever see."
Why this works
Shows the detection leads somewhere. A metric nobody acts on is a chart on a wall.
7
Say what you would not flag
Say it like this
"If someone swaps in their own real number inside a bullet, her actual four-minute wait-time cut instead of the model's guess, I wouldn't count that as a workaround. That bullet just got more true. I'd only flag a section where someone replaces the model's whole structure and voice, not one where they correct a fact."
Why this works
Shows judgment. Without this, the detector starts punishing the exact personalizing it should reward.
8
Close on the rule, in one line
Say it like this
"So the rule is: don't watch whether people open the feature, watch what they actually keep from it, section by section, and build the fix into the exact section the number points at."
Why this works
Closes on something reusable for a completely different feature you've never seen before.
If you remember one thing
A person opening your AI feature every day is not proof it's working. It only proves they're still standing in front of it. Watch what they keep, not how often they show up.
Let's learn
Here is a strange kind of success. A person opens your tool every single day, uses it for exactly the task it was built for, and still walks away with almost none of what it actually offers. That is what a working around looks like from the inside, and it is nearly invisible from a dashboard.
Say we build a tool that reads a job posting and someone's raw work history, and turns them into a tailored resume and cover letter. Wordbridge does this for Northstem Careers' users. Before a tool like this, someone hunting for a job spent about fifty minutes hand-writing and tailoring a resume and cover letter for each posting, which meant six or seven applications a week was close to the ceiling for a person doing it alone, on top of everything else in their week.
With Wordbridge running, that same person gets a full tailored draft, bullets rewritten to match the posting, an opening paragraph, a closing line, in about ninety seconds, then spends a few minutes checking it before sending. That is enough to push someone from six or seven applications a week to twenty.
Knowledge spark: what is a section similarity check?
A way to compare two pieces of text and score how close they are. Compare what the model wrote for a section against what the person actually submitted. High means she kept it close to the model's words. Low means she rewrote most of it herself.
Here is the turn. Twenty applications a week is not the problem. The real problem is what a person does with one specific part of what Wordbridge writes. Her bullets keep matching the model's draft closely, week after week. But her opening paragraph, the two or three lines meant to hook a recruiter before they read anything else, gets rewritten a little more each week, in a document Wordbridge never sees, then pasted back in right before she hits submit.
Section overlap with the model's own draft, by week three
A whole-document score would have averaged these two numbers into something that looked fine. Split by section, one number tells you exactly where the tool stopped being trusted.
That week-three snapshot only tells half of it. The drop didn't happen overnight, and it started well before anyone would have thought to look.
Opening line overlap, three weeks, against the flat bullet line
The drop starts weeks before anyone on the product team would have noticed anything, because submit counts kept climbing the entire time.
We did not lose twelve minutes of Odessa's evening. We lost the only place her real opening line was ever going to exist, where anyone but her could see it.
At its worst, this costs more than one person's evening. If this pattern holds across most of Wordbridge's users, the product that gets sold as a resume and cover letter assistant quietly becomes just a bullet-rewrite tool, with a summary generator nobody actually uses as intended, while the numbers leadership sees, applications sent, drafts generated, keep climbing and hiding the whole thing. A per-section check adds real cost: three similarity checks per submission instead of one. But a single whole-document check would have buried the exact section that broke inside an average that still looked healthy, which makes the extra cost worth paying.
The failure hiding under the hood
Wordbridge's opening line was tuned against Northstem's own internal quality grader, an automated score built to pick which prompt to ship. That grader rewards formal, keyword-dense summaries. It was never checked against what actually gets a recruiter to keep reading. The model got very good at a score that was never the real target, and nobody noticed because the score kept going up. The guardrail: sample real submitted openings against a small human-rated set, real hiring reviewers, not the internal grader, on a rolling basis, and watch for the two scores drifting apart.
The choice I would take back is not the opening paragraph's wording. It is that Wordbridge generates a whole application, bullets, opening, closing, as one model call with one "regenerate everything" button, and no way to keep four fifths of a draft while redoing the rest. That was a fine choice when Wordbridge's first users applied to two or three jobs a week and just wanted one clean draft, fast. It stopped being fine the moment someone started sending twenty a week and noticed the same paragraph shape every time.
What I would leave alone: a small dip in bullet-level overlap, when someone swaps the model's estimate for their own real number. Odessa's bullets moved from 91 percent overlap to 88 percent the week she corrected her actual wait-time savings into one line. That is the resume getting more honest, not less trusted, and chasing it back up to 100 would be chasing the wrong thing.
The lesson: a person who keeps opening your tool every morning has not told you it's working. She's only told you she's still looking for a job. Ask what she keeps, not how often she shows up.
Now here is the same thing as a story
The short version is above. Read on if you want to feel how ordinary the week was that this almost slipped past on.
Odessa Fairweather can tell you, in under thirty seconds, exactly how she trimmed four minutes off the average prescription wait at the pharmacy counter she ran for eight years at Corrivale Drug. The reorder trigger she changed, the pickup line she rearranged, the exact week it started working. When her district closed in the spring, she carried that same habit into her job search: she doesn't guess about her own work, she can name it precisely.
She found Wordbridge in her second week of looking. For about a month, 8am was the best part of her day. She'd paste in a posting for an operations coordinator role, and by 8:04 she had bullets rewritten to match it, an opening paragraph, a closing line. In week one she read every word of that opening carefully, changing one or two things. By week two she was skimming it, sending it mostly as written, because it kept being fine.
Then, on a Wednesday evening at a job-search meetup in a library conference room, a woman two seats over glanced at Odessa's laptop screen and said one line: "You're still letting it write your first paragraph? Mine used to sound exactly like that. Recruiters don't read past it."
Odessa didn't argue. She went home, opened her next application, and instead of editing Wordbridge's opening paragraph in place, she deleted it entirely, opened a blank document, and wrote her own three lines by hand. Then she pasted them back into Wordbridge before she submitted. She did that again the next morning. And the one after that. Every single time, from that Wednesday on.
The similarity score barely moved from week to week. The habit did not drift. It snapped once, on a Wednesday, and never went back.
She never had a number in her head for any of this. She had a feeling with exactly two settings: use what Wordbridge wrote, or step around it and write her own somewhere else. There was no version where she edited the paragraph a little more carefully. Once the switch flipped, it stayed flipped.
Wordbridge's team designed for a dial. What every job seeker actually got, the day the opening paragraph stopped landing, was a switch.
Months before any of this, when Serafina Vos's team scoped the first version of Wordbridge, the design meeting was short. The whole application, bullets, opening, closing, would come back from one model call, with one button to regenerate the entire thing if it wasn't right. Separate controls for each section were on the list. They got pushed to "v2, if anyone asks," because the earliest users, applying to two or three jobs a week, said having one clean draft fast mattered more than fiddly controls. Nobody in that room was picturing someone sending twenty applications a week and noticing the same paragraph shape on the fifteenth one.
What Serafina actually did, once the per-section similarity check flagged the pattern across dozens of accounts like Odessa's, not from a support ticket, since nobody ever filed one: she shipped a "regenerate just this" control on each section, and had every flagged opening routed to the model team as a labeled example instead of vanishing into someone's personal document. Three weeks later, the share of submitted applications with a near-empty opening-line overlap dropped from 61 percent to 18 percent. Odessa stopped leaving the app entirely. Her opening line still cost her about a minute of edits, not twelve, and for the first time since her search began, the paragraph she actually sent lived somewhere Wordbridge's own team could see it.
What I would tell myself, back in that first design meeting: one generate button feels like the simple choice. It stays simple right up until the moment someone needs to keep four fifths of what you made and rebuild the rest somewhere you can't see.
The five letters, if you want to carry them into the room
This is a detection question wearing a diagnosis question's clothes, which is exactly what makes FLIPS fit: find the person, find what they quietly stopped doing, and the flip itself is the signal you're actually being asked to catch.
The five moves in order. The middle one, the flip itself, is the only hard step. Everything else is setup and payoff.
F
Find the person. Whose morning is this?
Odessa Fairweather, eight weeks into her search, opening Wordbridge at 8am, twenty applications a week.
L
Locate the habit. What did she stop doing because it worked?
She stopped reading Wordbridge's opening paragraph closely, because for a month it kept being fine.
I
Identify the flip. What verb snaps, with no middle setting?
She goes from sending the model's opening line to always deleting it and writing her own in a separate document, then pasting it back in. Every time, from one Wednesday on.
P
Pinpoint the old decision. Which choice only made sense before?
Building the whole application as one model call with one regenerate-everything button, no separate control per section.
S
Show the replay. Same bad week, new design, better ending?
A per-section similarity check flags the opening line specifically. A "regenerate just this" control fixes it inside the app. Flagged rate falls from 61 percent to 18 percent in three weeks.
Knowledge spark: why pick the workaround flip family here, over pre-editing or concealment?
Pre-editing is about sanitizing the input before the model ever sees it. Odessa's input, her posting and her work history, never changed. Concealment is about hiding that a tool was used at all, from other people. Odessa never hid Wordbridge, she used it every morning in the open. What she actually did was keep the tool in her workflow and build a private, separate step around the one part it wrote badly. That is the workaround family's exact shape: the tool as the workflow, replaced by a personal process built around it.
And if you want to be sure it really works, try it somewhere else
Same method, a different flip family, a product nowhere near job hunting, so the method proves itself instead of repeating a story you happened to prepare.
Grenshaw Veterinary Network runs Scanlight, an AI tool that flags likely fractures and soft-tissue abnormalities on x-rays for rural vet clinics, so a vet tech can triage which images need a vet's eyes first.
F, find the person. Junia Marchant, a vet tech at a two-doctor clinic, uploads about thirty x-rays a week. L, locate the habit. She used to review every full image Scanlight flagged, edges included, before deciding what to escalate. I, identify the flip. A different family this time, pre-editing, not workaround. She learned Scanlight kept flagging harmless shadows near the image edges as false positives, so she started cropping every x-ray tight around the joint before uploading it, every single time, rather than uploading the full image. P, pinpoint the old decision. Scanlight was built to score the whole image at once, with no way to tell it "ignore the edges, I already checked those," so cropping the input was the only lever Junia actually had. S, show the replay. Scanlight adds a "mark this region as reviewed" tool instead of forcing a crop. Junia stops trimming images, Scanlight keeps seeing the full x-ray, and the edge cases she used to crop away, the ones a full image would have caught, stop disappearing before the model ever sees them.
Same shape, different stakes
At Wordbridge, the thing quietly disappearing was a paragraph a recruiter never got to read as the model wrote it. At Scanlight, it's the part of an x-ray a vet never gets flagged on at all. Different flip family, same lesson: watch what's missing from what the model actually sees or actually ships, not just whether the button got clicked.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the method: score each section against what's actually submitted, not the whole document at once, and say why usage numbers alone would miss it.
Cost: engineering says a per-section check triples the similarity calls versus one whole-document check. Don't cave and ship the cheaper coarse version. Say plainly what it would have hidden, the exact averaging that let Odessa's pattern hide for three weeks.
The model got better, for real: say Wordbridge's opening-line writing improves and overlap climbs back to 70 percent for most users. That's still not proof nobody's routing around anything. A better model can still get bypassed out of habit, which is exactly why the check stays on after the fix ships, not just during the incident.
Where people run it wrong.
They treat "usage is up" as proof the feature is working, instead of asking what fraction of the output actually survives to what gets shipped.
They build one score for the whole document, which averages away the one section that's actually broken.
They flag any edit at all as suspicious, which punishes honest personalizing and trains people to stop trusting the flag.
How to use it live. Say the reframe before naming a signal: "I wouldn't watch whether people open the feature. I'd watch what fraction of its output survives, section by section." That buys you room to give the real answer instead of reciting "track engagement" on reflex.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
What flip family is this, and what's its one-line shape?
Tap to flip
ANSWER
The workaround flip. The person uses the tool as the workflow, then quietly builds a private process around the one part it does badly, without ever leaving or hiding the tool itself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Odessa Fairweather, an eight-year pharmacy shift supervisor at Corrivale Drug, job hunting using Wordbridge, Northstem Careers' resume and cover letter assistant.
3 · THE HABIT
What did Odessa stop doing because it worked?
Tap to flip
ANSWER
She stopped reading Wordbridge's opening paragraph closely, skimming it and sending it mostly as written, because for about a month it kept being fine.
4 · THE FLIP
What's the two-setting switch in this story?
Tap to flip
ANSWER
Send the model's opening line as written, or always delete it and write her own in a separate document, pasting it back before submitting. No middle setting, and it never flipped back.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building the whole application as one model call with one regenerate-everything button. It made sense when early users applied to two or three jobs a week and wanted one clean draft fast, not fiddly per-section controls.
6 · THE NUMBER
Fill in the blank: by week three, Odessa's bullets held about ___ percent overlap with the model's draft, while her opening line had fallen to about ___ percent.
Tap to flip
ANSWER
88 percent for bullets, 9 percent for the opening line. A whole-document score would have averaged those into something that looked fine.
7 · THE REPLAY
Same pattern, new design. What changes, and what's the countable result?
Tap to flip
ANSWER
A per-section similarity check flags the opening line, a "regenerate just this" control fixes it inside the app. The flagged rate falls from 61 percent to 18 percent within three weeks.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Scanlight, Grenshaw Veterinary Network's x-ray triage tool. Its flip is pre-editing: a vet tech crops x-rays tight to dodge false-positive flags near the edges, hiding real cases from the model before it ever sees them.
Check yourself Score: 0 / 0
True or false
1. True or false: once Odessa noticed the opening line problem, she could have fixed it just as well by reading Wordbridge's draft a little more carefully instead of deleting it and writing her own elsewhere.
True
False
Show hint
A flip has exactly two settings and no middle. Ask whether "read it more carefully" is really a third option.
Show answer
False. There was no middle setting once the switch flipped. She either sent the model's line as written, or always deleted it and wrote her own somewhere else. "Read it more carefully" is a dial, and dials aren't what happened here.
Multiple choice
2. What was the flip in Odessa's story, and what were its two settings?
A. She went from applying to six jobs a week to applying to twenty.
B. She went from trusting Wordbridge completely to checking every single bullet by hand.
C. She went from sending the model's opening line as written to always deleting it and writing her own in a separate document before pasting it back in.
D. She went from using Wordbridge daily to only opening it once a week.
Show hint
The flip is a verb she performs on one specific section, not a change in how often she opens the app.
Show answer
C. A and D describe usage, which stayed healthy the whole time. B describes a different flip family entirely, checking everything. The real flip is narrower and quieter than either.
Fill in the blank
3. Within three weeks of shipping the per-section flag and a "regenerate just this" control, the share of submitted applications with a near-empty opening-line overlap fell from ___ percent to ___ percent.
Show hint
Check the end of the story section.
Show answer
61 percent to 18 percent. The fix worked because it removed the reason leaving the app was ever the easiest option, not because anyone told users to stop.
Short answer
4. What old product decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the design meeting memory, not a dial anyone could just turn back up.
Show answer
Model answer: Building the whole application, bullets, opening, closing, as one model call with a single "regenerate everything" button, and no separate control per section. It made sense because Wordbridge's earliest users applied to only two or three jobs a week and wanted one clean draft fast, not fiddly controls nobody had asked for yet.
Short answer, apply it yourself
5. Think of an AI feature you use yourself, an autocomplete, a summarizer, a chat assistant, anything. What's a part of its output you've quietly learned to always replace or route around, while still using the rest of it as intended?
Show hint
Look for a habit you built without ever filing a complaint about it.
Show answer
Model answer: A code assistant that writes solid function bodies but always generates a generic, over-explained comment above each one. A developer might learn to always delete and rewrite that comment by hand, while accepting the function body as written, every single time, without ever reporting it as a bug.
Short answer
6. Name a place in Wordbridge where a person editing the model's output a lot would NOT be a sign of a workaround. Why not?
Show hint
Look at "what I would leave alone" in the first section.
Show answer
Model answer: The bullets, when someone swaps in their own real number for the model's guess, the way Odessa corrected her wait-time savings. Her bullet overlap dipped from 91 to 88 percent that week, but the resume got more honest, not less trusted. Flagging that as a workaround would punish the exact personalizing the tool should want.
Before you close the answer
Why this works
Tests whether you'll reach for the usage metric everyone already tracks, or build the one signal that actually survives a person who never stops opening the app.
Follow-up traps
"Isn't a low similarity score just someone personalizing the resume, not a workaround?" Response: that's exactly why the check flags structure and voice replaced wholesale, not any edit at all, and why a small dip like Odessa's bullet score moving from 91 to 88 percent gets left alone on purpose.
"What if the model's writing genuinely improves and overlap climbs back up?" Response: keep the check running anyway. A better model can still get bypassed out of habit, and the check is what tells you whether trust actually came back or people just haven't noticed yet.
If pressed
The similarity check compares meaning, not just matching words, so a section that's been fully reworded in different phrasing still scores low. It runs in the background after the person submits, not before the send button, so it costs the backend real work but never makes anyone wait. And the flag line isn't one fixed number. It's a band, checked against a set of confirmed workarounds: under 30 percent gets a medium-confidence flag, under 15 percent gets a high-confidence one, and both get checked again against real outcomes every few months as the model changes.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.