Artifact critiqueIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #7

What is the honest way to describe your AI feature's limitations in marketing copy?

GUARD · the six months between one headline and a quiet cohort's churn

Pebblewise is Applewood Learning's AI tutoring app for kids, math and writing both. Magnus Jankowiak, who runs growth, tested a new headline: "One AI tutor. Every subject, automatically personalized." Florabel Eskander, Pebblewise's product manager, owns what the tutoring engine actually does on each side of that claim. Kazimiera Nowosielski is the parent who found out, four days before her son's scholarship deadline, that the two sides weren't the same.

The direct answer
Name the real scope in the copy itself. Say plainly what the feature does on its own and where it needs a person to check, for example "Math: instant and adaptive. Writing: a fast first pass, reviewed by you or a teacher before it's final," instead of one blanket line like "every subject, automatically personalized." Put that scoping in the headline, not a footnote, and treat it as something that earns trust, not something to hide.
Do this, in order
  1. Name the real scope in the headline, math versus writing, instead of one blanket claim.Why: a family reading "every subject, automatically" has no reason to expect a gap, so nobody looks for one until it costs them something real.
  2. Check what the model can actually back up, subject by subject, before any line ships.Why: Applewood only learned the real writing catch rate once someone ran the audit. Copy shouldn't get ahead of a number nobody's checked yet.
  3. Track churn and support tickets split by which feature a family actually leans on, not the blended average.Why: the blended number looked fine for months. The writing-only number told the real story weeks earlier.
  4. Put the scoped line in the headline and the app store copy, not a footnote or a terms page.Why: a parent reading marketing copy has no reason to go looking in the fine print for what it leaves out.
  5. Leave the math messaging exactly as confident as it already is.Why: it's earned. Scoping down a claim that's already true just to look consistent costs signups for nothing.
  6. Give the honest version real time to prove itself on trust, not just watch signups alone.Why: a scoped claim can cost a little conversion up front and still be the right call, if what it buys back is families who stay.

How to answer this, stage by stage

Nobody is grading whether you think marketing copy in general is too aggressive. They're grading whether you can turn "be honest about limitations" into an actual sentence you'd ship.

1
Scope it to one product and one piece of copy
Say it like this
"Let me make this concrete. Say a company builds an AI tutoring app for kids, math and writing both. Marketing writes one headline for the whole thing: 'every subject, automatically personalized.' That's the line I'll test, because 'be honest about limitations' means nothing until there's an actual sentence to point at."
Why this works
Keeps the answer from turning into a lecture about honesty in general.
2
Say your structure out loud
Say it like this
"I'll run this as GUARD. Groups: who's actually carrying risk here. Unequal: where that risk lands hardest. Ability to contest: can a parent reading the copy tell what's being left out. Reduce: the actual rewrite. Detect: how you'd know the copy drifted before a customer tells you."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about marketing.
3
Reframe what "honest" actually means here
Say it like this
"This isn't really about lying. Nobody wrote a false sentence. The problem is one true-sounding word, 'automatically,' doing two very different jobs, and the copy never says which."
Why this works
Separates this from a generic overselling answer and locates the real judgment: uneven model performance, not dishonesty in the moral sense.
4
Give the one decision
Say it like this
"Here's what I'd actually ship: name the real scope in the copy itself. 'Math: instant and adaptive. Writing: a fast first pass, reviewed by you or a teacher before it's final.' That goes in the headline, not a footnote."
Why this works
This matches the direct answer word for word. If it doesn't, the interviewer notices before you do.
5
Read the actual copy, before and after
Say it like this
"The line that shipped was: 'One AI tutor. Every subject, automatically personalized, so your child gets exactly the help they need, the moment they need it.' Here's what I'd put up instead: 'Pebblewise adapts instantly to your child's math level, from times tables to fractions. For essays and stories, Pebblewise gives a fast first pass, built to be reviewed by you or a teacher before it's the final word.'"
Why this works
An interviewer testing this exact question wants to hear the sentence rewritten, not a description of rewriting it.
Flawed: shipped six months ago
Landing page hero"One AI tutor. Every subject, automatically personalized, so your child gets exactly the help they need, the moment they need it."
App store screenshot caption"Pebblewise reviews your child's essay in seconds, so you don't have to."
Onboarding email"Sit back. Pebblewise has writing covered too. From times tables to term papers, Pebblewise adapts to your child in real time. You don't need to check the work. We've got it."
Honest: the rewrite, month five
Landing page hero"Pebblewise adapts instantly to your child's math level, from times tables to fractions. For essays and stories, Pebblewise gives a fast first pass, built to be reviewed by you or a teacher before it's the final word."
App store screenshot caption"Math: instant, adaptive feedback. Writing: a quick first read to get your child started, best paired with a teacher's look before a big deadline."
Onboarding email"Here's exactly what Pebblewise does for each subject. Math: fully automatic, and it gets sharper every time your child practices. Writing: fast, encouraging first notes. For anything graded, like a scholarship essay, have a teacher or you take one more look."
6
Prove it with the compressed failure
Say it like this
"Here's what happens without the rewrite. A parent signs up believing writing works exactly like math. Her son's scholarship essay gets called 'in great shape' by the app. His real teacher finds the opposite two days later. They've got four days left, not the quick polish they planned for."
Why this works
This is the story below, cut to four sentences, so the cost lands on a real family, not an abstraction.
7
Say what you'd watch, and what you'd leave alone
Say it like this
"I'd track churn and support tickets split by which feature a family actually uses, not the blended number, because the blended number can look fine for months while one cohort's already leaving. And I'd leave the math copy alone. It's not overclaiming anything."
Why this works
Shows judgment about where the real risk sits, not blanket caution slapped on every line of copy.
8
Close on something checkable
Say it like this
"You'll know it's working when a parent can read the headline and correctly guess which parts need their own attention before they ever open the app. You'll know it's broken when one cohort's churn keeps climbing while the average keeps looking fine."
Why this works
Ends on a test the interviewer could actually go verify, not a promise that it's handled.

Let's learn

Pebblewise lives on a tablet at the Nowosielski family's kitchen table, the one with a hairline crack across the corner where Tibor's little sister dropped it back in March. Two things happen on that tablet most weeks: times tables before dinner, and whatever essay is due that Friday.

Hand sketched labeled parts diagram titled Pebblewise dot com, month three. A central document icon labeled Pebblewise dot com, with five callouts around it: Headline says every subject automatic, Math screenshot marked instantly, Essay screenshot marked in seconds, No line about a review needed, One Start Free Trial button.
Five things on the page. Four of them say what the app does. None of them says where a parent still needs to check.

Pebblewise is an app that watches a kid work through math problems and writing drafts, and hands back help on both, right there on the screen. Applewood Learning checked how well it actually does that, on both sides, and the two answers are not close.

On 500 real math problems, checked against a certified answer key, Pebblewise caught the mistake and explained the right method correctly 486 times. That's 97 percent, and it's the number the whole company points to when anyone asks whether the tutoring really works.

Knowledge spark: why is grading an essay harder for a model than grading a math problem? A math answer has one right answer to check against. An essay doesn't. There's no single correct paragraph to compare a draft to, so the model has to judge quality instead of correctness, and a model trained to sound encouraging with kids tends to lean generous exactly where the answer is genuinely open-ended.

Writing is a different animal. Applewood pays a contracted literacy teacher, Zelig Kirchner, to hand-mark a sample every quarter. Her last batch was 40 real drafts from working students. She found an average of six real problems in each one: a thesis that never quite lands, a middle paragraph that wanders off, a claim with nothing backing it up. Pebblewise's own feedback caught an average of two of those six. In 14 of the 40 essays, on drafts Zelig had marked as needing real work, Pebblewise opened with something like "Great structure, nice work."

Hand sketched quadrant diagram titled Where Pebblewise is sure, where it's guessing. X axis, how checkable the answer is, from open ended to one right answer. Y axis, how well Pebblewise actually does, from weak to strong. Times tables, fractions, and spelling check sit in the strong, checkable corner. Essay structure and persuasive writing sit in the weak, open ended corner.
Two of these five sit on solid ground. Two of them are the model's best guess, dressed up exactly the same way as the other three.

Here's the turn. Missing four problems out of six isn't, on its own, the thing that hurt anyone. Magnus's headline is. "Every subject, automatically personalized" told every parent that writing worked the same way math did. So nobody went looking for a gap the copy never admitted was there.

We never wrote a false sentence. We let one true-sounding word cover two very different truths.

At its worst, this is Kazimiera Nowosielski two days before her son Tibor's scholarship deadline, reading a note from his actual teacher that flatly contradicts what Pebblewise told them a week earlier: "in great shape, just polish the ending." They had budgeted two hours for a polish. They had four days left, and a middle section that needed to be rebuilt.

The choice I would take back Magnus's headline was fine the day it shipped. Conversion was the only number anyone had, and the writing side hadn't been checked against a real teacher's read yet. It stopped being fine the month the essay-side numbers existed and kept getting worse for one cohort while the blended average kept looking fine.
What I would leave alone The math half of the copy never needed touching. It says exactly what it does, and 97 percent is a real, checked number, not a guess dressed up as one.

The lesson: "automatically personalized" is not one claim. It's two, wearing the same three words. If a model does two genuinely different jobs, the copy has to say so twice, not once, or somebody finds out the hard way, at exactly the moment it matters most.

Now here is the same thing as a story

Stage six above compresses this into four sentences. Here's the six months underneath, the part a stand-up answer skips.

Florabel Eskander has spent three years owning what Pebblewise's tutoring engine actually does, not what it's supposed to do. She can tell you, without opening a dashboard, which grade levels struggle with fractions and which ones struggle with topic sentences, because she reads the model's own accuracy numbers most mornings before her coffee's gone cold.

Six months ago, Magnus Jankowiak, who runs growth, tested a new headline on the landing page: "One AI tutor. Every subject, automatically personalized, so your child gets exactly the help they need, the moment they need it." Signups jumped. Trial-to-paid conversion went from 2.6 percent to 3.5 percent inside three weeks, the best number the page had ever posted. The line went everywhere: the hero, the app store screenshots, the welcome email. Nobody thought twice, because nothing in the numbers said to.

Hand sketched timeline diagram titled From one headline to a quiet cohort's churn. Five milestones: Headline ships, month one, signups jump. Math families thrive, months one to three. Writing families slip, months two to four. Lennox forwards five notes, month four, a Thursday, emphasized. Copy rewritten, tracked apart, month five.
Four of these five months looked completely fine, if the only thing you were watching was the average.

For three months, that looked exactly right. Math-heavy families reviewed the app well and stayed. Writing-heavy families seemed fine too, at least in the one number Florabel was watching then, a single blended churn rate for the whole app.

Parents who leaned on the writing feature did what new users always do at first: they read Pebblewise's notes on their kid's draft, then read the draft themselves, just to check. After a few weeks of the notes sounding right, most stopped reading the draft first. By month three, most weren't reading it at all. The headline had told them they didn't need to.

In month four, on an ordinary Thursday, Lennox Quirke, three weeks into the support job, forwarded Florabel a batch of five cancellation notes. He didn't flag it as urgent. He just wrote, "these all say almost the same thing, is that normal?" Four of the five used some version of the same sentence: the app said it was basically done.

Florabel pulled the churn number apart by feature that afternoon, not by the whole app. Math-heavy families were churning at 3 percent, flat, same as month one. Writing-heavy families had climbed from 4 percent in month one to 9 percent in month four, quietly, underneath a blended average that had barely moved.

Hand sketched comparison diagram titled Who holds the lever, who doesn't. Left, a person icon labeled Applewood Learning, caption writes the words, sets the claim, ships the update. Right, a person icon labeled A parent like Kazimiera, caption reads the words, no way to check what got left out.
Neither person here did anything wrong. Only one of them could see the whole page before it went out.

One of the five notes was from Kazimiera Nowosielski. Her son Tibor, eleven, had a scholarship application essay due that Friday. Pebblewise's feedback on his draft, five days earlier, opened with "In great shape, just polish the ending." His actual teacher read the same draft two days later and said the middle section didn't support his thesis at all. Kazimiera had budgeted two hours for a polish. She had four days left, and a real rewrite in front of her.

Florabel ran the writing audit that weekend, the one Zelig Kirchner runs every quarter. Forty real drafts, hand-marked, an average of six genuine problems per essay. Pebblewise's feedback caught two of those six, on average, and in fourteen of the forty cases opened with encouragement despite the real problems sitting right there in the draft.

We didn't take four days from Kazimiera. We took the two hours she thought she had.

The decision Florabel would take back sits six months earlier, in a fifteen-minute meeting about the new headline. Magnus showed the conversion lift. Someone asked if "automatically personalized" was accurate for writing too. Florabel said the writing model was "directionally fine" and the meeting moved on. Nobody asked directionally fine by what measure, because nobody had run the audit yet, so there was no number to ask about.

So here's the replay. Same six months, honest copy from month one: "Math: instant and adaptive. Writing: a fast first pass, reviewed by you or a teacher before it's final." Landing conversion settles at 3.1 percent instead of 3.5, a real but small cost. Writing-heavy churn never climbs past 4 percent, because parents like Kazimiera are already reading Tibor's draft themselves, the way the copy told them to. His scholarship essay gets a real teacher's read nine days before the deadline instead of two.

One design tells a parent nothing, so their trust has nowhere to break cleanly. The other tells them exactly where it can, so it never does, quietly, at the worst possible moment.

What I'd tell myself, back in that fifteen-minute meeting: "directionally fine" isn't an answer, it's a guess wearing a confident voice. If I didn't have the number, I should have said so, out loud, instead of letting the meeting move on.

What GUARD actually catches in one headline

This was never really about whether Applewood lied. GUARD is for naming who carries the real risk when one word covers two very different truths, and turning that into a rewrite you'd actually ship.

GGroups. Who's carrying real stakes, and what stakes.
Families like Kazimiera and Tibor Nowosielski, who subscribed on the strength of "every subject, automatically personalized" and had no way to know that meant something very different for essays than for times tables. And Applewood Learning itself, whose growth numbers are real and worth protecting, and whose writing-feedback model genuinely does help as a first pass, even though it isn't a finished grade.
Name both sides as real before picking a fix. The company's growth number and the family's trust are both worth something here.
UUnequal. Where the harm lands, and why that group specifically.
The harm doesn't spread evenly. Families whose kids mostly practice arithmetic get almost exactly what the copy promised: fast, mostly correct, automatic. Families whose kids lean on Pebblewise for essays, especially grades five through eight, hit the gap the copy never named. And it's invisible until it isn't: nothing on the screen tells a parent which side of that line their own family sits on.
The blended churn number stayed flat-looking because math-heavy families make up most of the base. The writing-heavy cohort's real number was climbing the entire time, underneath it.
AAbility to contest. Can a parent tell what's being left out.
A parent reading a landing page has no way to check what's actually being scoped out. The copy is written to persuade, not to disclose. Unless a parent independently keeps double-checking their kid's essay against a real teacher's read, the exact behavior the copy told them they no longer needed, they don't learn where the tool's real edge sits until a moment like Kazimiera's.
This is GUARD's sharpest question for this exact scenario: not "is this technically true," but "could the person reading it have told the difference."
Hand sketched flow diagram titled Where a parent's way to check should sit, and doesn't. Five connected boxes reading Copy says everything, She stops checking, Model over-praises, No warning shown, Found out too late, with the fourth box emphasized in orange.
Four of these five steps happened exactly as the copy was written to make them happen. The missing one, a plain warning, was never built.
RReduce. The actual rewrite, not a policy memo.
Write the copy to name the real scope: "Math: instant and adaptive. Writing: a fast first pass, reviewed by you or a teacher before it's final." Put it in the headline and the app store copy, not a footnote or a terms page. Treat the honesty as a trust asset, not a weakness to bury.
The alternative worth naming and rejecting: pull essay feedback out of the marketing entirely and only ever describe Pebblewise as a math tool, leaving writing as a quiet, unadvertised extra. That fully avoids the overclaim. It also ignores that 61 percent of weekly sessions for grades five through eight already lean on the writing feature, so families would depend on it just as much, they'd just walk in with no idea what to expect from it.
Hand sketched icon list diagram titled What honest, scoped copy actually says. Four numbered rows: names what's fully automatic plainly, names what still needs a human look, says it in the headline not a footnote, treated as a trust asset not a weakness.
Four lines. None of them apologizes for what Pebblewise can't do yet. They just say it.
DDetect. How you'd know, before a parent tells you.
Track churn and support tickets split by which feature a family actually uses, not the blended average. Before the fix, writing-heavy churn climbed from 4 percent to 9 percent across four months while math-heavy churn stayed flat at 3 percent, and the blended number for the whole app barely moved.
The failure worth naming plainly: a calm-looking average can mean everything's fine, or it can mean one cohort is quietly leaving while a bigger, healthier cohort holds the mean steady. Only the split number tells you which.
Monthly churn by which feature a family leans on, before and after the scoped rewrite
8% 4% 0 3% 3% Math-heavy families 8% 4% Writing-heavy families
Before the rewrite, average of months 3 and 4After the rewrite, month 6
Math-heavy families never moved. Writing-heavy families were churning at nearly triple the math rate before the copy named the real scope, and dropped back to nearly the same rate as math after it did.
Writing-feedback support tickets per week, six months, before and after the rewrite
20 10 0 copy rewritten Month 1 Month 2 Month 3 Month 4, peak Month 5 Month 6
Tickets mentioning missed or wrong writing feedback, per week
Nineteen tickets in one week is what finally reached Lennox's inbox. The climb underneath it had already been running for three months.
The trade-off, said out loud The scoped copy cost Applewood something real: landing conversion settled at 3.1 percent instead of the overclaim's 3.5 percent, a genuine 0.4-point loss on the way in. Applewood took that trade on purpose, for the cohort that was carrying real exposure, in exchange for a writing-heavy churn rate that stopped climbing toward the door.

And if you want to be sure it really works, try it somewhere else

Same five letters, a home-appliance diagnosis app instead of a tutoring one, and this time the unscoped word is "any."

Wrenchly, built by Farnsworth Diagnostics, reads a photo of an appliance's error code and a few follow-up questions, then tells a homeowner what's actually broken before a technician shows up. Gunnar Hakonsson owns Wrenchly's diagnosis model the way Florabel owns Pebblewise's tutoring engine.

Hand sketched decision tree diagram titled Writing copy for an AI appliance diagnosis app. Root: how do we describe what Wrenchly can diagnose. Three branches: claim it diagnoses any appliance instantly leads to overclaims on older off brand units it can't read. Bury the real limits in a footnote leads to still overclaims nobody reads footnotes. Name which brands and codes it reads well leads to honest, trust holds at the off brand dryer.
Same three branches Applewood faced, a different appliance. The honest branch is the only one that survives a customer's own off-brand dryer.

Guglielmo Grimaldi, who runs growth at Farnsworth, tested a new app store line: "Wrenchly diagnoses any appliance, instantly, no technician needed." Downloads jumped. What the line left out: Wrenchly reads digital error codes from six major brands correctly about 94 percent of the time. On older analog dials and off-brand units with no digital code at all, its guess rate drops to around 40 percent, barely better than naming the most common cause.

Same rank, mapped onto Wrenchly: name the real split in the app store copy itself, not a support article. "Reads digital error codes from major brands instantly. For older or off-brand appliances, Wrenchly gives its best guess, confirm with a technician before buying a part." Track refund requests split by appliance age, the same way Florabel split churn by subject, since a blended refund rate would hide an old-dryer cohort quietly asking for its money back.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the real scope in the headline, not the footnote. A blanket claim isn't free, it just charges you later, in churn.
Cost: no budget to rebuild the whole copy deck this quarter. Start with the two lines a parent, or a homeowner, actually reads before buying: the headline and the first screenshot. Scope those two first.
The model got better, for real: say Pebblewise's writing feedback jumps to catching five of six problems next quarter. That's a reason to widen the claim on purpose, with a new audit backing it up, not a reason it should have been claimed six months early.

Where people run it wrong.
They treat this as a moral question, honest versus dishonest marketing, instead of a scoping question about where the model is actually strong.
They fix the copy once, after a complaint, instead of tracking the split number that would have shown the gap widening months earlier.
They let growth alone decide what the copy claims, when the actual claim is a product judgment about what the model can back up.

How to use it live. Before answering, ask yourself out loud: "does this word mean the same thing in every case it's applied to?" Say the answer for the specific case in the question, and the honest scope almost always falls straight out of it.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
Which framework fits "what's the honest way to describe your AI feature's limits in marketing copy"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real test isn't whether the copy is honest in general, it's who gets hurt when a blanket claim meets a model that performs unevenly across subjects.
2 · THE PEOPLE
Who are the people this answer names?
Tap to flip
ANSWER
Florabel Eskander, Pebblewise's product manager. Magnus Jankowiak, who runs growth and wrote the original headline. Kazimiera and Tibor Nowosielski, the family who hit the gap. Lennox Quirke, the support rep who forwarded the tickets. Zelig Kirchner, the teacher who ran the writing audit.
3 · THE HABIT
What did parents using the writing feature stop doing, because the copy told them they didn't need to?
Tap to flip
ANSWER
They stopped reading their kid's essay draft themselves before it went in. "Automatically personalized" for every subject made checking feel redundant, until it wasn't.
4 · THE REAL GAP
What's the two-setting gap the copy never named?
Tap to flip
ANSWER
Math: 97 percent correct on a 500-problem checked set, genuinely automatic. Writing: caught 2 of 6 real problems per essay on average, and praised 14 of 40 flawed drafts. One word, "automatically," covered both.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Approving "every subject, automatically personalized" as the headline, in a meeting where "directionally fine" stood in for a real number on the writing side, because nobody had run the audit yet.
6 · THE NUMBER
Fill in the blank: on the writing audit, Pebblewise caught an average of ___ out of ___ real problems per essay, and praised ___ of the ___ essays Zelig had flagged as needing real work.
Tap to flip
ANSWER
2 out of 6, and 14 of 40. That gap between what the copy promised and what the model actually did is the whole story.
7 · THE REPLAY
Same six months, honest copy from month one. What changes?
Tap to flip
ANSWER
Landing conversion settles at 3.1 percent instead of 3.5. Writing-heavy churn never climbs past 4 percent. Kazimiera reads Tibor's essay herself, the way the copy told her to, and his teacher sees it nine days before the deadline instead of two.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and who plays the equivalent roles?
Tap to flip
ANSWER
Wrenchly, Farnsworth Diagnostics' appliance-repair app. Gunnar Hakonsson plays Florabel's role. Guglielmo Grimaldi plays Magnus's role, writing "diagnoses any appliance, instantly" over a model that's 94 percent solid on major brands and near a guess on old analog units.

Check yourself Score: 0 / 0

True or false
1. True or false: the problem in this story is that Applewood's marketing copy contained a factual lie.
  • True
  • False
Show hint
Check stage 3 of the walkthrough, the reframe.
Show answer
False. Every individual claim was arguably true in isolation. The problem is that "automatically personalized" meant something accurate for math and something misleading for writing, and the copy never said which.
Multiple choice
2. Why does tracking one blended churn number hide the real problem in this story?
  • A. Because blended numbers are always the wrong way to measure churn.
  • B. Because math-heavy families' flat 3 percent churn buries writing-heavy families' climb from 4 to 9 percent inside an average that barely moves.
  • C. Because churn isn't a real metric for a tutoring app.
  • D. Because Lennox miscounted the support tickets he forwarded.
Show hint
Check the Unequal and Detect steps in the GUARD recap.
Show answer
B. A blended number isn't wrong in general, that's option A's mistake. It's wrong here specifically because the larger, healthy cohort was masking the smaller, worsening one.
Fill in the blank
3. On the 500-problem math eval, Pebblewise correctly caught and explained the error ___ times, a ___ percent catch rate.
Show hint
Check flashcard 4, and the "Let's learn" section's math numbers.
Show answer
486 times, a 97 percent catch rate. This is the number that makes the math half of the headline true, and the reason it needed no rewrite.
Short answer, name the rejected alternative
4. Applewood considered one other fix besides rewriting the headline. What was it, and why did they reject it?
Show hint
Check the Reduce step in the GUARD recap.
Show answer
Model answer: Pull essay feedback out of the marketing entirely and only ever advertise math. Rejected because 61 percent of weekly sessions for grades five through eight already leaned on the writing feature, so hiding it from the copy wouldn't reduce how much families depended on it, it would just leave them walking in blind.
Short answer, apply it yourself
5. Think of a product or app you use that makes one confident-sounding claim covering several different things it does. Where would you expect that claim to actually hold up, and where would you expect it to quietly not?
Show hint
Look for a single word like "automatic" or "instant" doing double duty across two different tasks.
Show answer
Model answer: A budgeting app that says "automatically tracks every expense." It probably nails anything paid by card, since that's a clean data feed, and probably misses cash, checks, or split bills, since there's no clean signal for any of those at all.
Fill in the blank, work the number
6. Writing-heavy churn climbed from 4 percent in month one to 9 percent in month four, before the fix shipped, gaining roughly two points a month. If that same climb had continued for two more months instead of the rewrite shipping, about where would month six's churn have landed, and how does that compare to the actual result of 4 percent?
Show hint
Add roughly two points for month five, two more for month six, starting from 9 percent.
Show answer
Roughly 13 percent, versus the actual 4 percent. The fix didn't just slow the climb. It reversed it below where the climb even started.
Before you close the answer
Why this works
Tests whether you can turn "be honest in marketing" into a specific, checkable scoping decision instead of a values statement, and whether you actually know which part of the model is strong enough to promise and which part isn't.
Follow-up traps
"Isn't 'automatically personalized' technically true, since the model does personalize automatically?" Response: technically true and functionally misleading aren't the same thing. It personalizes math with a 97 percent catch rate and writing with a 33 percent one. One word covering both isn't a lie, but it isn't information either.

"Won't scoping the claim just cost you signups to competitors who don't bother?" Response: it cost about 0.4 points of landing conversion, 3.5 down to 3.1. That's a real cost, and Applewood accepted it because the alternative was losing writing-heavy families entirely once they found the gap themselves, at a worse moment than signup.
If pressed
The writing model's habit of over-praising weak drafts is a known pattern in models tuned to sound encouraging with kids. It shows up as confident wrongness, not silence, which is why usage numbers alone never caught it. The quarterly teacher-marked audit is the guardrail that actually does, since it checks the feedback against a real read, not just against how often families kept using the app.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more