Artifact critiqueIntermediateDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #13

What does honest marketing copy for an AI feature look like?

AUDIT the artifact is Pathforge's marketing page, an AI coding assistant Priyanka Reddy is evaluating for Verrick Systems

Verrick Systems runs a 40-person engineering team on a mixed Python, Go, and JavaScript codebase. Priyanka Reddy leads that team and was two weeks from recommending Pathforge, an AI coding assistant, on the strength of one number on its homepage.

The direct answer
Honest marketing copy names the eval set, the exact model version and date, and what the metric actually measures, not just a bare headline percentage. If a number can't survive "accurate at what, measured how, on what data, as of when," it's decoration wearing evidence's clothes, whether you're the one selling the claim or the one about to buy it.
Do this, in order
  1. State exactly what the metric measures, not just its name.Why: "98% accurate" and "98% of suggestions accepted" are wildly different claims wearing the same three words.
  2. Name the model version and the date the number was measured.Why: models change quietly, and an undated number can't be checked against anything, including its own later regression.
  3. Say what the number does not cover.Why: a claim's fine print usually reveals more than its headline ever does.
  4. Offer a way for a buyer to test the claim on their own real data.Why: a vendor confident in its own number should welcome exactly this, not resist it.
  5. Never claim a rate "improves as it learns" without saying whose data and how that's checked.Why: vague self-improvement claims are the easiest kind to write and the hardest kind to verify.

How to answer this, stage by stage

Nobody's asking you to write a style guide for marketers. They're asking whether you know what makes a number checkable versus just impressive.

Stage 1
Scope it to one real claim
Say it like this
"I'll answer this against Pathforge's homepage, an AI coding assistant claiming '98% accurate code suggestions.'"
Why this works
Turns "honest marketing" from an abstract virtue into a specific claim you can actually test.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself."
Why this works
Signals you're running a real checklist, not just declaring the claim suspicious on a hunch.
Stage 3
Ask who paid for it
Say it like this
"This is Pathforge's own number, about its own product, on its own homepage. That doesn't make it false. It does mean nobody neutral has checked it yet."
Why this works
Names the built-in incentive without accusing anyone of lying.
Stage 4
Uncover the eval set and demand the version
Say it like this
"The footnote says 'accuracy' but means 'suggestions accepted,' not 'suggestions later found correct.' There's no model version, no date, and no language scope attached to the number at all."
Why this works
Shows the actual gap between what the number sounds like and what it can prove.
Stage 5
Isolate what's missing
Say it like this
"Buried in the fine print, the number is Python-only. The hero image right above it shows JavaScript and Go. Nothing on the page tells you that up front."
Why this works
Names the exact gap between the visual pitch and the actual scope of the claim.
Stage 6
Test it yourself
Say it like this
"I ran Pathforge against our own repo for two weeks. Accepted-and-still-correct came out to 61 percent, not 98, and it varied hugely by language."
Why this works
Replaces trust in the vendor's number with a real, replicated one of your own.
Stage 7
Close on the one line
Say it like this
"A number nobody can check isn't a lie, necessarily. It's just not evidence yet. Honest marketing copy is the version of this claim that survives being checked."
Why this works
Restates the standard plainly, ready for whatever follow-up lands next.

Let's learn

What does 98% accurate actually mean, and why couldn't anyone at Pathforge answer that in one sentence when Priyanka finally asked?

Pathforge is an AI coding assistant. It suggests code as an engineer types, and its homepage leads with one number, in large type, above everything else: 98% accurate.

Knowledge spark: acceptance rate versus correctness rate Acceptance rate is how often a person clicks "accept" on a suggestion. Correctness rate is how often that accepted code is still right a week later, after review and tests. A suggestion can look plausible enough to accept and still be wrong. The two numbers are not the same claim.

Priyanka was two weeks from recommending Pathforge to her 40-person team, on the strength of that one headline number and a persuasive sales call.

The claim, and what Priyanka actually measured
100% 50% 0 Marketed: 98% Measured: 61%
The gap isn't a rounding error. It's the distance between "accepted" and "still correct a week later."

The turn: whether Pathforge's tool is actually good or bad was never really the question. The real problem is that the headline number couldn't be checked against anything at all, no eval set, no version, no scope, which means believing it and disbelieving it were both just guessing.

The decision I would take back Pathforge's marketing team published only the single most flattering number from internal testing, with no methodology attached, because a bare "98% accurate" headline tested better in ads than a fuller, honest sentence. That made sense while every competitor was doing the same bare-number marketing. It stopped making sense the moment a serious technical buyer asked "accurate at what, exactly?"

What I would leave alone: the rest of Pathforge's page, the product screenshots, the customer logos, the pricing table, doesn't need this treatment. The scrutiny belongs on quantified performance claims specifically, not on every sentence a vendor writes.

An unaudited number in marketing copy isn't really a claim about the product. It's a bet that nobody serious will ask where it came from.

The lesson: a vendor confident in its number should want you to check it. The moment a claim resists being checked, that resistance is the actual information, more than the number itself.

Hand sketched icon list titled The AUDIT checklist. Five items: who paid for this number, what eval set was used, which model version which date, what is missing from the claim, test it yourself on your own data.
Five questions, and Pathforge's homepage had a real answer to exactly none of them.

Now here is the same thing as a story

The short version above is what you'd say defending a "don't buy yet" recommendation to your own VP. Read this one for how Priyanka nearly didn't ask at all.

Before Pathforge: Priyanka's team wrote everything by hand, slower, and every review caught the mistakes a human made. After Pathforge's demo: a sales call so smooth she'd already started drafting the internal memo recommending it.

Priyanka has run a dozen tool evaluations in her career, and she's usually the skeptical one in the room. This time, the number did the convincing for her. Ninety-eight percent is a big number. It sounds finished. It sounds like someone else already did the hard part of checking.

Hand sketched timeline titled Priyanka's timeline. Four milestones: sales pitch 98 percent accurate headline number, colleague's remark nobody gets near that highlighted, she digs into the fine print no version no eval set, runs her own repo test 61 percent correct real number.
One remark, from someone with nothing to sell her, is what turned a headline number back into a question.

Then a former colleague, now at a different company, messaged her out of nowhere: "you're not actually buying that 98 percent number, are you? Nobody I know who's tried Pathforge gets anywhere near that."

Priyanka went back to the homepage that evening and read the footnote for the first time instead of skimming past it. It said "accuracy" but defined it, three lines down in smaller type, as suggestions accepted by a developer, not suggestions later confirmed correct. No model version. No date. No mention of which languages the number applied to at all.

Hand sketched decision tree titled Believe it or verify it. Root: a vendor publishes a headline number. One branch, no eval set or version named, leads to treat it as an ad. Other branch, eval set and version named, leads to worth testing further.
Pathforge's page sat entirely on the left branch. Priyanka had almost treated it like the right one.

What that would have cost, at its worst: a company-wide license purchase, an onboarding rollout across 40 engineers, and a real accuracy number discovered only after the fact, buried in code review backlash nobody could trace back to a marketing claim anyone had actually checked.

Hand sketched labeled parts diagram titled What is missing from most vendor claims. Center gauge icon labeled 98 percent headline, with four callouts: eval set name, model version, language scope, failure cases.
All four of these were missing from the same three words on Pathforge's homepage.

She spent two weeks running Pathforge against Verrick's own repo, a real mix of Python, Go, and JavaScript, and tracked every accepted suggestion against what code review and tests said a week later. The real number: 61 percent, and it varied hugely by language, far better on Python than on Go.

I almost recommended Pathforge off a number I never actually checked, because it was persuasive and the sales call was good. It took one skeptical remark from someone with nothing to sell me to remember that a number this important deserves the same scrutiny I'd give any claim in a pull request.

AUDIT, in one screenNot a fact-check ritual. AUDIT is what forces a checkable number apart from a persuasive one.

A
Ask who paid for it.
Pathforge's own marketing team, about its own product, with an obvious reason to pick the most flattering number available.
Names the incentive without assuming bad faith.
U
Uncover the eval set. The hard step.
"Accuracy" turned out to mean acceptance rate, not correctness, and no set or benchmark was named at all.
This is the gap the whole claim actually rests on.
D
Demand the version pin.
No model version, no date, on a homepage number for a product that ships new versions monthly.
An undated number can't be checked against a later regression.
I
Isolate what's missing.
Python-only, buried in fine print, under a hero image showing JavaScript and Go.
What a claim leaves out is usually louder than what it states.
T
Test it yourself.
Two weeks against Verrick's real repo, tracked for actual correctness, not clicks.
Replaces a vendor's number with your own, on your own real cases.
Pathforge's own benchmark score, by model version
100% 50% 0 71% 79% 98% 85% 87% 89% v1.0 v1.2 (marketed) v1.5 (current)
The marketed 98% was a one-time peak on v1.2, fourteen months ago. The version shipping today scores 89% on Pathforge's own benchmark, and 61% on Verrick's real repo.

The recap, one line per letter: ask who paid for it is Pathforge's own team with its own incentive, uncover the eval set is acceptance dressed up as accuracy, demand the version pin is a number with no date attached to a product shipping monthly, isolate what's missing is the Python-only scope buried in fine print, and test it yourself is the 61 percent that replaced the guess.

And if you want to be sure it really works, try it somewhere elseSame five letters, a streaming dubbing feature instead of a coding assistant. A different industry, and this time the claim is about a human ear, not a keyboard.

Cascabel Media, a streaming service, markets its AI dubbing feature as "indistinguishable from human dubbing in blind tests." Ruxandra Petrescu, a localization lead, is deciding whether to greenlight it for her catalog.

Mapped onto AUDIT: ask who paid for it is Cascabel's own PR team, who commissioned the blind test. Uncover the eval set is finding out the "blind test" used casual viewers, not professional linguists or voice-over reviewers who'd actually notice mismatched tone. Demand the version pin is realizing the claim only covers one language pair, Spanish-English, while the marketing page shows a world map implying all 22 supported languages. Isolate what's missing is the total absence of any mention of lip-sync accuracy or emotional tone preservation, just a single preference score. Test it yourself is Ruxandra running her own small panel of ten professional reviewers across three language pairs, landing on an "indistinguishable" rate near 45 percent, not the near-universal preference the marketing implies.

Hand sketched quadrant titled Sorting dubbing claims by verifiability. Axes how impressive sounding from modest to bold, how verifiable from unverifiable to checkable. Indistinguishable claim sits top left, bold and unverifiable. 22 language map sits near it. Sample reel dated sits bottom right, modest and checkable. Named test panel sits bottom right too, modest and checkable.
The boldest claims sit exactly where verifiability is lowest. That corner is where AUDIT earns its keep.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "name the eval set, the version, and the scope, or the number is an ad, not evidence," and stop.
Cost: there's no budget to run a full internal benchmark before every purchase decision. Say so honestly, and at minimum demand the vendor's own methodology in writing before signing anything.
The model gets better, for real: even if Pathforge's next version genuinely improves, that's still not a reason to trust an unscoped, undated headline number. A better model deserves a better-documented claim, not a free pass on the same vague one.

Where people run it wrong.
They treat a persuasive sales call as a substitute for a checkable number, when the two have nothing to do with each other.
They assume a vendor's own benchmark reflects their own real use case, when scope and language mix usually don't match at all.
They never actually test the claim themselves, so the purchase decision rests entirely on someone else's motivated number.

How to use it live. When asked what honest AI marketing looks like, ask yourself first: could I hand this exact number to a skeptical engineer and have it survive five follow-up questions? If not, it's an ad, dressed as a fact.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what does honest marketing copy for an AI feature look like"?
Tap to flip
ANSWER
AUDIT: ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Priyanka Reddy, engineering lead at Verrick Systems, evaluating Pathforge for a 40-person team.
3 · THE HABIT
What had Priyanka nearly stopped doing, that she normally does on every evaluation?
Tap to flip
ANSWER
Actually checking a vendor's headline number before trusting it. The 98% figure was persuasive enough that she'd almost skipped her usual skepticism.
4 · THE GAP
What's the real difference between "accuracy" and what Pathforge's number measured?
Tap to flip
ANSWER
The number measured acceptance rate, how often a suggestion got clicked, not correctness rate, how often it was still right a week later.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pathforge publishing only the single most flattering number, with no methodology, because a bare headline tested better in ads than an honest sentence.
6 · THE NUMBER
Fill in the blank: Priyanka's own two-week test on Verrick's real repo measured a correctness rate of ___ percent.
Tap to flip
ANSWER
61 percent, far from the marketed 98 percent, and it varied heavily by language.
7 · THE REPLAY
Same sales pitch, redesigned scrutiny. What changes?
Tap to flip
ANSWER
Priyanka reads the footnote closely before drafting any recommendation, and runs her own repo test before signing, instead of after a company-wide rollout.
8 · CROSS PRODUCT TRANSFER
Section 4 runs AUDIT again on a different product. Which one, and what's the equivalent of "test it yourself"?
Tap to flip
ANSWER
Cascabel Media's AI dubbing feature. There, it's Ruxandra's own panel of ten professional reviewers across three language pairs.

Check yourself Score: 0 / 0

Multiple choice
1. Why does this answer treat Pathforge's "98% accurate" claim as unverifiable, rather than simply calling it false?
  • A. Because Priyanka never actually tested the tool herself.
  • B. Because 98 percent is mathematically impossible for any AI tool.
  • C. Because the claim has no named eval set, model version, or scope, so it can't be checked against anything at all.
  • D. Because Pathforge's sales rep admitted the number was invented.
Show hint
Look at the "uncover the eval set" and "demand the version pin" stages.
Show answer
C. AUDIT's point isn't proving a number false. It's proving whether a number can be checked at all.
True or false
2. True or false: this answer says every sentence on a vendor's marketing page needs the same level of scrutiny.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The scrutiny belongs specifically on quantified performance claims, not screenshots, testimonials, or pricing.
Fill in the blank
3. Fill in the blank: the marketed 98 percent figure actually came from Pathforge's version ___, measured fourteen months before Priyanka's evaluation.
Show hint
Look at the line chart of Pathforge's benchmark score by version.
Show answer
v1.2. The version shipping today, v1.5, scores 89 percent on Pathforge's own benchmark, and 61 percent on Verrick's real repo.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Publishing only the single most flattering number with no methodology, since a bare headline tested better than a fuller sentence, while every competitor did the same.
Short answer, apply it yourself
5. Find an AI product's marketing claim you've seen recently (an app store listing, an ad, a homepage). Run it through AUDIT in one sentence: what's missing that would make it checkable?
Show hint
Look for a percentage or superlative with no eval set, version, or scope attached.
Show answer
Model answer: Most people land on "it never says what the number is measured against, or as of when," which is exactly the gap AUDIT is built to find.
Before you close the answer
Why this works
Tests whether you can tell a persuasive number apart from a checkable one, and whether you'd actually go verify a vendor's claim before betting a purchase decision on it.
Follow-up traps
"Isn't it unrealistic to expect every vendor to publish full methodology?" Response: the bar isn't a research paper, it's one sentence naming the eval set, the version, and what the metric measures, which costs a vendor almost nothing to add.

"Couldn't Priyanka's own 61 percent test just be measuring something different than Pathforge's benchmark?" Response: possibly, which is exactly why the fix isn't "trust my number instead," it's "demand enough detail from the vendor's number that the two can actually be compared."
If pressed
Priyanka's real test tracked accepted suggestions against three separate tests, whether it compiled, whether it passed unit tests, and whether a reviewer flagged it in code review a week later, since a single pass/fail check would have missed the slower kind of wrongness that only review catches.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more