ConceptIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #11

What early signals tell you an AI project is being oversold internally?

LEAD · why a reused demo outruns its own evidence, tested on a brand sentiment tool called Chatterglass

Chatterglass is Vespergate's brand sentiment tool. It reads what people post about a brand online and tags each mention positive, negative, or mixed. Delyth Pettibone owns its roadmap. Bartram Colefax, Vespergate's VP of Growth, showed the same three sarcastic and mixed sentiment posts in every deck for four months straight. Nobody had asked, in that whole time, what the real number behind them was.

The direct answer
Track the gap between how confidently a capability gets described and when it was actually last checked on real data, not the demo. Before a failure shows up, that gap shows up first: a claim moves from "exploring" to "a core feature" while the same few examples get reused for months with no new number beside them. When you see that gap, ask for the current number next to the exact claim being made, and treat "we haven't checked lately" as evidence something's wrong, not evidence nothing is.
Do this, in order
  1. Ask for the current eval number attached to the exact claim being made, every time.Why: a claim said with confidence needs no number. A claim that's actually true can always name one.
  2. Watch the gap between how confident the claim sounds and when it was last checked, not the wording itself.Why: one confident sentence sounds reasonable on its own. Only the widening gap over months shows the drift.
  3. Treat "we haven't measured that lately" as a finding, not a stall.Why: silence on a specific number is information. Waiting politely just lets the gap grow another month.
  4. Notice when the same demo gets reused across many audiences with no updated number next to it.Why: a capability that's actually improving earns a fresh number. A stalled one gets the same three examples instead.
  5. Separate the benchmark that's easy from the one the claim actually needs.Why: a real, high number on the easy population is not proof of anything on the hard one the claim is actually about.
  6. Leave alone any claim that already carries a dated number right next to it.Why: chasing proof from claims that already show their work wastes the goodwill you need for the ones that don't.

How to answer this, stage by stage

Nobody is grading whether you can spot a liar. They're grading whether you can name the actual gap oversell lives in, before it ever reaches a client.

1
Pin the question to one real product
Say it like this
"Let's ground this in one real case. Chatterglass is Vespergate's brand sentiment tool, and I'll answer using Delyth Pettibone, who owns its roadmap, not the idea of an oversold project in general."
Why this works
Naming a real product and a real owner stops "signs of oversell" from turning into vague office politics advice.
2
Name the four part method out loud
Say it like this
"I'll run this as LEAD. Link, what oversell actually damages if nobody catches it. Early signal, the concrete tells that show up before the damage does. Abuse, how those tells get created without anyone really lying. Decision, what I'd actually do the moment I see them."
Why this works
Two seconds of structure shows the interviewer a method, not a hunch about office politics.
3
Answer the literal question first, in one line
Say it like this
"Short version: watch the gap between how confidently a capability gets described and when it was last actually checked on real data. That gap is the tell, and it shows up weeks before anything visibly breaks."
Why this works
This gives the direct answer plainly, before any story, so the interviewer never has to dig for it.
4
Reframe what the question is really testing
Say it like this
"This isn't really asking whether I can spot a liar. It's asking whether I can tell a claim that's been checked recently from one that's just been said with more confidence than last time, because that's the actual shape oversell takes inside a real company."
Why this works
Shows the interviewer you understand the mechanism, drift, not deceit, which is what the question is actually probing.
5
Name the three tells, in the words you'd actually hear
Say it like this
"Three things I'd listen for. One, someone says 'the model handles X' in a status update with no hedge, when six weeks ago the same team said 'we're exploring X.' Two, the exact same demo gets shown to a new audience, a board, a prospect, a renewal call, and nobody's updated the number beside it. Three, I ask about a specific failure case and get 'still tuning it,' the same answer, for months, with no before and after number to back it up."
Why this works
This is the early signal step said out loud, specific enough that a listener could go spot it in their own company tomorrow.
6
Prove it with the real four months, numbers first
Say it like this
"Here's what actually happened at Vespergate. Cyrek Marchbank's team measured sarcasm and mixed sentiment accuracy at 61 percent, once, at kickoff. Over the next sixteen weeks, Bartram's deck language went from 'a stretch goal' to 'a core differentiator,' six confident claims total, while the eval got rerun zero times. When a client renewal deck needed a current number, there wasn't one newer than four months old."
Why this works
Two real numbers, sixteen weeks apart, do more work than any amount of talk about spotting hype.
7
Give the decision, name what you'd leave alone, then close
Say it like this
"So here's what I'd actually do. Every claim in a deck gets the eval number and the date it was last checked, right next to it, on that exact claim, not a borrowed number from an easier test. I wouldn't chase a claim that already carries its own dated number, that's not where the risk lives. Watch the gap, not the wording, and 'we haven't checked lately' is the answer, not a placeholder for one."
Why this works
Naming a place you wouldn't apply the scrutiny is what proves this is judgment, not paranoia.

Let's learn

What does it mean for a claim inside a company to quietly stop being true?

Vespergate's Chatterglass reads what people post about a brand, on social platforms and review sites, and tags each mention positive, negative, or mixed. A marketing team wakes up to one score instead of a thousand posts.

Knowledge spark: what is mixed sentiment? A single post that's good and bad at once, like a review that loves the flavor but hates that the box arrived crushed. Tag it wrong and a brand loses track of the exact kind of feedback it actually needs to act on.

Before Chatterglass, a five person brand team at Hollowreed, a snack brand and one of Vespergate's clients, read maybe 200 of their loudest posts a week by hand, and missed the rest completely. Chatterglass reads all 40,000 posts Hollowreed gets in a month, in about ten minutes, tags every one, and rolls it into a single score. On plain, clear cut posts, no sarcasm, no mixed feelings, it gets it right 94 percent of the time. That part of the product is real and it works.

Sarcasm and mixed sentiment are the harder, newer part. Six months ago, Cyrek Marchbank's team built a second eval set: 900 real posts pulled straight off social platforms and review sites, sarcasm and mixed feelings left in, not cleaned out. At kickoff, Chatterglass scored 61 percent on that set. Nobody has scored it again since.

Hand sketched comparison diagram titled What the sarcasm claim actually rested on. Left panel, a document icon labeled The demo, caption the same three sarcastic mentions, shown for sixteen weeks. Right panel, a scale icon labeled The eval set, caption 900 real messy mentions, last scored at kickoff.
One of these got shown every week. The other one got measured exactly once.

Bartram Colefax, Vespergate's VP of Growth, has three posts he shows in every deck: a sarcastic complaint about a hold time, a gushing review of a new flavor, and a mixed review that loves the packaging and hates that it arrived crushed. Chatterglass tags all three correctly, every time. So he kept using them, in every board update, every prospect call, every renewal review, for four months straight.

Hand sketched timeline titled Four monthly decks, one claim getting louder, with four milestones. Month 1, kickoff deck, caption sarcasm named a stretch goal. Month 2 update, caption early results look promising. Month 3 update, caption catches sarcasm, said as fact. Month 4, renewal deck, this milestone emphasized, caption a core differentiator now.
Four decks, four months. The number behind the claim never moved because nobody asked it to.
Confident capability claims made vs. eval refreshes run, by month
8 4 0 1 6, renewal deck 0, the whole time Month 1 Month 2 Month 3 Month 4
Confident claims made, no hedgeEval refreshes run on the sarcasm set
Claims climbed from one to six over four months. The eval behind them was never run again after kickoff. Nobody was watching this gap, which is exactly why it got to keep widening.
A confident sentence is not evidence. It's just a sentence that hasn't been checked yet.

Here's the turn. A wrong tag on one sarcastic post was never the real problem. The real problem is what happens next: a VP tells a client the model catches it, a slide calls it a differentiator, a renewal gets sold on a capability nobody has actually re-measured since a single afternoon four months back.

Chatterglass accuracy: the number quoted vs. the number that applied
100% 50% 0% 94% Overall accuracy, quoted every deck 61% Sarcasm accuracy, last verified
Number quotedNumber that applied to the claim
94 percent was real, measured on the easy, mostly clear cut posts. It was never the number behind the sarcasm claim. That number sat at 61 percent the whole time, and nobody said so out loud.
Hand sketched icon list titled Which one actually proves the claim, with three rows. Row one, document icon, the same three demo mentions shown again, in a warning amber color. Row two, person icon, confident wording one notch louder each deck, in a warning amber color. Row three, scale icon, a dated eval number sitting next to the claim, in a healthy teal color.
Two of these cost nothing to repeat. Only one of them can actually be checked.

What it costs at its worst: a client builds a whole campaign around a sentiment score that's quietly wrong on exactly the posts, sarcastic complaints, angry customers being sarcastic, that most needed catching. Chatterglass ends up worse than no tool at all, because a wrong score said with confidence is more dangerous than an honest "we don't know yet."

The choice I would take back Four months earlier, nobody wrote down that a capability claim needed a number and a date attached before it went in a deck. That was fine when Chatterglass only made claims about its plain sentiment tagging, which really was checked often, since that's most of the product. It stopped being fine the moment a much newer, much harder capability, sarcasm and mixed sentiment, started getting talked about in exactly the same confident voice.
Hand sketched labeled parts diagram titled What Delyth's claim ledger holds. A central document icon labeled Claim ledger, with four labeled callouts: the exact claim made, the eval number behind it, the date it was checked, who can see it above her.
Four parts, every time a claim goes in a deck. Leave one out and it goes back to being a sentence nobody can check.

What I would leave alone: Chatterglass's plain sentiment tagging, positive, negative, neutral, on a post with no sarcasm in it. The 94 percent clean set number is honest and gets rechecked with every model release. Making Bartram attach a fresh date to that claim every single time would be a real cost for no real risk.

The lesson: a claim that keeps getting repeated without anyone re-checking it isn't proof it's true. It's proof nobody's checked lately, and the two get mistaken for each other more easily than you'd think.

Now here is the same thing as a story

The short version above is what you'd actually say in an interview. Read this one for the four months it took before anyone asked what the sarcasm number actually was.

Delyth Pettibone has owned product roadmaps for six years, three of them at Vespergate. She's good at the unglamorous part of the job: reading a deck before it goes out, catching the sentence that promises more than the data behind it. She caught two of those in her first year at Vespergate, quietly, before they ever reached a client.

Chatterglass launched with a real, working core: plain sentiment tagging, 94 percent accurate on ordinary posts, checked and rechecked with every model update. For its first year, that was the whole product, and every claim Bartram made about it held up, because someone always had a current number for it.

Then Cyrek's team added sarcasm and mixed sentiment tagging, the feature Vespergate's biggest prospects kept asking for. It worked, sometimes beautifully, on the right examples. Bartram found three posts that always landed in a demo, an angry complaint dressed as praise, a gushing review, a review that loved and hated the same box in one breath. He put them in the kickoff deck. The room loved them.

Hand sketched timeline titled Delyth's four business reviews with Bartram, with four milestones. Review one, caption same three mentions, she nods. Review two, caption wording firms up, she nods again. Review three, caption a core claim now, she almost asks. QBR prep, week 16, this milestone emphasized, caption the analyst asks one real question.
Nobody decided to stop asking. A climbing number of nods just made it feel like there was nothing left to ask.

Delyth's first monthly business review with Bartram went fine. He clicked to the sarcasm slide, called it "a stretch goal we're exploring," and the room moved on. Fine, she thought. Early days.

Second review, six weeks later, the wording had firmed up: "early results on sarcasm are promising." She almost asked what promising meant in numbers. She didn't, because the deadline for a different launch was eating her week, and Bartram sounded certain.

Third review, ten weeks in: "Chatterglass catches sarcastic reviews other tools miss." No hedge left in it at all. She noticed the shift this time. She almost asked. Then a board member asked Bartram a follow up question about pricing instead, and the moment passed.

She wasn't hiding from the question. She just never had a reason, on any single one of those Tuesdays, to think the sentence she'd just heard was any different from the one before it. Each version sounded like a small step up from the last, and small steps up sound like progress, not drift.

Week sixteen, a new customer success analyst was building the renewal deck for Hollowreed, Vespergate's snack brand client, and needed a number for the sarcasm slide. She pinged Delyth: "What's the current accuracy on this, so I can put a real number next to the claim instead of just the demo screenshots?"

Delyth went looking. She asked Cyrek's team for the latest sarcasm eval score. There wasn't one newer than kickoff, four months back. Sixty one percent, the same number that had been sitting quietly under six months of steadily more confident decks, untouched.

We did not lose four months to a hard problem. We lost four months to a number nobody thought to ask for again.

I want to say the problem was that sarcasm is hard to catch. It is. That's not really the story though. Bartram never once made up a number. He just repeated the last confident sentence a little more confidently each time, because the room had nodded at it before and nobody had corrected it. Each version, on its own, sounded reasonable. The cumulative picture, four months later, had nothing behind it but three rehearsed posts and an eval score that had never moved.

The decision Delyth would take back sits further back than any of the four reviews. It sits in the kickoff meeting itself, where nobody wrote down that a capability claim needed a fresh number and a date attached before it went anywhere near a client. That rule had never been needed before, because Chatterglass's only claims used to be about the plain tagging, and that number got checked constantly. Nobody updated the rule the day a much harder, much newer claim started using the exact same confident voice.

Delyth pulled Cyrek's team in that week. She asked one question first, not "can we fix sarcasm fast," but "what's the real number, right now, and can you rerun it before Thursday." They reran the 900 post eval set. Sixty three percent. Barely moved in four months, because nobody had actually been improving it, they'd been talking about it instead.

She rewrote the renewal slide herself. Instead of "a core differentiator," it read: "sarcasm and mixed sentiment detection, 63 percent accurate on real posts today, our roadmap gets it to 80 by next quarter." Then she built the rule that should have existed from the start: every capability claim in any external deck carries its eval number and the date it was last checked, on that exact claim, next to it, not borrowed from an easier test.

Hollowreed renewed anyway. The honest number, and a real roadmap behind it, cost far less trust than a confident claim quietly falling apart in front of them would have.

What I'd tell myself, back at that first review with the stretch goal slide: a claim that sounds a little more confident than last month isn't progress by itself. It's just a sentence. Someone still has to go check whether the world agrees with it.

LEAD, for catching a claim before it outruns its evidence

Not a way to prove Bartram was dishonest. LEAD is what forces you to name the actual gap oversell lives in, and to catch it before a client is the one who finds it.

LLink. The business outcome that actually matters.
Not whether a demo lands well in a room. What actually matters is whether the company's decisions, roadmap bets, client commitments, board updates, are still anchored to real, current eval evidence. The moment that anchor slips, the outcome that eventually breaks is a client catching a claim that was never actually true, or a team burning a quarter chasing a capability that doesn't exist yet.
Chatterglass's real risk was never one wrong tag on one post. It was Hollowreed building a campaign on a sentiment score that was quietly wrong on the exact posts that mattered most.
EEarly signal. The thing that moves before the outcome does.
Three concrete tells, all visible weeks before anything breaks. A capability gets described in confident, near term language when it was hedged language a few weeks earlier. The same demo gets shown to a new audience with no updated number beside it. A specific failure question gets answered with "still tuning it," the same words, for months, with no before and after evidence.
Bartram's deck went from "a stretch goal" to "a core differentiator" over four reviews, six confident claims total, while the eval behind it got rerun exactly zero times.
AAbuse. How oversell happens without anyone really lying.
Not one dramatic lie. A slow drift where each individual claim felt reasonable in the moment. The mechanism at Vespergate was specific: 94 percent was a real, honest number, just measured on the easy, clean set. It got carried into a claim about a much harder population, sarcasm and mixed sentiment, that it was never validated for, and each speaker who repeated it assumed the last person had checked.
"94 percent accurate" and "catches sarcasm other tools miss" sound like the same sentence. Only one of those numbers was ever measured on sarcasm at all.
DDecision. What you'd actually do differently.
Attach the eval number and its date to every capability claim, on that exact claim, in every external facing deck. Treat "we haven't checked lately" as a finding worth chasing down immediately, not a placeholder to wait politely on. Tie the re-check cadence to when a claim actually gets reused externally, a renewal call, a board update, not to a fixed calendar nobody enforces.
One rerun, one honest number, one rewritten slide turned a renewal that could have imploded into a client who renewed on a roadmap they actually trusted.
Hand sketched comparison diagram titled Which clock rings first. Left panel, a gauge icon labeled A client's own trust, caption only breaks once they test the claim themselves, months away. Right panel, a gauge icon labeled The claim evidence gap, caption visible in every deck, the same week it opens.
One of these is ready to read the same week a deck goes out. The other one waits for a client to find out the hard way.

The recap, one line per letter: link oversell to whether the company's real decisions are still anchored to checked evidence, not to how good a demo feels in a room. The early signal is the gap between confident wording and a last verified date, because that gap widens for months before a client ever notices. Name the abuse plainly, a real number borrowed from an easier test and quietly reused for a harder claim, repeated by people who each assumed someone else had checked it. And the decision is what makes it real: a number and a date on every claim, and "not yet checked" treated as the finding it actually is.

Two things worth saying outright, since the real judgment sits here. Delyth considered a simpler fix first, just asking Bartram to stop using the same three demo posts. She rejected it, because the demo was never the actual problem, the same three posts with a current, honest number beside them would have been fine. Banning the demo without fixing the missing evidence would have just meant a different demo showed up next quarter with the same gap behind it. The AI specific failure worth naming by name is benchmark mismatch: a model's real accuracy on one population, clean, mostly clear cut posts, gets quoted as if it applies to a harder, different population, sarcastic and mixed sentiment posts, that it was never actually scored against. The guardrail that catches it is a segment specific eval number for every claim about a specific segment, never one aggregate accuracy figure standing in for all of them. And the trade off was real and accepted on purpose: checking a claim's number before every external deck costs Cyrek's team real hours that could otherwise go into improving the model, and Delyth took that cost on purpose, because a claim nobody can back with a number costs Vespergate far more the day a client checks it themselves.

And if you want to be sure it really works, try it somewhere else

Same four letters, a claims drafting tool instead of a sentiment score, and this time the harder population is a multi car pileup instead of a sarcastic tweet.

Penhollow runs Briefline for its adjusters. An adjuster uploads photos and notes from a car accident, and Briefline drafts the claim summary a human reviewer then checks and files. Horatio Oatlands owns its roadmap, and hit a close cousin of Delyth's exact problem four months into Briefline's rollout.

Hand sketched flow diagram titled The same drift, a claims drafting tool this time, with five connected steps: Simple demo, Shown monthly, Claimed ready, Regulator asks, this step emphasized, Only 54 percent, stale.
A different claim, a different eval gap. The same shape of silence let it sit unchecked for four months.

Briefline's demo, shown at every underwriting review since launch, was always a single vehicle fender bender: one car, one driver, one clean set of photos. It drafted those claims well, every time. Leadership decks started calling Briefline "ready for complex, multi vehicle claims" as a near term fact, three months before anyone had actually scored it on one.

The decision Horatio would take back Penhollow's rollout plan never separated the claim "handles single vehicle claims" from the much bigger claim "handles complex, multi vehicle claims." Both got filed under one accuracy number, so a real result on the easy case quietly stood in for a claim about the hard one nobody had tested yet.

The trigger wasn't a client this time. A state regulator, ahead of a routine audit, asked Penhollow's compliance team for Briefline's current accuracy on multi party claims specifically. Nobody had a number newer than four months old: 54 percent, measured once at kickoff, never rerun since. Horatio delayed the audit response by a week, reran the eval on 400 real multi vehicle claims, and told the regulator the honest current number instead of the implied one.

Mapped onto LEAD, the shape holds. The link is whether Penhollow's actual claims decisions stay anchored to checked evidence, not whether a demo impresses an underwriting review. The early signal is the same gap, confident wording outrunning its last verified date, visible in every monthly deck for months before the audit ever asked. The abuse Horatio found was the same mechanism too, a real number on an easy case quietly standing in for an unproven claim about a hard one. And his decision matched Delyth's exactly: a number and a date on every claim, and a compliance review that now asks for the segment specific score before any capability gets called ready.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: watch the gap between confident wording and a last verified date, and ask for the number attached to the exact claim.
Cost: no budget this quarter for a full eval refresh before every deck. Whoever owns the claim states the last real number and its date out loud, in the deck itself, an honest old number beats a confident unverified one.
The model got better, for real: say sarcasm accuracy genuinely climbs to 85 percent overnight. Keep the number and date rule anyway, because a stronger model that nobody rechecks starts the exact same drift one release later.

Where people run it wrong.
They treat a good demo as proof, instead of a demonstration of three cases that happen to work.
They let one aggregate accuracy number stand in for a claim about a specific, harder segment it was never actually scored against.
They wait for someone to volunteer a correction instead of asking for the current number the moment a claim's language firms up.

How to use it live. When an interviewer asks how you'd catch a project being oversold, buy yourself a second by asking out loud: "when was the last time anyone checked the exact thing this claim says, on real data." That question previews the whole answer.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks how you'd catch a project being oversold internally?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Built for metric questions, it ties "oversold" to something observable, the gap between claim confidence and a last verified eval date, instead of a feeling.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Delyth Pettibone, who owns Chatterglass's roadmap at Vespergate, and Bartram Colefax, the VP of Growth who reused the same three demo posts for four months.
3 · THE LINK
What should "the project is being oversold" actually be measured against?
Tap to flip
ANSWER
Whether the company's real decisions stay anchored to checked eval evidence, not whether a demo lands well or a claim sounds confident in the room.
4 · THE EARLY SIGNAL
What are the three concrete tells that show up before a project's oversell becomes visible?
Tap to flip
ANSWER
Aspirational future tense claims said as near term fact, the same demo reused across audiences with no updated number, and "still tuning it" repeated for months with no before and after evidence.
5 · THE OLD DECISION
What decision would Vespergate take back?
Tap to flip
ANSWER
Never requiring a capability claim to carry its eval number and check date before it went in a deck, a rule that was never needed while every claim was about the well checked plain tagging feature.
6 · THE NUMBER
Fill in the blank: the overall clean set accuracy quoted in every deck was ___ percent, but the real sarcasm and mixed sentiment accuracy, last verified at kickoff, was ___ percent.
Tap to flip
ANSWER
94 percent quoted; 61 percent actually applied. A real number, borrowed from the wrong population and never corrected.
7 · THE REPLAY
Same renewal deadline pressure, but the claim gets a number and a date attached before it ships. What changes?
Tap to flip
ANSWER
Cyrek's team reruns the eval, landing at 63 percent. The slide states the real number and a roadmap to 80 percent instead of calling it a core differentiator. Hollowreed renews on the honest version.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the parallel oversell there?
Tap to flip
ANSWER
Briefline, a claims drafting tool at the insurer Penhollow, run by Horatio Oatlands. There a single vehicle demo quietly stood in for an unproven "handles complex multi vehicle claims" claim, caught when a regulator asked for a number nobody had rechecked in four months.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Vespergate's sarcasm claim keep sounding more confident every month, even though nobody actually lied?
  • A. Bartram invented a fake accuracy number for each deck.
  • B. Each claim repeated the last one a little more confidently, and nobody rechecked the real eval number behind it for four months.
  • C. Cyrek's team kept retraining the model every week without telling anyone.
  • D. Vespergate didn't own an eval set for sarcasm at all.
Show hint
Check the "A, abuse" step in the framework recap.
Show answer
B. This is the slow drift LEAD's abuse step names: each individual claim felt reasonable, and the gap between confidence and evidence just kept widening because nobody was watching it.
True or false
2. True or false: the 94 percent accuracy number Bartram quoted was fabricated.
  • True
  • False
Show hint
Look at the bar chart comparing the quoted number to the number that applied.
Show answer
False. 94 percent was a real, honestly measured number, just on the easy clean set of posts. The problem was that it got carried into a claim about sarcasm, a much harder population it was never scored against.
Fill in the blank
3. When the customer success analyst finally asked for a current number, the sarcasm eval hadn't been rerun since kickoff, ___ months earlier, and it was still sitting at ___ percent.
Show hint
Look at stage 6 of the walkthrough, or the "Let's learn" section's numbers.
Show answer
Four months earlier; 61 percent. The same number that had been sitting quietly under six months of increasingly confident decks.
Short answer, name the reversal
4. What old decision would Delyth take back, and why did it make sense the first time nobody made that rule?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Never requiring a capability claim to carry its eval number and check date before going in a deck. It made sense while every Chatterglass claim was about the plain tagging feature, which really was checked constantly. It stopped making sense the moment a newer, harder claim, sarcasm, started using the same confident voice.
Short answer, apply it yourself
5. Think of a claim you've heard repeated inside a company you've worked for or used a product from. What question would have told you whether it was still true, or just still being said?
Show hint
Think about asking for a specific, dated number attached to the exact claim, not a general "is this still accurate" question.
Show answer
Model answer: If a support tool kept getting called "fully automated" in every update, asking "what percent of tickets closed with zero human touch last week, and when was that last measured" would separate a real, current claim from one that had just been repeated long enough to sound true.
Short answer, work the number
6. If Chatterglass's clean set accuracy had been 99 percent instead of 94, would that have made the sarcasm claim any safer to state as fact? Why or why not?
Show hint
Ask what population that higher number would actually have been measured on.
Show answer
No, not safer at all. A higher clean set number still says nothing about sarcasm and mixed sentiment posts, since that's a different, harder population the claim was never actually scored against. The gap between the two numbers is the whole problem, not the size of either one alone.
Before you close the answer
Why this works
Tests whether you can tell an evidence gap from a trust problem. Most candidates answer with vague advice about being skeptical of leadership, and never name the actual, checkable thing that moves before a project visibly fails: the distance between how a claim is worded and when it was last verified.
Follow-up traps
"Isn't this just about being more skeptical of your VP of Growth?" Response: no, it's about missing evidence, not a bad person. Bartram never lied, each version of the claim felt reasonable given the one before it, which is exactly why suspicion of him wouldn't have caught it. Only a number attached to the exact claim would.

"What if a fresh eval number is expensive to get every time, so nobody can rerun it constantly?" Response: then tie the cadence to when the claim actually gets reused externally, a renewal call, a board deck, not to a fixed calendar. The real cost was never refreshing the number often, it was staking a client relationship on a claim nobody had paid to check even once since kickoff.
If pressed
The 900 post eval set deliberately oversamples sarcastic and mixed sentiment posts, since raw platform data is mostly plain sentiment and a random sample wouldn't hold enough hard cases to grade the claim being made. The final score gets reweighted back to the real mix of post types afterward, so the oversampling doesn't quietly fake a lower or higher number than reality.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more