What early signals tell you an AI project is being oversold internally?
Chatterglass is Vespergate's brand sentiment tool. It reads what people post about a brand online and tags each mention positive, negative, or mixed. Delyth Pettibone owns its roadmap. Bartram Colefax, Vespergate's VP of Growth, showed the same three sarcastic and mixed sentiment posts in every deck for four months straight. Nobody had asked, in that whole time, what the real number behind them was.
- Ask for the current eval number attached to the exact claim being made, every time.Why: a claim said with confidence needs no number. A claim that's actually true can always name one.
- Watch the gap between how confident the claim sounds and when it was last checked, not the wording itself.Why: one confident sentence sounds reasonable on its own. Only the widening gap over months shows the drift.
- Treat "we haven't measured that lately" as a finding, not a stall.Why: silence on a specific number is information. Waiting politely just lets the gap grow another month.
- Notice when the same demo gets reused across many audiences with no updated number next to it.Why: a capability that's actually improving earns a fresh number. A stalled one gets the same three examples instead.
- Separate the benchmark that's easy from the one the claim actually needs.Why: a real, high number on the easy population is not proof of anything on the hard one the claim is actually about.
- Leave alone any claim that already carries a dated number right next to it.Why: chasing proof from claims that already show their work wastes the goodwill you need for the ones that don't.
How to answer this, stage by stage
Nobody is grading whether you can spot a liar. They're grading whether you can name the actual gap oversell lives in, before it ever reaches a client.
Let's learn
What does it mean for a claim inside a company to quietly stop being true?
Vespergate's Chatterglass reads what people post about a brand, on social platforms and review sites, and tags each mention positive, negative, or mixed. A marketing team wakes up to one score instead of a thousand posts.
Before Chatterglass, a five person brand team at Hollowreed, a snack brand and one of Vespergate's clients, read maybe 200 of their loudest posts a week by hand, and missed the rest completely. Chatterglass reads all 40,000 posts Hollowreed gets in a month, in about ten minutes, tags every one, and rolls it into a single score. On plain, clear cut posts, no sarcasm, no mixed feelings, it gets it right 94 percent of the time. That part of the product is real and it works.
Sarcasm and mixed sentiment are the harder, newer part. Six months ago, Cyrek Marchbank's team built a second eval set: 900 real posts pulled straight off social platforms and review sites, sarcasm and mixed feelings left in, not cleaned out. At kickoff, Chatterglass scored 61 percent on that set. Nobody has scored it again since.
Bartram Colefax, Vespergate's VP of Growth, has three posts he shows in every deck: a sarcastic complaint about a hold time, a gushing review of a new flavor, and a mixed review that loves the packaging and hates that it arrived crushed. Chatterglass tags all three correctly, every time. So he kept using them, in every board update, every prospect call, every renewal review, for four months straight.
Here's the turn. A wrong tag on one sarcastic post was never the real problem. The real problem is what happens next: a VP tells a client the model catches it, a slide calls it a differentiator, a renewal gets sold on a capability nobody has actually re-measured since a single afternoon four months back.
What it costs at its worst: a client builds a whole campaign around a sentiment score that's quietly wrong on exactly the posts, sarcastic complaints, angry customers being sarcastic, that most needed catching. Chatterglass ends up worse than no tool at all, because a wrong score said with confidence is more dangerous than an honest "we don't know yet."
What I would leave alone: Chatterglass's plain sentiment tagging, positive, negative, neutral, on a post with no sarcasm in it. The 94 percent clean set number is honest and gets rechecked with every model release. Making Bartram attach a fresh date to that claim every single time would be a real cost for no real risk.
The lesson: a claim that keeps getting repeated without anyone re-checking it isn't proof it's true. It's proof nobody's checked lately, and the two get mistaken for each other more easily than you'd think.
Now here is the same thing as a story
The short version above is what you'd actually say in an interview. Read this one for the four months it took before anyone asked what the sarcasm number actually was.
Delyth Pettibone has owned product roadmaps for six years, three of them at Vespergate. She's good at the unglamorous part of the job: reading a deck before it goes out, catching the sentence that promises more than the data behind it. She caught two of those in her first year at Vespergate, quietly, before they ever reached a client.
Chatterglass launched with a real, working core: plain sentiment tagging, 94 percent accurate on ordinary posts, checked and rechecked with every model update. For its first year, that was the whole product, and every claim Bartram made about it held up, because someone always had a current number for it.
Then Cyrek's team added sarcasm and mixed sentiment tagging, the feature Vespergate's biggest prospects kept asking for. It worked, sometimes beautifully, on the right examples. Bartram found three posts that always landed in a demo, an angry complaint dressed as praise, a gushing review, a review that loved and hated the same box in one breath. He put them in the kickoff deck. The room loved them.
Delyth's first monthly business review with Bartram went fine. He clicked to the sarcasm slide, called it "a stretch goal we're exploring," and the room moved on. Fine, she thought. Early days.
Second review, six weeks later, the wording had firmed up: "early results on sarcasm are promising." She almost asked what promising meant in numbers. She didn't, because the deadline for a different launch was eating her week, and Bartram sounded certain.
Third review, ten weeks in: "Chatterglass catches sarcastic reviews other tools miss." No hedge left in it at all. She noticed the shift this time. She almost asked. Then a board member asked Bartram a follow up question about pricing instead, and the moment passed.
She wasn't hiding from the question. She just never had a reason, on any single one of those Tuesdays, to think the sentence she'd just heard was any different from the one before it. Each version sounded like a small step up from the last, and small steps up sound like progress, not drift.
Week sixteen, a new customer success analyst was building the renewal deck for Hollowreed, Vespergate's snack brand client, and needed a number for the sarcasm slide. She pinged Delyth: "What's the current accuracy on this, so I can put a real number next to the claim instead of just the demo screenshots?"
Delyth went looking. She asked Cyrek's team for the latest sarcasm eval score. There wasn't one newer than kickoff, four months back. Sixty one percent, the same number that had been sitting quietly under six months of steadily more confident decks, untouched.
I want to say the problem was that sarcasm is hard to catch. It is. That's not really the story though. Bartram never once made up a number. He just repeated the last confident sentence a little more confidently each time, because the room had nodded at it before and nobody had corrected it. Each version, on its own, sounded reasonable. The cumulative picture, four months later, had nothing behind it but three rehearsed posts and an eval score that had never moved.
The decision Delyth would take back sits further back than any of the four reviews. It sits in the kickoff meeting itself, where nobody wrote down that a capability claim needed a fresh number and a date attached before it went anywhere near a client. That rule had never been needed before, because Chatterglass's only claims used to be about the plain tagging, and that number got checked constantly. Nobody updated the rule the day a much harder, much newer claim started using the exact same confident voice.
Delyth pulled Cyrek's team in that week. She asked one question first, not "can we fix sarcasm fast," but "what's the real number, right now, and can you rerun it before Thursday." They reran the 900 post eval set. Sixty three percent. Barely moved in four months, because nobody had actually been improving it, they'd been talking about it instead.
She rewrote the renewal slide herself. Instead of "a core differentiator," it read: "sarcasm and mixed sentiment detection, 63 percent accurate on real posts today, our roadmap gets it to 80 by next quarter." Then she built the rule that should have existed from the start: every capability claim in any external deck carries its eval number and the date it was last checked, on that exact claim, next to it, not borrowed from an easier test.
Hollowreed renewed anyway. The honest number, and a real roadmap behind it, cost far less trust than a confident claim quietly falling apart in front of them would have.
What I'd tell myself, back at that first review with the stretch goal slide: a claim that sounds a little more confident than last month isn't progress by itself. It's just a sentence. Someone still has to go check whether the world agrees with it.
LEAD, for catching a claim before it outruns its evidence
Not a way to prove Bartram was dishonest. LEAD is what forces you to name the actual gap oversell lives in, and to catch it before a client is the one who finds it.
The recap, one line per letter: link oversell to whether the company's real decisions are still anchored to checked evidence, not to how good a demo feels in a room. The early signal is the gap between confident wording and a last verified date, because that gap widens for months before a client ever notices. Name the abuse plainly, a real number borrowed from an easier test and quietly reused for a harder claim, repeated by people who each assumed someone else had checked it. And the decision is what makes it real: a number and a date on every claim, and "not yet checked" treated as the finding it actually is.
Two things worth saying outright, since the real judgment sits here. Delyth considered a simpler fix first, just asking Bartram to stop using the same three demo posts. She rejected it, because the demo was never the actual problem, the same three posts with a current, honest number beside them would have been fine. Banning the demo without fixing the missing evidence would have just meant a different demo showed up next quarter with the same gap behind it. The AI specific failure worth naming by name is benchmark mismatch: a model's real accuracy on one population, clean, mostly clear cut posts, gets quoted as if it applies to a harder, different population, sarcastic and mixed sentiment posts, that it was never actually scored against. The guardrail that catches it is a segment specific eval number for every claim about a specific segment, never one aggregate accuracy figure standing in for all of them. And the trade off was real and accepted on purpose: checking a claim's number before every external deck costs Cyrek's team real hours that could otherwise go into improving the model, and Delyth took that cost on purpose, because a claim nobody can back with a number costs Vespergate far more the day a client checks it themselves.
And if you want to be sure it really works, try it somewhere else
Same four letters, a claims drafting tool instead of a sentiment score, and this time the harder population is a multi car pileup instead of a sarcastic tweet.
Penhollow runs Briefline for its adjusters. An adjuster uploads photos and notes from a car accident, and Briefline drafts the claim summary a human reviewer then checks and files. Horatio Oatlands owns its roadmap, and hit a close cousin of Delyth's exact problem four months into Briefline's rollout.
Briefline's demo, shown at every underwriting review since launch, was always a single vehicle fender bender: one car, one driver, one clean set of photos. It drafted those claims well, every time. Leadership decks started calling Briefline "ready for complex, multi vehicle claims" as a near term fact, three months before anyone had actually scored it on one.
The trigger wasn't a client this time. A state regulator, ahead of a routine audit, asked Penhollow's compliance team for Briefline's current accuracy on multi party claims specifically. Nobody had a number newer than four months old: 54 percent, measured once at kickoff, never rerun since. Horatio delayed the audit response by a week, reran the eval on 400 real multi vehicle claims, and told the regulator the honest current number instead of the implied one.
Mapped onto LEAD, the shape holds. The link is whether Penhollow's actual claims decisions stay anchored to checked evidence, not whether a demo impresses an underwriting review. The early signal is the same gap, confident wording outrunning its last verified date, visible in every monthly deck for months before the audit ever asked. The abuse Horatio found was the same mechanism too, a real number on an easy case quietly standing in for an unproven claim about a hard one. And his decision matched Delyth's exactly: a number and a date on every claim, and a compliance review that now asks for the segment specific score before any capability gets called ready.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: watch the gap between confident wording and a last verified date, and ask for the number attached to the exact claim.
Cost: no budget this quarter for a full eval refresh before every deck. Whoever owns the claim states the last real number and its date out loud, in the deck itself, an honest old number beats a confident unverified one.
The model got better, for real: say sarcasm accuracy genuinely climbs to 85 percent overnight. Keep the number and date rule anyway, because a stronger model that nobody rechecks starts the exact same drift one release later.
Where people run it wrong.
They treat a good demo as proof, instead of a demonstration of three cases that happen to work.
They let one aggregate accuracy number stand in for a claim about a specific, harder segment it was never actually scored against.
They wait for someone to volunteer a correction instead of asking for the current number the moment a claim's language firms up.
How to use it live. When an interviewer asks how you'd catch a project being oversold, buy yourself a second by asking out loud: "when was the last time anyone checked the exact thing this claim says, on real data." That question previews the whole answer.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a fresh eval number is expensive to get every time, so nobody can rerun it constantly?" Response: then tie the cadence to when the claim actually gets reused externally, a renewal call, a board deck, not to a fixed calendar. The real cost was never refreshing the number often, it was staking a client relationship on a claim nobody had paid to check even once since kickoff.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Managing stakeholder expectations and AI hype
- #1 Your CEO saw a demo on social media and wants that feature in six weeks. Structure your response.
- #2 How do you set expectations about AI capability without sounding like you are blocking?
- #3 Describe the difference between a demo and a product, using a concrete example.
- #4 Your board asks why competitors ship AI features faster. Prepare your answer.
- #5 Write the three sentences you would use to reset expectations after an overpromised launch date.
- #6 How do you handle a sales team that has already sold a capability you do not have?