ConceptIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #11
How do you evaluate a vendor's roadmap credibility?
TRACE the line that read "next quarter" in eight straight decks, unchanged
Thermorun Services keeps rooftop HVAC units running for commercial buildings across its region. Naledi Ferreira, an ops analyst there, was asked to check Airloop Predictive's roadmap credibility before a multi-year renewal. Airloop's shipped product already predicts compressor failure well. Its promised next feature had not moved in two years.
The direct answer
Don't read a roadmap item, test it. Pull the vendor's last several decks and see if the same promise repeats unchanged, split their track record by how novel each capability actually is, and ask for a rough prototype on your own real data before trusting a date. A vendor that can show a number, even a bad one, on your own data has real work behind a promise. A vendor that goes quiet when asked doesn't.
Do this, in order
Ask for a working prototype on your own data before trusting any roadmap date.Why: a real number, even a rough one, is the one thing a wishcast promise can't fake.
Line up several old roadmap decks and check if the same promise repeats unchanged.Why: a capability that's been "next quarter" for two years isn't close, no matter what the current deck says.
Split the vendor's track record by how new each capability actually is.Why: a vendor that ships integrations on time can still be unreliable specifically on new model capabilities.
Don't count a relabeled old metric as evidence of new progress.Why: "enabled for 200 buildings" can describe the feature that already shipped, not the one still promised.
Trust a roadmap item that names its own real blocker over one that just repeats a date.Why: a vendor that can say exactly what's hard is doing real work; one that can't is guessing out loud.
How to answer this, stage by stage
Nobody is scoring whether you can say "be skeptical of roadmaps." They're scoring whether you can name the one check that actually tells a stalled promise from a hard one.
Stage 1
Scope it to one real roadmap item
Say it like this
"I'll answer this for Thermorun Services, checking Airloop Predictive's promised auto-diagnosis feature, not vendor roadmaps in general."
Why this works
Keeps "roadmap credibility" from turning into a generic due-diligence lecture.
Stage 2
Say your structure out loud
Say it like this
"I'll use TRACE. Timeline the promise, recut by category, assume nothing about the vendor's own framing, name three cause candidates, then run the one evidence test."
Why this works
TRACE is built for diagnosis, and a stalled promise is exactly that: something you have to diagnose, not just react to.
Stage 3
Reframe what the question is really testing
Say it like this
"This isn't really asking if I trust the vendor's sales team. It's asking whether I can tell a genuinely hard problem from a promise nobody ever scoped."
Why this works
Separates a real answer from a vague "vendors always overpromise" line.
Stage 4
Give the one check that matters most
Say it like this
"Ask Airloop to run auto-diagnosis, even roughly, on our own forty confirmed compressor failures. A real number, even a bad one, tells us this has real engineering behind it."
Why this works
This is the direct answer, said the way you'd actually say it out loud.
Stage 5
Prove it with the compressed story
Say it like this
"The same 'auto-diagnosis, next quarter' line sat in eight straight decks. When we finally asked for a rough prototype, we got three weeks of silence."
Why this works
Turns "be careful with roadmaps" into one specific, checkable pattern.
Stage 6
Say what you'd still leave alone
Say it like this
"Airloop's data-pipeline and reporting roadmap items ship on time about nine times in ten. I wouldn't apply this same suspicion there."
Why this works
Shows the caution is specific to novel-capability promises, not the whole vendor.
Stage 7
Close on the one line
Say it like this
"Don't read the roadmap, test it. A prototype on your own data separates a hard promise from one nobody ever really started."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Let's learn
Say a vendor sells AI that predicts when a rooftop compressor will fail, and promises a second capability, telling you exactly which part will fail, still "coming next quarter."
Before Airloop, Thermorun's technicians walked each rooftop once a season, listening to the compressor and logging a condition score by hand, about two hours per unit. That caught a real problem about half the time, roughly three months before it actually failed.
This is the job Airloop's shipped product already improved on. The roadmap question is about the part it hasn't built yet.
Airloop's shipped product predicts compressor failure thirty to forty five days out, correctly, seventy eight percent of the time, a real, working improvement Thermorun had verified for two years. Its roadmap promised something bigger: auto-diagnosis, telling a technician exactly which part inside the compressor would fail, not just that something would.
The real question was never whether Airloop's shipped product worked. It was whether the thing they kept promising next had any real engineering behind it at all.
On-time delivery rate, by roadmap item category
Auto-diagnosis lives in the worst-performing bucket. That alone doesn't prove it's fake, but it earns a closer look.
At its worst, signing a multi-year renewal on the strength of a feature that never ships doesn't just waste a budget line. It risks Thermorun planning its own staffing and pricing around a capability that was never technically real, and finding out only after the contract locks it in.
The choice I would take back
Thermorun's old vendor-review process only ever looked at the current roadmap deck, not the history of decks before it. That was fine when Airloop was a new, small vendor with no track record yet. It stopped being fine two years and eight decks into the relationship, once there was a real pattern to check against.
What I would leave alone: Airloop's data-pipeline and reporting roadmap items don't need this scrutiny. They ship on time about nine times in ten, and treating every promise from this vendor with equal suspicion would waste real review time on the parts that were never the problem.
The lesson: a roadmap item earns your trust by surviving a test, not by sounding confident on a slide. The vendors worth trusting are the ones willing to run that test.
Now here is the same thing as a story
The short version above is what you'd say defending this call to Thermorun's leadership. Read this one for how the pattern actually got noticed.
Naledi Ferreira had reviewed vendor contracts at Thermorun for four years, and she was the one who always asked for last year's numbers before trusting this year's pitch.
Eight quarters of the same sentence. Nobody had lined the decks up side by side before.
Airloop's account team was easy to work with, quick with data-pipeline fixes, quick with reporting tweaks. Its quarterly roadmap deck always looked the same too: a slide near the bottom reading "auto-diagnosis of failure type: coming next quarter," next to a bar showing the feature at "76% complete."
Knowledge spark: what's the difference between research risk and wishcasting?
Research risk means a real team is working on a genuinely hard, unsolved problem, and progress is slow because the problem is slow. Wishcasting means a feature was written onto a roadmap to help a sale close, with no real technical work ever scoped behind it. From the outside, both look like "still not shipped." Only a real test tells them apart.
Thermorun's finance team pulled the last three years of contract renewal decks for an unrelated budget review, and someone noticed the auto-diagnosis slide read almost word for word the same in all of them, "76% complete" in one deck, "72%" in another, never moving toward a hundred.
Airloop wasn't an unreliable vendor. It was a reliable vendor with one specific kind of promise it never actually started.
Naledi had considered simply trusting Airloop's "enabled for 200 buildings" claim as proof the new feature was close. She checked first, and found that number described the already-shipped failure-prediction product, not auto-diagnosis at all, relabeled to sound like fresh progress on the thing still missing.
All three suspects looked equally plausible from the outside. Only a real test could tell them apart.
She asked Airloop for one thing: run a rough version of auto-diagnosis, even an early prototype, against Thermorun's own last forty confirmed compressor failures, where the real failure type was already known and could be checked against the model's guess.
Three weeks passed. No number came back. No blocker was named either, just a note that the team was "still finalizing the approach."
Airloop landed on the branch Naledi least wanted to see, and it was the one that actually explained two years of an unmoving slide.
That silence was the answer. A team facing genuine research risk can usually show a rough, ugly number and explain exactly where it breaks. A team blocked on engineering can usually name the missing dataset. A team with neither, after two years and eight quarters, most likely never scoped the feature as real work in the first place.
Thermorun renewed its contract for the working failure-prediction product, at its proven seventy eight percent accuracy, and struck the auto-diagnosis promise from the contract entirely, telling Airloop to bring it back only once they could run it on Thermorun's own data first.
TRACE, in five movesNot a lecture on trusting vendors less. TRACE is what separates a hard, real promise from one that was never actually started.
T
Timeline. When the promise first appeared, and what's shipped since.
Auto-diagnosis first appeared eight quarters ago as "next quarter" and hasn't moved since.
A promise's age is the first honest fact about it, before anyone's opinion gets involved.
R
Recut. Slice the vendor's track record by category.
Data pipeline and reporting items ship on time 85 to 90 percent of the time. New ML capability items ship on time only 20 percent of the time.
An average reliability score hides that the risk is concentrated in one specific bucket.
A
Assume nothing. Check the vendor's own framing first.
"Enabled for 200 buildings" described the already-shipped product, not the promised one, once actually checked.
A relabeled old metric can look exactly like new progress on a slide.
C
Cause candidates. Three named hypotheses, not a list of everything possible.
Genuine research risk, real engineering risk needing more data, or marketing wishcasting with no technical work behind it.
This is the hardest step: naming exactly the theories that would explain the same stalled slide.
E
Evidence test. The one check that separates the top candidates.
Ask for a rough prototype on Thermorun's own forty confirmed failures. Silence, with no blocker named, pointed straight at wishcasting.
This is the strongest move in the whole method: one request that a real team can answer and an empty one can't.
Airloop's auto-diagnosis promise had none of these four. That absence was the whole finding.
The recap, one line per letter: timeline is checking how long a promise has actually existed, recut is splitting the track record by how new each capability is, assume nothing is checking whether a claimed metric is really about the new feature, cause candidates is naming research risk, engineering risk, and wishcasting as the three real options, and evidence test is the one request, a prototype on your own data, that tells them apart.
And if you want to be sure it really works, try it somewhere elseSame five moves, a restaurant equipment supplier instead of an HVAC service company. A different evidence test breaks the second roadmap open.
Skilletbrook Restaurant Supply services deep fryers for a chain of commercial kitchens. Its vendor, FryWise, sells AI that tracks fryer-oil degradation and schedules replacement, already shipped and working well. Its roadmap promised a second feature, auto-recommending replacement timing across different fryer brands with different oil chemistries, "next quarter," for six straight quarterly decks. Mapped onto TRACE: timeline is finding the cross-brand promise first appeared eighteen months ago with no movement since. Recut is noticing FryWise's single-brand features ship on time about eighty percent of the time, while cross-brand promises have a zero percent on-time rate across three attempts. Assume nothing is checking that FryWise's "500 kitchens using our recommendations" claim describes the single-brand feature already live, not the cross-brand one still promised. Cause candidates is the same three: real research risk in modeling unfamiliar oil chemistries, engineering risk needing more cross-brand data, or wishcasting.
Skilletbrook ran the same test Thermorun did, and got the same silence back.
The evidence test here is different from Airloop's: Skilletbrook asked FryWise to run the cross-brand model, even roughly, on oil-change logs from just two fryer brands instead of all of them, a smaller ask than Thermorun's own request. FryWise couldn't produce even that narrower result, which pointed to the same conclusion: not a hard research problem yet being chipped away at, but a promise that had never been scoped as real engineering work.
Airloop's own reported completion percentage for auto-diagnosis, across eight quarterly decks
A percentage that wanders between 65 and 76 for two years isn't tracking real progress. It's a guess dressed up as a metric.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask for a prototype on your own data before trusting any roadmap date," and stop.
Cost: there's no time to line up years of old decks before a renewal decision is due. Say so honestly, and at minimum ask the one evidence-test question live in the meeting instead of skipping it.
The model gets better, for real: if Airloop's next deck finally shows a real number on Thermorun's own data, that's the credibility signal actually arriving, and it deserves real weight, not more suspicion out of habit.
Where people run it wrong.
They read the current roadmap deck in isolation, never checking it against the same promise from a year ago.
They accept a vendor's own usage numbers as proof of progress on a specific unshipped feature, without checking what those numbers actually describe.
They never ask for a prototype on their own data, so a wishcast promise and a hard, real one look identical from the outside.
How to use it live. The moment a vendor's roadmap includes something ambitious that's slipped before, ask: can you show me a rough version of this running on our own real data today? What you get back, a number, a named blocker, or silence, is the answer.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a diagnosis-shaped question like judging roadmap credibility?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Built for diagnosing what's really behind a stalled claim.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Naledi Ferreira, an ops analyst at Thermorun Services, who always checks a vendor's history before trusting this year's pitch.
3 · THE PATTERN
What pattern gave the stalled promise away?
Tap to flip
ANSWER
The exact same "auto-diagnosis, next quarter" line, with a completion percentage that wandered between 65 and 76, across eight straight decks.
4 · THE EVIDENCE TEST
What single request separated a hard promise from a fake one?
Tap to flip
ANSWER
Asking the vendor to run a rough prototype on the buyer's own real, confirmed data. A number or a named blocker means real work. Silence means neither.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Reviewing only the current roadmap deck each year, instead of lining it up against older decks to check if a promise was actually moving.
6 · THE NUMBER
Fill in the blank: new ML capability roadmap items at Airloop shipped on time only ___ percent of the time, versus 85 to 90 percent for other categories.
Tap to flip
ANSWER
20 percent, concentrating nearly all of Airloop's real roadmap risk in one specific bucket of claims.
7 · THE REPLAY
Same renewal decision, but Naledi runs the evidence test at the start instead of the ninth quarter. What changes?
Tap to flip
ANSWER
Thermorun catches the wishcast promise in one quarter instead of two years, and never plans a budget or staffing decision around a feature that was never really being built.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different company. Which one, and what's different about its evidence test?
Tap to flip
ANSWER
Skilletbrook Restaurant Supply's FryWise. Its evidence test asked for a smaller, two-brand prototype rather than a full cross-brand run, and still got only silence back.
Check yourself Score: 0 / 0
Multiple choice
1. What was the single strongest signal that Airloop's auto-diagnosis feature was wishcasting, not real engineering delay?
A. Airloop's shipped failure-prediction product wasn't perfectly accurate.
B. The roadmap deck used a purple color scheme.
C. The vendor produced neither a number nor a named blocker when asked for a prototype on real data.
D. Thermorun's own technicians preferred manual inspection.
Show hint
Look at the E step and the decision tree.
Show answer
C. A team with real research or engineering progress can usually show a rough number or name a specific blocker. Airloop showed neither.
Fill in the blank
2. Fill in the blank: Airloop's reported completion percentage for auto-diagnosis wandered between 65 and ___ percent across eight quarterly decks, never approaching 100.
Show hint
Look at the line chart in Section 4.
Show answer
76 percent. A number that moves in place for two years is not tracking real progress toward completion.
True or false
3. True or false: Airloop's shipped, already-proven failure-prediction product is unreliable in the same way its roadmap promises are.
True
False
Show hint
Look at "what I would leave alone" and the recut step.
Show answer
False. The shipped product is verified and reliable. The risk is specific to the promised, unshipped, novel capability.
Short answer, where it wouldn't matter
4. Name a category of Airloop's roadmap where this level of suspicion genuinely doesn't apply.
Show hint
Look at "what I would leave alone" and the bar chart.
Show answer
Model answer: Data pipeline and UI or reporting roadmap items, which ship on time 85 to 90 percent of the time.
Short answer, apply it yourself
5. Think of a promised feature you've seen slip repeatedly, from any vendor or team. What single test would have told you if it was real progress or not?
Show hint
Ask whether anyone ever requested a rough, working version on real data instead of trusting the next date on a slide.
Show answer
Model answer: Usually asking for even a rough prototype run on real data, which a genuinely in-progress effort can produce and a wishcast one can't.
Short answer, name the reversal
6. What decision does the Skilletbrook version of this story take back, and how is its evidence test different from Thermorun's?
Show hint
Look at Section 4's evidence test description.
Show answer
Model answer: Skilletbrook asked for a smaller, two-brand version instead of the full cross-brand model, a lower bar than Thermorun's test, and FryWise still couldn't meet it.
Before you close the answer
Why this works
Tests whether you can turn "check the roadmap's credibility" into a concrete, repeatable test, rather than a vague instinct about trusting or not trusting a vendor.
Follow-up traps
"Isn't it unfair to judge a hard research problem by whether it's shipped on schedule?" Response: yes, which is exactly why the test isn't "did it ship on time," it's "can you show any real, even rough, progress or name the actual blocker."
"What if the vendor just says the prototype is confidential?" Response: that's still more than silence; ask for a redacted or aggregate number instead of a full technical readout, and treat continued refusal the same as no answer at all.
If pressed
Thermorun's renewed contract added a clause tying any future price increase tied to auto-diagnosis to a working demo on Thermorun's own data first, not just a roadmap mention, so the next promise has to clear the same bar before it can affect the price.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.