CaseAdvancedAI Opportunity & Model Strategy / Competitive analysis in fast-moving AI / #12

How do you evaluate a competitor's quality claims without access to their evals?

BOUND a press release number with nowhere to check it against

Picture a press release with a big, confident number in it, and no way for you to check where that number came from. Netreef builds an AI tool that identifies fish species from dock and vessel photos, for the catch reports fleets are legally required to file. Soren Kallio is the marine data analyst who tracks fishing-fleet compliance for Netreef, and this is what happened the week a rival, Tallyfin, published a species-identification accuracy claim nobody outside their own building could verify.

The direct answer
Treat the published number as an upper bound, not a fact. Break it down: how was it measured, is it blended across easy and hard cases, and what does it look like against your own honestly-measured number on the same kind of problem. Discount it by a stated, reasoned amount for self-report optimism, and give a range instead of accepting or rejecting the claim outright. Say plainly which single missing fact would sharpen that range the most.
Do this, in order
  1. Treat the published number as an upper bound, never the true figure.Why: a self-reported claim with no methodology can only overstate quality, never understate it, by design.
  2. State your discount assumption out loud, with a number attached.Why: "probably lower" isn't an estimate; a stated point value is something a colleague can actually argue with.
  3. Sanity-check it against your own honestly-measured number.Why: a claim that beats your own peer-reviewed-quality result, on a harder dataset you built over years, should raise real doubt.
  4. Give a range, not a single guessed number.Why: false precision on an unverifiable claim is just as dishonest as accepting it at face value.
  5. Name the one missing fact that would sharpen the range the most.Why: knowing exactly what you'd need to ask for next is what separates a real estimate from a shrug.

How to answer this, stage by stage

Nobody is scoring whether you believe the competitor's number. They're scoring whether you can build a defensible range around it, out loud, with the arithmetic showing.

Stage 1
Scope it to one real claim
Say it like this
"I'll answer this for Tallyfin's published 99.2 percent species-identification claim specifically, not for competitor claims in general."
Why this works
Anchors an abstract question in one real number that can actually be broken down.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, own my numbers, use a range, nail a sanity check, and name what would move it most."
Why this works
Signals a real estimation method instead of a gut reaction dressed up as analysis.
Stage 3
Reframe the question
Say it like this
"The real question isn't 'is 99.2 percent true or false.' It's 'what's the most honest range I can defend, using only what's outside their building.'"
Why this works
Moves from a yes-or-no judgment to a defensible estimate, which is what BOUND is actually for.
Stage 4
Give the one estimate
Say it like this
"Their published claim is an upper bound. Discounted for self-report optimism and an unstated blend across easy and hard species, I'd put their real, like-for-like accuracy somewhere between 75 and 90 percent."
Why this works
This is the direct answer, and it's a range, not a guessed point, which is exactly what the question is testing.
Stage 5
Show the arithmetic
Say it like this
"Self-reported AI accuracy claims with no published method typically run 15 to 20 points optimistic against blind, independent replication. 99.2 minus that lands around 80 to 84, which is already close to our own honestly-measured 91.4 percent, on a harder, more balanced test set."
Why this works
Shows visible numbers building to the estimate, not a confident tone standing in for math.
Stage 6
Run the sanity check
Say it like this
"If their 99.2 percent really held on a hard, balanced species mix, it would beat our own five-year, 12,000-image result. That's possible, but it's the kind of claim that needs real evidence before I'd believe it over ours."
Why this works
Compares the claim to something known, instead of evaluating it in a vacuum.
Stage 7
Name what would move the estimate most
Say it like this
"The single fact that would sharpen this range the most is whether their 99.2 percent is blended across all species or reported per species. If it's blended, the hard cases are probably hiding well below it."
Why this works
Shows you know exactly what to ask for next, which is what a good estimator does and a bad one doesn't.
Stage 8
Close on the one line
Say it like this
"Treat their number as an upper bound, discount it out loud with a real assumption, and hand over a range instead of a verdict, because a range is the only honest thing you can produce without their evals."
Why this works
Restates the direct answer in one breath, exactly what a live follow-up rewards.

Let's learn

Here is what happens when a competitor publishes a quality number with no way for anyone outside their building to check it.

Hand sketched icon list titled What quality actually decomposes into. Four items: measured how self report or third party, blended average or per species, how hard the test set really was, what independent signal exists at all.
A single accuracy number is four different questions wearing one costume.

Netreef's own species-identification model runs against a golden set of 12,000 labeled catch photos across 40 species, tested blind, on data the model has never seen. Overall, it scores 91.4 percent. Broken out, it scores 97 percent on common, visually distinct species like salmon and cod, and only 68 percent on visually similar juvenile groundfish, a group that makes up just 8 percent of total catch volume but causes most of the actual regulatory violations.

Netreef's own honest accuracy, blended versus broken out by difficulty
100% 50% 0 91.4% Overall, blended 97% Common species 68% Juvenile groundfish
One blended number hides a 29-point gap. This is exactly what a rival's single published figure could be hiding too.

Here's the turn: Tallyfin's marketing page claimed 99.2 percent species-identification accuracy, with no methodology, no third-party validation, and no mention of whether that number was blended or measured per species. The claim itself wasn't necessarily false. But treated at face value, it implied Tallyfin had beaten Netreef's own five-year, 12,000-image result, on a problem where the hardest cases matter most, using nothing anyone outside Tallyfin could check.

The choice I would take back Years earlier, Netreef itself told customers its accuracy was "lab-grade reliable" without ever publishing the actual eval methodology behind that phrase. That was fine when Netreef was the only real option in the market. It stopped making sense the moment a rival made an equally unverifiable claim, and customers had no way to compare either number, including Netreef's own.

What I would leave alone: a small rival with no public accuracy claim at all doesn't need this treatment. There's nothing to discount or sanity-check yet; the honest answer there is simply "no claim exists to evaluate."

The lesson: a number nobody can check isn't evidence. It's a starting point for an estimate, and refusing to build that estimate isn't caution, it's just letting the loudest number in the room win by default.

Now here is the same thing as a story

The short version above is what you'd say defending this estimate to Soren's own director. Read this one for how a simple question from someone new turned into an actual method.

His name is Soren Kallio. He has tracked fishing-fleet compliance data for five years, long enough to know which dockside lighting makes a young hake look like a young cod in a bad photo, and which regional fleets file catch reports late every single quarter.

Hand sketched flow diagram titled How the question actually got asked, third box highlighted in red. Five boxes in sequence: New hire joins, Reads rival's claim, Asks Soren why, Nobody has an answer, Soren builds the range.
Nothing dramatic happened here. Someone just asked a question the team had quietly stopped asking itself.

A new analyst joined Netreef's team in her second week and read Tallyfin's marketing page out loud in a meeting: "99.2 percent species-identification accuracy." She turned to Soren and asked, simply, "How do we know that's real?" Soren opened his mouth to answer and realized he didn't actually have one. Nobody did. The number had been sitting in a competitor-tracking slide for two months, unchallenged, mostly because it sounded confident and nobody had built a way to push back on it.

Hand sketched comparison titled Self-reported versus independently confirmed. Left panel, self-reported claim, document icon, no methodology no third party. Right panel, independently confirmed, scale icon, a real checkable number.
One of these is a number. The other is a number wearing a number's clothes.
Knowledge spark: what's self-report optimism? The tendency for a number reported by the same team that built the thing being measured to run higher than an independent, blind test would find, not necessarily through dishonesty, but because the easiest version of the test is the one that gets run and published first.

Soren spent the afternoon building an actual estimate instead of a shrug. He started with Netreef's own honestly-measured number, 91.4 percent overall, achieved on a golden set built over five years, tested blind. He assumed, based on published industry patterns for self-reported AI accuracy claims with no released methodology, that such numbers typically overstate real, independently-checked performance by 15 to 20 points. Applied to Tallyfin's 99.2 percent, that lands the estimate around 80 to 84 percent, which is suspiciously close to, and possibly below, Netreef's own number.

A number nobody can check isn't proof of anything. It's the top of a range, and the honest work is figuring out how far down that range really goes.
Hand sketched labeled parts diagram titled What's inside Netreef's own honest number. A gauge icon at the center labeled 91.4 percent overall, with four callouts: common species 97 percent, juvenile groundfish 68 percent, 12000 image golden set, blind held out test.
This is what a real number looks like once you're willing to show your own hard cases too.

Here is the decision Netreef would take back. Telling customers years ago that the system was "lab-grade reliable," with no published breakdown, was fine when nobody else was making claims to compare it to. It stopped being fine the exact week a rival made an equally vague, equally unverifiable claim, because now two unverifiable numbers were sitting side by side, and neither company could point to anything solid.

Soren's team started publishing their own species-level breakdown the same quarter, the 97 and 68 percent figures included, not hidden. Six weeks later, a prospective customer asked both companies for their per-species numbers. Tallyfin didn't answer. Netreef did, in a single afternoon, because the number had already been broken down the way Soren had learned to do it for a rival's claim.

BOUND, in one screenNot a verdict on whether the number is true. BOUND is what turns an unverifiable claim into a defensible range.

B
Break it down. State the equation out loud.
Estimated real accuracy equals the published claim, minus a self-report optimism discount, adjusted for whether the number is blended or per species.
Naming the equation first is what keeps the rest of the estimate honest.
O
Own numbers. State each assumption, and where it came from.
Self-reported claims with no published method typically run 15 to 20 points optimistic against blind, independent replication, based on published patterns across similar unverifiable AI claims.
This is the hardest step: naming a specific number you're willing to defend, not just a vague "probably lower."
U
Use a range, not false precision.
Estimated real accuracy: 75 to 90 percent, not a single confident number pretending to more certainty than the evidence supports.
A single guessed number implies confidence nobody actually has here.
N
Nail the sanity check.
99.2 percent on a genuinely hard, balanced species mix would beat Netreef's own 91.4 percent, achieved on a harder, five-year dataset. That's the kind of claim that needs real evidence before it's believed over your own.
Comparing to something known catches an estimate that doesn't survive contact with reality.
D
Direction. Which assumption would move it most.
Whether the 99.2 percent is blended across all species or reported per species. If blended, the hard cases are almost certainly hiding well below it.
Naming the single most valuable missing fact is what a good estimator does that a bad one skips.
What would change this estimate the most
Self-report optimism 16 pts Blended vs per-species 14 pts Test-set difficulty 9 pts points the estimate could swing
Knowing whether Tallyfin's number is blended or per species would sharpen the estimate more than almost anything else available.
Hand sketched decision tree titled What to tell a customer who asks, root Can we verify their claim independently. Four branches: yes third party exists leads to cite it directly, no self report only leads to give our honest range instead, partial some reviews exist leads to share range plus reviews, claim contradicts physics leads to say so plainly.
The same four branches decide what Soren actually says out loud, once the estimate is built.

The recap, one line per letter: break it down is stating the equation before touching numbers, own numbers is a stated 15-to-20-point self-report discount, use a range is 75 to 90 percent instead of one guessed figure, nail the sanity check is comparing it against Netreef's own harder-won 91.4 percent, and direction is naming that blended-versus-per-species reporting would move the estimate the most.

And if you want to be sure it really works, try it somewhere elseSame five letters, a small-town news service's AI writing assistant instead of a fisheries tool. A different claim, the same method.

Millbrook Wire runs an AI tool that drafts short local-news articles from public records for small papers. A rival service claimed its AI drafts were "98 percent factually accurate" in a funding announcement, with no published fact-checking methodology. Mapped onto BOUND: break it down is separating factual accuracy on simple record-based facts, like a permit number, from accuracy on inferred claims, like a quote's tone. Own numbers assumes self-reported factual-accuracy claims in AI writing tools typically run 10 to 15 points optimistic against a human fact-checker's blind review, based on published newsroom AI pilots. Use a range lands the estimate around 83 to 88 percent. Nail the sanity check compares it against Millbrook Wire's own fact-checked rate of 89 percent on simple record facts alone, before ever touching inferred claims, which the rival's number never separated out. Direction says the single most valuable missing fact is whether the 98 percent excludes inferred claims entirely, since record facts alone are far easier to get right than anything requiring judgment.

Hand sketched quadrant titled Which competitor claims are worth investigating, axes How loud and public the claim and How consequential to our customers. Tallyfin's 99.2 percent claim sits loud and high consequence. Small rival vague claim sits quiet and low consequence. Legacy tool no AI claim sits quiet and very low consequence. New entrant quiet pilot sits medium loud, high consequence.
Not every claim earns a full BOUND estimate. This is what decides which ones do.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "treat the number as an upper bound, discount it with a stated assumption, and give a range, not a verdict," and stop.
Cost: there's no time to research real self-report optimism rates before the next customer call. Say so honestly, and use a conservative, clearly-labeled placeholder discount until better data exists, rather than skipping the estimate entirely.
The model gets better, for real: if a rival eventually does publish real, third-party-verified methodology behind their number, that's not a threat to this method, it's the evidence test finally landing, and the range collapses into a real number worth trusting.

Where people run it wrong.
They either accept a rival's number at face value or dismiss it outright, instead of building an actual estimate in between.
They give a single guessed number instead of a range, pretending to more precision than the evidence supports.
They never check the claim against their own honestly-measured result, missing the sanity check that would have caught an implausible number.

How to use it live. The moment someone quotes you a rival's unverifiable number, ask yourself: what would I have to assume about how they measured this for that number to be real? Say the assumption out loud, and build the range from there.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an estimation question like this one?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. It shows arithmetic instead of a confident guess.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Soren Kallio, the marine data analyst at Netreef who built the actual estimate after a new hire asked how anyone knew a rival's claim was real.
3 · THE HABIT
What habit changed once the question got asked out loud?
Tap to flip
ANSWER
The team stopped letting an unverifiable competitor number sit unchallenged in a slide, and started building a real, stated estimate instead.
4 · THE OWN NUMBER
What discount assumption does this answer use for self-reported claims with no published method?
Tap to flip
ANSWER
15 to 20 points optimistic compared to blind, independent replication, based on published patterns in similar unverifiable AI claims.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling customers Netreef's system was "lab-grade reliable" with no published methodology, fine when there was no rival to compare against, wrong once a rival made an equally vague claim.
6 · THE NUMBER
Fill in the blank: the final estimated range for Tallyfin's real accuracy is ___ to ___ percent.
Tap to flip
ANSWER
75 to 90 percent, built by discounting the published 99.2 percent and sanity-checking it against Netreef's own 91.4 percent.
7 · THE REPLAY
A prospective customer asks both companies for a per-species breakdown. What happens now?
Tap to flip
ANSWER
Netreef answers within an afternoon, because the breakdown already exists from building the estimate. Tallyfin doesn't answer at all.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the key missing fact there?
Tap to flip
ANSWER
Millbrook Wire, a small-town news AI writing tool. The key missing fact is whether the rival's 98 percent claim excludes inferred claims entirely, not just record-based facts.

Check yourself Score: 0 / 0

Multiple choice
1. According to the direct answer, how should a self-reported, unverifiable competitor claim be treated?
  • A. As false, until the competitor proves otherwise.
  • B. As true, since there's no way to disprove it.
  • C. As an upper bound, to be discounted into a defensible range.
  • D. As irrelevant, since it can't be verified either way.
Show hint
Look at the direct answer and the B and U steps.
Show answer
C. The claim can only overstate quality, never understate it, which is exactly why it functions as an upper bound.
Fill in the blank
2. Fill in the blank: Netreef's own honestly-measured accuracy on juvenile groundfish, the hardest species group, is ___ percent, far below its 91.4 percent blended average.
Show hint
Look at the bar chart of Netreef's own accuracy broken down by difficulty.
Show answer
68 percent. This 29-point gap between the blended and hardest-case numbers is exactly what a rival's single published figure could also be hiding.
True or false
3. True or false: this answer concludes that Tallyfin's 99.2 percent claim must be a lie.
  • True
  • False
Show hint
Look at the sanity check step, and what it says the claim would need to be true.
Show answer
False. The claim isn't declared a lie, it's treated as an unverified upper bound, and a specific range is built around it instead of a verdict.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Calling Netreef's own system "lab-grade reliable" with no published methodology. It made sense with no rival to compare against, and stopped making sense once a rival made an equally vague claim.
Short answer, apply it yourself
5. Think of a product that advertises a quality claim you've never been able to verify, like a battery life estimate or a cleaning product's "removes 99.9 percent of germs." What range would you estimate for the real number, and why?
Show hint
Think about how the number was likely measured, and under what conditions it would be hardest to reproduce.
Show answer
Model answer: A "12-hour battery life" claim is usually measured under minimum brightness and no active use, so a real range under normal use might honestly be 6 to 9 hours.
Before you close the answer
Why this works
Tests whether you can build a real, numbered estimate under real uncertainty, instead of either accepting a competitor's marketing at face value or dismissing it as unknowable.
Follow-up traps
"Isn't a 15-to-20-point discount just a number you made up?" Response: it's a stated, checkable assumption, not a fact, and the whole point of BOUND is showing that assumption out loud so someone else can argue with the specific number instead of the vague feeling behind it.

"What if Tallyfin's claim is actually true, and you've unfairly discounted a better product?" Response: then the direction step already points at the fix, ask them directly whether the number is blended or per species, and if they answer, the range collapses into a real number instead of an estimate.
If pressed
Netreef's own 91.4 percent figure came from a golden set stratified to match real catch-volume proportions, not an even split across all 40 species, which is exactly the kind of methodology detail Tallyfin's press release never specified either way.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more