Artifact critiqueAdvancedAI Opportunity & Model Strategy / Competitive analysis in fast-moving AI / #11

Build the competitive teardown template you would use for AI products specifically.

SPARK a template built for the day a competitor's claim can't be checked at all

Say we build a document that lays out exactly how a rival's AI feature stacks up against ours, before a leadership meeting. Cartlogic sells route-optimization and contamination-detection software to waste haulers, using a camera on the truck to flag non-recyclables before they contaminate a whole load. Priyanka Deol is the product manager who had to build the actual template her team uses to size up a competitor, after leadership finally asked to see ten past ones side by side.

The direct answer
Build the teardown as six fixed, independently fillable sections, not one free-form narrative doc: claim versus evidence, model and data moat, failure modes and guardrails, pricing and segment bet, roadmap signal log, and a final verdict with an explicit confidence score. Every claim gets tagged confirmed, inferred, or unverified, so a rival's unverifiable eval claim doesn't stall the whole teardown, it just gets marked honestly and compared anyway.
Do this, in order
  1. Use six fixed sections for every competitor, never a blank page.Why: a free-form doc can't be compared to another free-form doc, no matter how well either one is written.
  2. Tag every claim confirmed, inferred, or unverified.Why: an AI competitor's quality claims are often genuinely unverifiable, and pretending otherwise makes the whole teardown dishonest.
  3. Give the model-and-data-moat section its own space, separate from features.Why: a feature list looks the same whether it's backed by proprietary data or a thin wrapper on someone else's model.
  4. Scale the depth to the threat, not to whoever has time that week.Why: a low-threat competitor doesn't need six full sections, and a high-threat one can't get away with two.
  5. Don't build an automated scraping pipeline for this on day one.Why: the template's value is the structure, not the tooling, and tooling can wait until the structure's proven itself.

How to answer this, stage by stage

Nobody is scoring whether your template looks polished. They're scoring whether it survives the one thing every real teardown eventually hits: a claim nobody can actually check.

Stage 1
Scope it to one real team
Say it like this
"I'll build the actual template Priyanka Deol's team at Cartlogic would use, for real AI competitors, not a generic 'how to do competitive analysis' slide."
Why this works
Grounds the artifact in a real team's real problem instead of a textbook exercise.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, what the team does today. Payoff, the habit I want the template to build. Anchor, the one structural decision. Risk, what happens the day a claim can't be checked. Keep out, what I won't build yet."
Why this works
Shows you're designing an artifact deliberately, not just listing sections that sound thorough.
Stage 3
Reframe the question
Say it like this
"The real ask isn't 'list what the competitor does.' It's 'which comparison would a room of leaders actually trust enough to make a call from.'"
Why this works
Moves the answer from a checklist to a decision-support tool.
Stage 4
Give the anchor, the one decision
Say it like this
"Six fixed sections, every time: claim versus evidence, model and data moat, failure modes, pricing and segment bet, roadmap signal log, verdict with a confidence score."
Why this works
This is the direct answer, specific enough a team could adopt it tomorrow.
Stage 5
Prove the anchor survives its own risk
Say it like this
"If Rubbleye claims 99 percent contamination-catch accuracy and there's no way to check it, the section still gets filled: claim, 99 percent; evidence, none found; tag, unverified. The comparison still happens."
Why this works
Shows the template doesn't break the first time a rival's claim can't be confirmed, which is the normal case, not the exception.
Stage 6
Say what you'd measure
Say it like this
"I'd track how long a teardown takes to produce, and whether every new one uses all six sections, not just the ones an analyst felt like filling in."
Why this works
Shows you're checking whether the template actually gets adopted, not just whether it looks good once.
Stage 7
Say what you're deliberately leaving out
Say it like this
"No automated scraping pipeline, no live alert system, no paid market-data subscription. This is a structured document first. The tooling can come once the structure's earned its keep."
Why this works
Shows judgment about what a v1 doesn't need, instead of a wish list of everything that could exist.
Stage 8
Close on the one line
Say it like this
"Build the template around six fixed sections and a confidence tag on every claim, because the day you can't verify something is the day a free-form doc quietly falls apart and a structured one keeps working."
Why this works
Restates the direct answer in one breath, exactly what a live follow-up rewards.

Let's learn

Here is what happens when a team builds ten competitor reviews the same way you'd write ten different essays: with feeling, but with no shared shape.

Cartlogic's team had been writing "competitor reviews" for two years, one loose document per rival, whoever had time that week. When leadership finally pulled ten of them at random to compare, the reviews ranged from 400 words to 4,000. Only two of the ten drew any line between a claim the competitor made and evidence that backed it up. Only one mentioned, anywhere, how confident the writer actually was in what they'd written.

Hand sketched flow diagram titled Priyanka's old routine, third box highlighted in red. Four boxes in sequence: Rival ships something, One meeting held, Notes taken once, Never reopened.
This is the whole process, before the template. It produced ten documents that couldn't be laid next to each other.

Here's the turn: the extra length of some reviews wasn't the problem, and the shortness of others wasn't either. The real problem was that no two of them measured the same thing, so leadership couldn't actually compare Rubbleye to any of the other nine rivals, no matter how well any single review was written. Ten documents, and zero real comparisons.

Ten past competitor reviews, before the template existed
100% 50% 0 6.0 hrs avg time to produce 20% used claim/evidence split 10% mentioned confidence
Six hours to write, on average, and still only one in ten reviews told the reader how much to actually trust it.
Hours per teardown, four quarters after the template shipped
6 hrs 3 hrs 0 1.8 hrs before Q4 after
The template didn't just make reviews comparable. It made them four times faster to write, because nobody was reinventing the shape of the document every time.

At its worst, a pile of well-meaning but incomparable competitor reviews doesn't just waste the hours spent writing them. It lets leadership make a roadmap call based on whichever review happened to be the most vividly written, not the one that was actually the most reliable.

The choice I would take back Each analyst wrote one big, free-form document per competitor, start to finish, all or nothing. That made sense when only one or two competitors mattered and one person owned the whole picture. It stopped making sense once ten rivals needed reviewing by four different analysts, none of whom had agreed on what a "complete" review even looked like.

What I would leave alone: a competitor with no visible AI feature at all, purely a manual, non-AI tool competing on price, doesn't need the full six-section teardown. A two-line note in the roadmap signal log is enough; the moat and failure-mode sections don't apply to a product with no model in it.

The lesson: a competitive review that reads beautifully and can't be compared to any other review is not a document. It's an opinion with good production values.

Now here is the same thing as a story

The short version above is what you'd say defending this template to Priyanka's own VP. Read this one for how the gap actually surfaced, in a room full of people who'd each written one of the ten.

Every time a competitor shipped something new, Priyanka Deol's team held one meeting, took one set of notes, and never looked at them again. It wasn't neglect. It just never occurred to anyone that the notes from March needed to line up with the notes from July, because nobody had ever tried to put them side by side.

Then leadership asked for exactly that: ten reviews, laid out on one table, to decide whether Cartlogic's contamination-detection roadmap needed a real course correction. The room went quiet fast. One review was four pages on Rubbleye's pricing with barely a sentence on its actual detection accuracy. Another was two pages entirely about a smaller rival's UI, with no pricing at all. A third simply said "seems similar to ours" and stopped.

Knowledge spark: what's a data moat? The part of a product that's hard to copy because it depends on data a company built up over time, not just a clever model. A camera that's seen two million contaminated loads has a moat a brand-new competitor's camera doesn't, even with the exact same underlying model.

Someone finally asked the question that mattered: "Which of these ten are we actually supposed to trust?" Nobody had an answer, because trust had never been part of the format. A four-page review wasn't more trustworthy than a two-line one. It was just longer.

Hand sketched comparison titled The day a claim cannot be verified. Left panel, claim 99 percent accurate, question mark box icon, tagged unverified no evidence found. Right panel, verdict stays usable, scale icon, compared on confidence not silence.
The old format had no way to say "we don't know" out loud. The new one is built to say it plainly and keep going anyway.
Ten reviews didn't fail because they were badly written. They failed because nothing about them could be laid side by side.

Priyanka built the six-section template that week: claim versus evidence, model and data moat, failure modes and guardrails, pricing and segment bet, roadmap signal log, and a final verdict with a confidence score from one to five. Every section, every competitor, every time. When Rubbleye's marketing claimed "99 percent contamination-catch accuracy" with no published methodology anywhere, the template didn't stall. It logged the claim, tagged it unverified, and moved on to the moat section, where the real gap turned out to be: Rubbleye's camera had only been live at eleven facilities, against Cartlogic's four hundred.

Hand sketched labeled parts diagram titled The anchor close up, what the template holds. A document icon at the center labeled Teardown Template, with six callouts around it: claim versus evidence, model and data moat, failure modes, pricing and segment bet, roadmap signal log, verdict confidence score.
Six sections, the same six, whether the competitor is loud, quiet, well-funded, or barely real yet.

Within one quarter, all fourteen new teardowns used the same six sections, every one of them. Average production time fell from six hours to two and a half, because nobody was staring at a blank page anymore, they were filling in a shape. And when the next roadmap debate came up, it took one meeting instead of three, because everyone was finally arguing from the same document.

Hand sketched timeline titled The quarter after the template shipped, week 4 emphasized in red. Week 1, template rolled out. Week 4, all new teardowns match. Week 8, first cross rival compare. Week 12, roadmap decided in one meeting.
Every point on this line used to take a fresh argument about format before anyone could even start comparing content.

SPARK, the template built to survive an unverifiable claimNot a features checklist with AI written on it. SPARK is what keeps a missing eval report from stalling the whole page.

S
Situation. What the team does today, honestly.
Ten free-form reviews, wildly different lengths, only one in ten mentioning any confidence at all.
Naming the real starting point, mismatch included, is what makes the new design credible.
P
Payoff. The habit you want the template to build.
Leadership compares rivals on the same six axes every time, instead of judging whichever review happened to be the best written.
The habit, not the document's polish, is the actual thing being shipped.
A
Anchor. The one structural decision.
Six fixed sections, every competitor, every time, with a confirmed, inferred, or unverified tag on every single claim.
This is the hardest step, and the one the whole template actually turns on.
R
Risk. What happens the day a claim can't be checked.
Rubbleye's 99 percent claim gets logged, tagged unverified, and the teardown keeps moving into the moat and pricing sections instead of stalling.
Proves the anchor was built to survive exactly the situation that happens most often, not the exception.
K
Keep out. What this template won't build yet.
No automated scraping pipeline, no live alert system, no paid data subscription on day one.
Naming what's deliberately absent is what separates judgment from a wish list.
Hand sketched icon list titled What we left for later. Three items: an automated live scraping pipeline, a real time competitor alert system, third party market data subscriptions.
All three of these are real ideas. None of them were the actual problem this quarter.

The recap, one line per letter: situation is naming that ten reviews couldn't be laid side by side, payoff is teaching leadership to compare on the same six axes every time, anchor is the six fixed sections with a confidence tag on every claim, risk is proving an unverifiable claim doesn't stall the page, and keep out is refusing to build tooling before the structure has proven itself.

And if you want to be sure it really works, try it somewhere elseSame five letters, a port-operations vessel-scheduling tool instead of a waste-hauling one. A different old decision breaks the second template.

Harborlist sells AI-based berth-scheduling software to port operators. Its first competitor-review process had the opposite problem from Cartlogic's: a single locked spreadsheet template, one row per competitor, that every analyst filled in identically, with no room to note anything that didn't fit a cell. Mapped onto SPARK: situation is a spreadsheet so rigid that a competitor's genuinely novel pricing model got squeezed into a "monthly fee" column that didn't fit it. Payoff is teaching the team to compare structure without losing texture. Anchor is adding one open "context" cell per section, capped at three sentences, so nuance has exactly one place to live without swallowing the whole format again. Risk is a competitor whose real threat is a partnership, not a feature, which the spreadsheet's feature-focused rows would have missed entirely without the context cell. Keep out is still refusing a fully free-text section, since that's the exact mistake Harborlist was trying to recover from. The old decision here is a merged-steps reversal: Harborlist had collapsed "record the fact" and "explain what it means" into one identical cell, removing the room for judgment that a pure fact needed to become an insight.

Hand sketched quadrant titled How deep a teardown should go, axes How big a threat and How unverifiable their claims are. Rubbleye sits high threat, hard to check. Small regional player sits low threat, easy to check. New entrant loud claims sits medium threat, very hard to check. Legacy competitor sits medium threat, easy to check.
Not every rival earns the same six full sections. The quadrant is what decides how deep to go, before writing a word.
Hand sketched decision tree titled How deep to go per competitor, root How big a threat is this rival. Four branches: low threat leads to quick scan two sections, medium threat leads to standard six section teardown, high threat loud claims leads to full teardown plus evidence test, customer already asking leads to emergency teardown this week.
The same four branches decide how much of the template to fill in, before anyone starts writing.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "six fixed sections, a confidence tag on every claim, and a teardown that survives a claim you can't verify," and stop.
Cost: there's no time to build all six sections for every past competitor before the next meeting. Say so honestly, and backfill only the two or three highest-threat rivals first, using the quadrant to pick them.
The model gets better, for real: if a rival's claim turns out to be genuinely confirmed once you check it, that's not a reason to drop the tag, it's the template doing exactly its job, telling you which claims to actually worry about.

Where people run it wrong.
They let teardown depth track whoever has free time that week instead of the actual size of the threat.
They treat an unverifiable claim as a reason to skip the section entirely, instead of tagging it honestly and moving on.
They build the scraping tool before proving anyone will actually use the six sections by hand first.

How to use it live. The moment someone asks you to design a competitive teardown, ask yourself: what happens on this page the day I can't verify the scariest claim on it? Build the template around surviving that day, not around looking thorough on a good one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "design the artifact you'd build" questions like this one?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It runs forward, designing a decision that has to survive its own worst case.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priyanka Deol, the product manager at Cartlogic who built the six-section teardown template after leadership compared ten inconsistent reviews.
3 · THE HABIT
What habit did the six-section template build in the team?
Tap to flip
ANSWER
Comparing every rival on the same six axes, instead of judging whichever review happened to be the longest or best written.
4 · THE ANCHOR
What's the one structural decision the whole template hangs on?
Tap to flip
ANSWER
Six fixed sections for every competitor, with a confirmed, inferred, or unverified tag on every single claim.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing one free-form document per competitor, a call that made sense with one or two rivals and one owner, and stopped making sense at ten rivals and four analysts.
6 · THE NUMBER
Fill in the blank: average teardown production time fell from 6 hours to ___ hours after the template rolled out.
Tap to flip
ANSWER
2.5 hours, because analysts were filling in a shape instead of staring at a blank page.
7 · THE REPLAY
A rival makes an unverifiable claim again. What changes with the new template in place?
Tap to flip
ANSWER
The claim gets logged and tagged unverified instead of stalling the whole review, and the teardown moves on to the moat and pricing sections where the real gap usually shows up.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Harborlist's vessel-scheduling tool. The reversal is a merged-steps call: recording a fact and explaining its meaning had been squeezed into one identical spreadsheet cell.

Check yourself Score: 0 / 0

Multiple choice
1. According to the anchor step, what should happen when a competitor's claim can't be verified?
  • A. The teardown stops until someone finds proof.
  • B. The claim gets left out of the document entirely.
  • C. The claim gets logged and tagged unverified, and the teardown continues.
  • D. The team assumes the claim is false by default.
Show hint
Look at the risk step in the SPARK recap.
Show answer
C. The template is built to survive exactly this case, honest tagging instead of stalling or guessing.
True or false
2. True or false: this answer recommends building an automated scraping and alert system as part of the template's first version.
  • True
  • False
Show hint
Look at "keep out" and the icon list of what got left for later.
Show answer
False. The template is a structured document first; scraping, alerts, and paid data feeds are explicitly left for later.
Fill in the blank
3. Fill in the blank: before the template, only ___ percent of past reviews mentioned any confidence level at all.
Show hint
Look at the bar chart of the ten past reviews.
Show answer
10 percent. Only one of the ten past reviews said anything about how confident the writer actually was.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Writing one free-form document per competitor. It made sense with one or two rivals and one owner, and broke down once ten rivals needed reviewing by four different analysts.
Short answer, apply it yourself
5. Think of a comparison document you've seen at work or school, a vendor evaluation, a college shortlist. What fixed sections would have made it easier to compare against a different one written by someone else?
Show hint
Think about what got left out of one version that the other version happened to include.
Show answer
Model answer: A vendor comparison usually needs fixed cost, support-response-time, and contract-lock-in sections, or two reviewers will end up weighing completely different things.
Short answer, where it wouldn't matter
6. Name a kind of competitor this answer says doesn't need the full six-section teardown.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A purely manual, non-AI competitor with no model in its product. The moat and failure-mode sections don't apply, so a short note in the roadmap log is enough.
Before you close the answer
Why this works
Tests whether you can design a comparison tool that stays honest about what it doesn't know, instead of one that only works when every fact happens to be checkable.
Follow-up traps
"Doesn't tagging half your claims 'unverified' make the whole document look weak?" Response: no, it makes the document honest, and a leadership team that's been burned by an overconfident review once will trust a document that admits its limits far more than one that hides them.

"What if two analysts disagree on whether a claim counts as confirmed or inferred?" Response: that disagreement is useful signal on its own, and the verdict section's confidence score is exactly where that tension gets surfaced instead of silently resolved by whoever wrote the section.
If pressed
The confidence score in the verdict section isn't a single number, it's the lowest of the five section-level confidence tags, so one badly-evidenced section can't get averaged away by four strong ones.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more