Which competitors matter more: incumbents adding AI or AI-native startups? Defend it.
Ground it in Castlemere Trust Bank, which runs SentinelAI, a fraud-detection layer bolted onto its own fifteen-year-old rules engine. Iwan Sokolov reviews flagged transactions there. A younger rival, Falcondrift Labs, builds fraud detection with no legacy system underneath it at all.
- Track startups' technical trajectory, not just their press releases.Why: a startup's real threat shows up in a model change or a new data source, months before it shows up in marketing.
- Know exactly what your own AI feature is bolted onto, and what that ceiling is.Why: a legacy system underneath your own AI defines what it structurally cannot do, no matter how good the model on top gets.
- Set a kill criterion up front: what would make incumbents matter more.Why: if a jurisdiction requires fully explainable fraud decisions, a compliant incumbent tool can beat a more accurate black-box one regardless of raw capability.
- Treat a rival incumbent's loud AI announcement as informational, not urgent.Why: it's usually bounded by the same legacy constraints your own team already understands well.
- Budget engineering time for a possible rebuild, not just a bolt-on feature.Why: catching up to a startup's real capability sometimes requires rebuilding the data architecture underneath, not adding a feature on top.
How to answer this, stage by stage
Nobody is scoring whether you pick a side. They're scoring whether you can defend it when the interviewer pushes on the exact case where you'd be wrong.
Let's learn
The artifact here is a queue: a list of flagged transactions on Iwan Sokolov's screen at Castlemere Trust Bank, ranked by how confident SentinelAI is that each one is fraud. For years, that queue moved at a predictable pace, a few hundred flags a week, and Iwan reviewed every single one by hand.
SentinelAI itself is an AI layer bolted onto Castlemere's fifteen-year-old rules engine: velocity checks, blacklists, thresholds tuned over a decade. Adding the AI layer on top improved the queue's accuracy without touching what sat underneath it. That was, for years, exactly enough.
Here's the turn: a wave of synthetic-identity fraud, fabricated people built from real stolen data fragments, started hitting the industry. Falcondrift Labs, an AI-native rival with no legacy system to work around, had built its fraud model around a relationship graph from day one: whose data fragments connect to whose. Peer banks running Falcondrift's tool caught seventy-one percent of the new fraud type. SentinelAI, unable to plug a graph signal into a rules engine that had never been built to hold one, caught twelve percent.
At its worst, the gap doesn't stay a fraud-detection statistic. It shows up in real dollars lost, and eventually in a regulator's question about why a known fraud pattern took six weeks to catch.
What I would leave alone: I wouldn't rebuild SentinelAI from scratch for every fraud type. Most fraud still fits the rules-engine's existing signals just fine; the rebuild only matters for the specific pattern the legacy system structurally can't see.
The lesson: an incumbent's own AI feature inherits every constraint of whatever it's bolted onto, and a startup with nothing underneath it can reach a capability an incumbent structurally cannot match without tearing something out first.
Now here is the same thing as a story
The short version above is what you'd say defending this position under interview pressure. Read this one for how the six weeks actually felt on the floor.
The fraud operations floor at Castlemere Trust Bank gets loud around two in the afternoon, when the day's flagged transactions pile up before the evening processing cutoff. Iwan Sokolov had worked that floor for ten years. He could glance at a flagged transaction and tell, before reading a single detail, whether it smelled like a stolen card or something stranger.
The habit thinned in beats that felt reasonable each time. When the synthetic-identity ring hit, SentinelAI's flag volume quadrupled overnight, mostly low-confidence flags the model couldn't resolve on its own. Iwan and his team, unable to review four times their normal volume, narrowed to only the high-confidence flags within three weeks, a sensible triage under real pressure.
What they didn't know was that the synthetic-identity cases were landing almost entirely in the low-confidence pile, exactly the pile the team had just stopped reviewing, because SentinelAI's rules-engine backbone had no graph signal to raise its confidence on a pattern it had never been built to recognize in the first place.
A trigger arrived sideways: a compliance officer at a peer bank running Falcondrift's tool mentioned, at an industry roundtable, that they'd caught an entire ring "off one weird graph connection" three weeks earlier. Iwan's manager, sitting in the same session, asked what a graph connection even was.
The old decision, to bolt SentinelAI onto the existing rules engine instead of rebuilding the underlying signal architecture, had been made three years earlier, in a room where rules-engine fraud was still catching the overwhelming majority of real cases. Rebuilding then would have meant months of engineering time spent on a problem that, at the time, barely existed. The call was sound. It stopped being sound the exact week synthetic identities became the dominant new fraud pattern, and nobody had a trigger built in to notice the moment that happened.
The replay: same peer-bank comment, same conference table, but Castlemere already tracks startups' technical trajectory, not just their press releases, as a standing practice. Falcondrift's graph-based approach gets flagged as a real threat the quarter it first ships, not the quarter a competitor mentions it in passing. Castlemere budgets three months to add a lightweight graph signal alongside the existing rules engine, catching the next synthetic-identity wave at closer to fifty percent instead of twelve, before six weeks of losses force the question.
What Iwan's floor took from it wasn't "the startup is smarter." It was that an incumbent's own AI is only ever as capable as whatever it's bolted onto, and the rival worth watching closely is the one with nothing underneath to hold it back.
PICK, the four letters that defend a positionNot a coin flip between two kinds of competitor. PICK is what makes the pick defensible under pressure.
The recap, one line per letter: position is startups over incumbents for where real capability comes from, impact is a competitor-deck line versus six real weeks of unseen fraud losses, cost asymmetry is that the startup miss is the one that's quiet and structurally hard to undo, and kill criteria is a hard regulatory explainability requirement that would hand the advantage back to incumbents regardless of capability.
And if you want to be sure it really works, try it somewhere elseSame four letters, municipal waste routing instead of bank fraud. A different old decision breaks this one.
A mid-size city's waste-management department compares its longtime incumbent hauler, which just bolted a route-optimization dashboard onto its existing dispatch software, against a newer, smaller hauler built entirely around live sensor data from bin fill-levels. Mapped onto PICK: position is that the AI-native hauler matters more to watch, since its routing improves as sensor coverage grows, while the incumbent's dashboard is capped by dispatch software that was never built to ingest live fill-level data at all. Impact is that missing the incumbent's dashboard update costs nothing, since it changes little about actual route efficiency; missing the AI-native hauler's real advantage costs the city a contract renewal built on stale assumptions about which routes are actually full. Cost asymmetry is that the incumbent's miss is visible and reversible at the next contract cycle, while the AI-native hauler's real edge compounds quietly for years as its sensor network grows, getting harder to catch up to the longer it's ignored. Kill criteria is that if the city's contracting rules require a hauler with ten years of continuous local service history, the incumbent wins regardless of routing quality, and the pick reverses. The old decision here isn't a bolted-on rules engine, it's a delegation choice: the incumbent hauler delegated route planning entirely to dispatchers' own judgment for years, so when it finally added AI, nobody on staff had the sensor-reading habits or trust calibration to know when to override a route the model got wrong, and drivers either followed bad routes blindly or ignored the tool completely.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "watch AI-native startups more closely, because their real capability jumps are quiet and expensive to miss, while incumbents are loud and bounded by their own legacy systems," and stop.
Cost: there's no budget to monitor every small startup in the space. Focus tracking on the two or three with genuinely novel data access or architecture, not headcount or funding size.
The model gets better, for real: if an incumbent's bolted-on AI genuinely closes the capability gap through a real architecture rebuild, not just a feature add, that's the moment to reconsider, since it means the legacy constraint itself got removed.
Where people run it wrong.
They track competitors by headline size, not by what's actually holding their capability back or open.
They assume "startup" automatically means "more dangerous," without checking what data or architecture actually backs the claim.
They forget to name a kill criterion, turning a defensible position into a permanent bias that never gets to be wrong.
How to use it live. The moment an interviewer asks you to pick a side, ask yourself: which miss would cost more to catch up from, six months from now? Pick that side, and say exactly what would change your mind.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't Castlemere just acquire Falcondrift instead of rebuilding?" Response: possibly, but that's still a response to having correctly identified the threat early, which is the actual point of watching startups' technical trajectory, not their funding announcements.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Competitive analysis in fast-moving AI
- #1 How do you run competitive analysis in a market where the landscape changes monthly?
- #2 What is the difference between a competitor's feature and a competitor's advantage in AI?
- #3 Describe how you would test a competitor's AI feature to find its real limitations.
- #5 How do you assess whether a competitor's capability is a moat or a thin wrapper?
- #6 What does it mean when your competitor and you both build on the same model provider?
- #7 Explain how to compete when a model provider could ship your feature natively.