ConceptIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #22
What role should competitor pressure play in AI opportunity selection?
PICK · Duskrun's own model already proved a 14% earnings lift. A rival's unproven number still won five weeks on the roadmap
Duskrun tells rideshare drivers where and when to drive next, using each driver's own trip history instead of one citywide guess. Bevis Osterhaus owns EarnPath, the model behind those suggestions. Saffron Duncannon runs growth, and watched a rival's new feature pull more trade-press coverage in a week than Duskrun had gotten all year. Bevis had to decide how much of the roadmap that one launch should actually move.
The direct answer
Competitor pressure is real evidence, not a deciding vote. A rival shipping something tells you what they believed in enough to build. It does not tell you the thing is actually working for their users, or that it fits yours. Weigh it like any other signal, and check it before it moves your roadmap: does the rival show any visible sign their move is helping their own users, real engagement or retention, not just a launch. No evidence, no priority bump. At Duskrun, that one check would have taken an afternoon. Skipping it cost five weeks.
Do this, in order
Treat a competitor's launch as a clue, not a verdict, before it touches the roadmap.Why: this is the whole decision, and skipping it is exactly what let Duskrun build blind.
Check for real evidence the rival's move is actually working before reacting to it.Why: a launch announcement is not proof, driver reviews or any public number are the closest thing to it.
Never let a roadmap item skip the evidence bar just because it's tagged "competitive response."Why: that exact exemption is what let Peak Router skip the one check that would have caught it.
Weigh what a wrong call costs on each side before committing engineering time.Why: the cost only runs one direction, and it is not the direction that feels urgent in the moment.
Protect your own validated work from getting bumped by a rival's unvalidated one.Why: EarnPath had a proven 14% number and still lost five weeks to a feature with no number behind it at all.
Move fast and skip the gate on genuinely low-stakes parity, like a UI change with no model behind it.Why: not every competitor-inspired change carries the same risk, only the ones that touch what the model actually decides.
How to answer this, stage by stage
Nobody in the room is grading whether you sound calm about competitors. They're grading whether you can name the one check that separates weighing a rival's move from being run by it.
1
Ground it in one product and one real decision
Say it like this
"Let's ground this in something real. Duskrun tells rideshare drivers where and when to drive next, using each driver's own trip history. Bevis owns EarnPath, the model behind those suggestions. His growth lead, Saffron, just watched a rival's feature get more trade-press coverage in a week than Duskrun got all year, and she wants a matching feature on the roadmap by quarter's end."
Why this works
Keeps the interviewer grading a real decision, not a theory about competition in general.
2
Name the method out loud
Say it like this
"I'll use PICK. Position, the real role competitor pressure should play. Impact, what breaks on each side if you get that wrong. Cost asymmetry, which mistake is cheap and which is expensive. Kill criteria, the one test for whether a rival's move should actually move the roadmap."
Why this works
Two seconds of structure tells the room a method is coming, not a mood.
3
State the position plainly, before any story
Say it like this
"My position: competitor pressure is real evidence a need might exist, worth weighing. It should never be the main reason a feature gets built. A rival shipping something tells you what they believed. It doesn't tell you the thing is actually working, or that it fits your users."
Why this works
This is the direct answer, said early enough that the story can't blur it later.
4
Tie the line to something only a model raises
Say it like this
"This only matters because the rival's feature is a model, not a fixed screen. A generic surge-prediction model that's supposedly working somewhere else doesn't automatically transfer to your own driver base. It was trained on someone else's zones, someone else's traffic, someone else's driver mix."
Why this works
Keeps the answer anchored to model behavior, not generic advice about watching competitors.
5
Bring the real number
Say it like this
"Here's what happened. Duskrun's own EarnPath model had already shown a 14 percent earnings-per-hour lift for drivers who used it, in an eight-week pilot with 400 drivers. PulseShift's Surge Sync launched on one number: a self-reported 30 percent cut in idle time, no independent data behind it. Saffron pushed to match it anyway, and the roadmap gate that normally asks for proof got waived, because the item was tagged 'competitive response.'"
Why this works
A real number the whole argument would fall apart without, not a hypothetical.
6
Weigh what's lost on each side
Say it like this
"Ignore every competitor move, and you risk missing a real signal your users want something you haven't built. Chase every competitor move without checking it first, and you build reactive features off someone else's unproven bet instead of your own validated data. That second one is exactly what happened here."
Why this works
Naming both losses stops the answer collapsing into "always trust the rival" or "always ignore them," neither of which is a decision.
7
Name the cost asymmetry, then hand over the test
Say it like this
"Checking PulseShift's own driver reviews would have taken one afternoon. A sample of 90 reviews had only 21 describing any real change. Building Peak Router blind took three engineers five weeks, about 550 hours, and it landed at 4 percent weekly repeat-use against Duskrun's own 15 percent bar, while EarnPath's next validated update sat shelved the whole time. My test: does the rival show any real evidence the move is working for their own users, not just that it shipped. No evidence, no priority bump."
Why this works
Ends on something countable, and hands over a test instead of a feeling.
Let's learn
What does it actually mean when a rival ships something you don't have yet?
Duskrun tells rideshare drivers where and when to drive next, using each driver's own trip history instead of one citywide guess. About 26,000 drivers across dozens of mid-size cities open it before a shift.
Before any of this went sideways, EarnPath had already earned its keep. An eight-week pilot with 400 drivers showed that anyone who followed EarnPath's suggestions at least three times a week earned 14 percent more per hour than drivers who ignored it. That's a real, validated number, sitting in Duskrun's own data.
Five steps, and the whole argument sits inside the fourth one: the one moment nobody asked for proof.
Then PulseShift, a rival scheduling app, launched Surge Sync: a live heatmap of predicted surge zones, with a button that auto-suggests moving there. The press release claimed it cut driver idle time by 30 percent. That number came from PulseShift itself, in its own launch materials, with no independent data behind it anywhere.
One of these three had already earned its spot. One of them hadn't, and it's the one that got built first.
Here's the turn. Matching a competitor fast was never really the mistake. The mistake was matching one with no evidence behind it, at the exact moment Duskrun's own roadmap process was supposed to ask for some. Normally, any new roadmap item needs a one-line answer before it enters a sprint: what's the proof this solves a real driver problem. Items tagged "competitive response" were exempt from that question, on the idea that speed mattered more than paperwork for those, and because competitive-response items used to show up maybe once a year, the exemption's cost looked small on paper.
We didn't lose to a better model. We lost to a launch we never checked.
Knowledge spark: what does "self-reported" actually mean here?
It means PulseShift measured its own feature, using its own method, and published the number with nothing outside the company to check it against. That's not automatically false. It's just not the same thing as an independent result, and treating the two as equal is exactly what let a 30 percent claim skip the same bar Duskrun's own numbers had to clear.
The choice I would take back
Duskrun's roadmap gate exempted anything tagged "competitive response" from the same evidence question every other item had to answer. That made sense when those items were rare, maybe one a year, so the shortcut's real cost never showed up. It stopped making sense once a rival launch could arrive most quarters, and it walked Peak Router straight through the one check that would have caught it.
What I would leave alone: a few months earlier, PulseShift redesigned its own earnings summary screen, and Duskrun matched the layout within two weeks, no evidence gate, no debate. That was the right call. A screen redesign has no model behind it and no way to mislead a driver about how much they'll earn. Matching it fast cost almost nothing and risked even less.
The lesson: a model that's supposedly working somewhere else is not proof it will work on your own users. It's trained on someone else's zones, someone else's traffic, someone else's drivers. You have to check that, not assume it, and the check is usually smaller than the fear of falling behind makes it feel.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for what it cost Duskrun to learn it the slow way.
Bevis Osterhaus can tell, from one week of a driver's trip history, whether they're leaving real money on the table by working the wrong two hours of the day. He'd spent three years building EarnPath around that one fact: every driver's best shift is a little different, and a model that treats the whole city the same way is guessing on purpose.
The pilot had proven it. Four hundred drivers, eight weeks, and the ones who actually followed EarnPath's suggestions three or more times a week earned 14 percent more per hour than the ones who didn't. Bevis had the chart. He'd shown it to leadership twice. It was, by a wide margin, the best number his team had ever produced.
Saffron Duncannon ran growth, and she'd been chasing driver-side parity with PulseShift for two quarters, the kind of thing that shows up in a board update as a real logo, a real feature, a real headline. When Surge Sync launched, with a heatmap and a big claimed idle-time cut, three separate driver forums lit up with posts about it inside a week. Saffron brought it to Bevis with a deadline attached: something comparable, live, before quarter's end.
Five weeks of building sat between the launch nobody checked and the question somebody finally asked.
The team built fast, and mostly it built the right thing. The one place the plan and the calendar disagreed was the evidence question, since answering it properly meant reading through PulseShift's own driver reviews before writing a line of code, and that felt like a delay nobody had budgeted for. Someone did float a smaller version, a beta shown to just 5 percent of drivers before a full rollout, to catch a bad idea early and cheap. Saffron turned it down. A rideshare-driver newsletter published a quarterly "best app" comparison in six weeks, and a 5 percent beta wouldn't read as parity in time for that issue. So Peak Router, Duskrun's version of Surge Sync, got built for everyone, all at once: three engineers, five weeks, about 550 hours.
In week four, a new data scientist named Jethro Ingleside joined the team and sat in on his first standup. He listened to the sprint update, then asked the question nobody in the room had an answer for: had anyone actually checked whether Surge Sync was working for PulseShift's own drivers, or were they just matching a launch?
Neither of them was wrong about what they were holding. The launch date was real. So was the 14 percent. They just weren't the same kind of proof.
Nobody answered him that morning. The build kept going, because stopping it felt like the bigger risk with five weeks already spent.
We had a validated number sitting in our own data. We built five weeks off a number that belonged to someone else's press release.
Peak Router shipped in week five, to every driver at once. Two weeks later, the numbers came in: 4 percent of drivers who saw the feature used it more than once in a week, against the 15 percent weekly repeat-use bar Duskrun normally set for keeping any new feature past its first month. Earnings-per-hour for the drivers who did use it moved by 0.4 percent, which was inside normal week-to-week noise. Meanwhile, EarnPath's own next update, the one expected to push that validated 14 percent lift even higher, had sat untouched the entire five weeks, because the same three engineers were the ones building Peak Router.
I want to say the model was the problem. It wasn't, not really. Duskrun's engineers built exactly what the brief asked for. The real gap sat one step earlier, in the meeting where someone raised the evidence question and it got answered with a deadline instead of a check.
Months before, in the planning meeting where the "competitive response" exemption first got written into the roadmap process, an engineer had asked what would stop it from becoming the default excuse for skipping the gate whenever things felt urgent. The answer at the time was that competitive-response items were rare enough not to matter. Nobody was being careless. They were weighing a real cost, a slower roadmap process, against a risk that hadn't happened yet, which is exactly the kind of bet that looks fine until it doesn't.
Here's the replay. Same six weeks, same rival, same deadline. This time Jethro's question gets asked on day one instead of week four. Bevis and an analyst spend one afternoon reading through PulseShift's own driver reviews and forum posts: 90 mentions of Surge Sync, and only 21 describe any real change in earnings or idle time. That's thin evidence, not zero evidence, so instead of a five-week full build, the team runs a two-week limited pilot, a lightweight version of Peak Router shown to 8 percent of drivers, watching real earnings-per-hour and repeat-use instead of press coverage. Cost: about 150 hours. It doesn't clear the 15 percent kill line either, landing around 5 percent, so it gets dropped for a fraction of the cost, and EarnPath's validated update ships on schedule that same quarter.
One version of this story spends 550 hours proving what an afternoon and a two-week pilot would have shown for about a quarter of the cost. The other spends that same six weeks shipping the update that already had a real number behind it. Same rival, same deadline pressure, same model owner. The only thing that changed was whether the evidence question got asked before the build, or after.
What I'd tell myself, sitting in the meeting where "competitive response" first got its own exemption: a rule that only applies when nobody's paying close attention is not a small rule. It's the one that decides what actually happens the next time things feel urgent.
The four checks a competitor's launch actually has to pass
Not a reason to ignore rivals, and not a reason to chase every one of them. PICK only earns its keep here if it turns "they shipped it" into a real test before anyone touches a sprint board.
One path costs an afternoon and answers the question. The other costs five weeks and still doesn't.
PPosition. The real role competitor pressure plays.
Competitor pressure is real evidence a genuine need might exist. It's worth weighing the same way you'd weigh a support ticket, a survey, or a usage pattern. It should never be the main reason a feature gets built, because a rival shipping something proves what they believed in, not that it's actually working, and not that it fits your own users.
This isn't advice to ignore rivals. Duskrun's own dark-mode-style parity matches, the low-stakes ones, moved fast without any of this scrutiny, and that was correct.
State the position before any story, so it doesn't look reverse-engineered from what already went wrong.
IImpact. What's lost each way.
Ignore competitor moves completely, and you risk missing a real signal that a genuine, validated need exists, one your own users might actually be feeling too.
Let competitor pressure alone drive the roadmap, and you build reactive, me-too features chasing someone else's unproven bet, instead of your own validated user insight. At Duskrun that meant five weeks and 550 hours spent on a feature that never cleared the bar EarnPath had already cleared twice over.
Naming both losses stops the answer from collapsing into "trust every rival" or "ignore every rival," neither of which is a real decision.
CCost asymmetry. The heart of it.
A brief, deliberate check of whether a competitor's move reflects real, validated demand costs little time. Reading PulseShift's own driver reviews took one afternoon, about six hours for two people. Building a reactive feature purely because a competitor shipped something similar, with no validation of your own, costs real roadmap time on something that might not fit your users at all. Peak Router cost about 550 hours, and it landed at 4 percent weekly repeat-use, well under the 15 percent bar Duskrun uses to decide whether a new feature earns its keep. Start from the cheap mistake. Only commit real build time once evidence, not a launch date, says it's worth it.
KKill criteria. The one test.
Does the competitor's move come with any visible evidence it's actually working for their users, real engagement or retention, not just that it shipped. For Surge Sync, the answer was thin: 21 of 90 sampled driver reviews described any real change, the other 69 said some version of "same as the old surge map." Bevis's team considered one shortcut instead of reading the reviews themselves: just checking Surge Sync's average app-store star rating. It got rejected, because an average blends every complaint into one number and would have hidden the exact 69-of-90 pattern that actually mattered. Without real evidence, competitor pressure alone shouldn't be the deciding factor, whatever the deadline says.
Four branches, one root question: has anyone actually looked, or did the launch date do all the deciding.
Cost, by the numbers: validating first versus building the clone blind
Cheap, catches it earlyExpensive, catches nothing sooner
A two-week pilot to 8 percent of drivers would have cost about 150 hours and answered the question. Building Peak Router for every driver at once, with no evidence check first, cost about 550 hours, roughly three and a half times as much, and reached the same negative answer five weeks later.
The kill line, charted: Peak Router's weekly repeat-use rate, weeks 1 to 6
Below the kill lineCut at week 6
Duskrun tracks weekly repeat-use specifically, not first-week curiosity clicks, because a feature that gets tried once and abandoned looks fine on day one and empty by week three. Peak Router never got within twelve points of the bar.
The trade worth saying out loud: checking first costs a short, visible delay, an afternoon at minimum, sometimes a two-week pilot, right when a deadline makes speed feel like the only thing that matters. That delay is worth paying, because the alternative, building real engineering weeks off a rival's self-reported number, only shows its true cost after the hours are already spent and the validated work that got bumped is five weeks further from shipping.
And if you want to be sure it really works, try it somewhere else
Same four letters, a freelance translator's inbox instead of a driver's dashboard, and the unproven claim is about bids won instead of idle time cut.
WordRelay helps freelance translators decide which job listings to bid on, using each translator's own win-rate and turnaround-speed history by subject area, legal, medical, marketing, and so on. A rival platform, LinguaBoost, ships "AutoBid," a feature that submits bids automatically on every open job using flat, marketplace-average rates, with no personalization to any translator's own specialty. LinguaBoost's launch claims, self-reported, that AutoBid wins 40 percent more jobs. Someone on WordRelay's team proposes cloning it within the month.
Position: competitor pressure here is a real signal that translators might want less manual bidding. It shouldn't be the reason WordRelay builds a copy of AutoBid specifically, since AutoBid's claim has no independent number behind it. Impact: ignore the signal entirely, and WordRelay risks missing a genuine want for less manual bidding work. Clone AutoBid without checking, and translators could get auto-bid onto jobs outside their specialty, a legal contract bid by someone who translates marketing copy, which is a real quality risk, not just a wasted feature. Cost asymmetry: a morning spent reading translator forum comments about AutoBid costs about four hours; a sample of 40 comments found only 9 describing a real increase in jobs won, most said bid volume went up while win rate stayed flat or dropped, because the bids didn't match anyone's specialty. Building a full AutoBid clone blind would cost two engineers three weeks, about 200 hours, before finding out the same thing. Kill criteria: does LinguaBoost show any real, independent evidence AutoBid improves win rate for its own translators, not just bid volume. Today it doesn't, so WordRelay holds the clone and instead builds a smaller, personalized suggestion: which open jobs best match each translator's own specialty and speed, the thing its own data already says drivers, or here, translators, actually respond to.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the test: does the rival show real evidence, not just a launch, before anything else gets said.
Cost: no time or budget for even the afternoon check before a decision has to be made. Fine, but say the test out loud and make whoever owns the roadmap answer it, rather than letting the deadline answer it by default.
The model got better, for real: say a later version of Surge Sync posts real, independently verifiable retention gains. That's the moment it earns a look, not before it.
Where people run it wrong.
They let the loudest recent launch skip the same evidence bar every other roadmap item has to clear.
They treat "no evidence yet" as "no evidence ever," and never revisit a rival's move once it's actually proven.
They copy the whole feature instead of the one part of it, if any, that real evidence actually backs.
How to use it live. If you're ever asked whether a rival's launch should move your roadmap, buy yourself a second with one plain question, said out loud: "has anyone checked if it's actually working for their users, or are we just matching a headline?" That question is the whole method, asked instead of stated.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question about how much weight one outside factor should carry in a decision?
Tap to flip
ANSWER
PICK: state the real position, name the impact on each side, find which mistake is cheap versus expensive, then give the one test that actually decides it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bevis Osterhaus, who owns EarnPath, the personalized shift-recommendation model at Duskrun, a scheduling app for rideshare drivers.
3 · THE POSITION
What's the real role competitor pressure should play in opportunity selection?
Tap to flip
ANSWER
Real evidence a need might exist, worth weighing. Never the main reason a feature gets built, since a rival shipping something proves what they believed, not that it works or fits your users.
4 · THE GAP
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
Duskrun's own EarnPath pilot showed a validated 14 percent earnings lift. PulseShift's Surge Sync launched on a self-reported, unverified 30 percent claim with no independent data behind it.
5 · THE REVERSAL
What old decision would Bevis take back?
Tap to flip
ANSWER
The roadmap gate's exemption for anything tagged "competitive response." It made sense when those items were rare, maybe once a year. It stopped making sense once nearly every quarter had one, and it let Peak Router skip the one check that would have caught it.
6 · THE NUMBER
Fill in the blank: the evidence check would have taken about ___ hours. Building Peak Router blind took about ___ hours and reached ___ percent weekly repeat-use, against Duskrun's own ___ percent bar.
Tap to flip
ANSWER
About 6 hours, one afternoon. About 550 hours. 4 percent. 15 percent.
7 · THE KILL TEST
What's the one test for whether a rival's move should actually move the roadmap?
Tap to flip
ANSWER
Does the rival show any real, independent evidence the move is helping their own users, real engagement or retention, not just that it shipped. No evidence, no priority bump.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of PulseShift's unproven claim?
Tap to flip
ANSWER
WordRelay, a job-bid recommender for freelance translators. The role goes to LinguaBoost's AutoBid, which claims to win 40 percent more jobs with no independent data behind the number, the same shape as Surge Sync's 30 percent claim.
Check yourself Score: 0 / 0
Multiple choice
1. What actually explains why Peak Router failed to move the needle?
A. The model behind it was poorly built by Duskrun's engineers.
B. It launched a week later than planned, missing the trade newsletter's deadline.
C. Nobody had checked whether Surge Sync was actually working for PulseShift's own drivers before Duskrun copied it.
D. Drivers were never told the new feature existed.
Show hint
Think about what was missing before the build started, not what happened during it.
Show answer
C. Nothing was wrong with the engineering. The gap was upstream: the roadmap gate that would normally ask for evidence got waived, so a self-reported, unverified claim drove five weeks of work.
True or false
2. True or false: since Duskrun skipped its evidence gate to match Surge Sync, it should apply that same gate before matching any small, low-stakes competitor change, like a rival's redesigned earnings summary screen.
True
False
Show hint
Ask whether the change involves a model making a decision, or just a layout.
Show answer
False. A screen redesign with no model behind it carries almost no risk of misleading a driver. The gate exists for changes that touch what a model decides, not for cosmetic parity, and Duskrun correctly matched that kind of change fast, with no gate at all.
Fill in the blank
3. The evidence check would have taken about ___ hours. Duskrun built Peak Router blind instead, which took about ___ hours and reached only ___ percent weekly repeat-use, against Duskrun's own ___ percent bar for keeping a new feature.
Show hint
Look at stage 7 of the walkthrough, and the cost-asymmetry step in the recap.
Show answer
6 hours. 550 hours. 4 percent. 15 percent. The gap between those numbers is the entire cost asymmetry the answer turns on.
Short answer, name the rejected alternative
4. What old decision would Bevis take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back" in the Let's learn section.
Show answer
Model answer: Duskrun's roadmap gate exempted anything tagged "competitive response" from the same evidence question every other item had to answer. It made sense when those items were rare, maybe once a year, so the exemption's real cost never showed up. It stopped making sense once a rival launch could arrive most quarters.
Multiple choice
5. Which statement best matches PICK's position on competitor pressure in AI opportunity selection?
A. Ignore every competitor move; your own data is always the better source of truth.
B. Match every competitor move fast; speed matters more than proof.
C. Weigh a competitor's move as real evidence, but don't let it be the deciding factor without some independent sign it's actually working.
D. Only trust a competitor's move once it has been live for at least a year.
Show hint
Look at the direct answer's first two sentences.
Show answer
C. Competitor pressure is weighed, not obeyed and not dismissed. The kill criteria, real evidence the move is helping their own users, is what decides whether it earns a place on the roadmap.
Short answer, apply it yourself
6. Think of a product you use that recently copied a feature from a competitor. What evidence would you want to see before deciding that copy was the right call?
Show hint
Ask what would tell you the original feature was actually working, not just that it existed.
Show answer
Model answer: A grocery delivery app added a rival's "instant reorder" button shortly after the rival launched one. I'd want to know whether the rival's version actually raised repeat orders or basket size for their own users, not just that people tapped it once out of curiosity, before assuming the copy was worth building.
Before you close the answer
Why this works
Tests whether you treat a competitor's outside claim with the same evidence bar you'd demand of an internal one, or whether a launch date alone can override your own validated data. Most candidates either dismiss competitors entirely or chase every launch on reflex. The real judgment is checking before reacting, and knowing which changes are cheap enough not to need the check at all.
Follow-up traps
"Isn't checking first just slower, and doesn't that let competitors win?" Response: the check is an afternoon, not a quarter. The five weeks Duskrun actually lost came from skipping it, not from doing it.
"What if PulseShift's claim turns out to be true later, with real numbers behind it?" Response: then it earns a place on the roadmap once that evidence shows up, that's exactly what the kill criteria is for. Checking first doesn't rule out building it later, it just rules out building it blind.
If pressed
EarnPath is trained on each driver's own trailing 12 weeks of trips, refreshed weekly. That's why a citywide model like Surge Sync can't just be assumed to transfer: it's optimized for an average across all of PulseShift's drivers, not for any one person's own zones and hours, so a claim about its average performance says very little about whether it would actually help a specific Duskrun driver.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.