ConceptIntermediateAI Opportunity & Model Strategy / Competitive analysis in fast-moving AI / #18
How do you track model provider announcements as a competitive input?
SPARK a support-chatbot team that scrolled the same three feeds every morning until nobody could say which announcement actually mattered
Picture this product before anyone has touched a single setting on it: a support chatbot for mid-size retailers, and a small product team trying to keep up with how often the ground shifts under it. Ember Desk builds that chatbot. Yusuf Adeyanju runs product there, and for a long time, "keeping up with model providers" meant one person scrolling three feeds before coffee.
The direct answer
Build a small, tiered watch process, not a habit of scrolling feeds. Score each provider announcement on two things: how big the capability delta actually is, and how soon it could touch your specific product surface. Route big-and-near items to a same-week team review, small-or-distant ones to a watchlist nobody has to act on yet, and name one person who owns the triage so it survives someone being on vacation. Skip chasing every rumor; that's not day one, and it never needs to be.
Do this, in order
Score every announcement on capability delta and time-to-impact, not on how loud it is online.Why: a flashy demo with no real delta and a quiet paper with a big one deserve opposite amounts of attention.
Name one owner for the triage, even in a two-person team.Why: an unowned watch process quietly stops happening the first busy week, and nobody notices until it's cost you a quarter.
Set a real cadence: a short weekly scan, a longer monthly deep dive, a roadmap check once a quarter.Why: without a cadence, watching becomes whatever's left over after everything else, which in practice is nothing.
Route only the big, near-surface items to a real team review; log the rest to a watchlist.Why: reacting to everything burns the same attention a genuine roadmap-breaking announcement will need later.
Keep one lightweight artifact, not a full write-up per announcement.Why: a process that takes longer to run than the announcements take to matter won't survive contact with a busy quarter.
How to answer this, stage by stage
Nobody is scoring whether you can name five AI labs. They're scoring whether your tracking process would survive a normal busy month, not just the week you set it up.
Stage 1
Scope it to one real team
Say it like this
"I'll design this for a real team: Ember Desk, a support chatbot for retailers, small enough that nobody has a full-time job just watching model providers."
Why this works
A resourced, specific team keeps the answer honest about what's actually sustainable.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK: situation, how they track it today; payoff, the habit I want to build; anchor, the one structural decision; risk, what happens the day a provider surprises them; keep out, what I won't build yet."
Why this works
Shows you're designing a system, not just listing sources to check.
Stage 3
Name the situation, plainly
Say it like this
"Right now, tracking provider announcements at Ember Desk means Yusuf scrolling three feeds most mornings, reacting late, with no record of what got checked and what didn't."
Why this works
Names the real, unglamorous starting point instead of assuming a mature process already exists.
Stage 4
Give the anchor, the one decision
Say it like this
"The anchor is a simple two-axis score on every announcement, capability delta and time-to-impact, with a named owner and a fixed weekly slot, not an open-ended habit of scrolling."
Why this works
This is the direct answer, specific enough that someone could build it by Friday.
Stage 5
Prove the anchor survives its own risk
Say it like this
"The day a provider ships a surprise capability that touches our roadmap, the score already flags it as big-and-near within the week's scan, and the named owner escalates it that same day, instead of the team finding out from a customer first."
Why this works
Shows the design was built to survive exactly the situation it's meant for.
Stage 6
Name what you're deliberately leaving out
Say it like this
"I'm not chasing every rumor, and I'm not writing a full benchmark comparison for every minor point release. That's real effort spent on things that will never touch our roadmap."
Why this works
Shows judgment about what a lean team doesn't need yet, not a wish list of everything possible.
Stage 7
Close on the one line
Say it like this
"Score announcements by capability delta and time-to-impact, give one person the job of triaging them on a fixed cadence, and you'll catch the one that matters without burning attention on the hundred that don't."
Why this works
Restates the direct answer in one breath, exactly what a live follow-up rewards.
Let's learn
Here is what happens when a small team tries to track a fast-moving field with nothing but attention and good intentions.
Before Ember Desk built any real process, tracking model provider announcements meant one person, usually Yusuf, checking a couple of feeds most mornings, roughly fifteen minutes, sometimes skipped entirely during a busy sprint. It caught the big, loud announcements everyone already knew about within a day or two.
This is the whole process before any structure existed. It only ever caught what was already loud.
Here's the turn: the missed loud announcements were never actually the problem. The real cost was the quiet ones, a smaller capability update that happened to land exactly on Ember Desk's own roadmap bet, that nobody flagged for three weeks because nothing about the ad hoc scroll ever separated "big and relevant" from "interesting but irrelevant."
Major model provider announcements per quarter, across the field
Roughly four times the volume in eighteen months, on a process that was never designed to scale past a quick morning scroll.
Detection lag: days between an announcement landing and Ember Desk flagging it as relevant
Four months of an unscored habit held lag near three weeks. One scored, owned process cut it to two days almost immediately.
At its worst, an unscored, unowned watch habit doesn't just miss things occasionally. It quietly stops happening altogether during exactly the weeks a team is busiest, which are usually the weeks something important is also shipping elsewhere.
The choice I would take back
Ember Desk once kept a lightweight weekly "watch" document, a few bullet points on anything notable that week, then cut it after two quarters because it felt like busywork next to a genuinely good internal eval suite. That made sense when the eval suite was catching every quality regression that mattered. It stopped making sense once the risk wasn't quality regression at all, it was a competitor's capability shift the eval suite was never built to see coming.
What I would leave alone: Ember Desk's internal eval suite itself didn't need to change. It was never the thing meant to catch a competitive shift; asking it to do that job was the actual mistake, not the eval suite's design.
The lesson: tracking a fast-moving field isn't about reading more. It's about scoring what you read so the one thing that matters doesn't get buried in the hundred things that don't.
Now here is the same thing as a story
The short version above is what you'd say defending this process to your own engineering lead. Read this one for the whiteboard that quietly told the truth before anyone read it.
Yusuf Adeyanju is good at spotting which customer complaints are really the same complaint wearing different words, a skill that made him a natural fit for triage of any kind. At the end of most weeks, someone photographed the small team's shared whiteboard before wiping it, a habit that started as a joke and stuck.
For the first year, that whiteboard photo mostly showed customer bugs and feature requests. Provider announcements, when Yusuf noticed one at all, got a quick mention in Slack and nothing more.
This is the one design decision the whole answer hangs on, made inspectable instead of just described.
Then a new hire, two weeks into the job, asked a question in a planning meeting nobody could answer cleanly: "didn't a provider ship something like our whole roadmap item three months ago? Why are we still building it like it's novel?" Nobody had connected that announcement, half-noticed in passing, to the actual feature Ember Desk was six weeks into building.
Knowledge spark: what's a capability delta?
The real, measurable gap between what a model could do before an announcement and what it can do after. Not the size of the press release, the size of the actual change in what's possible.
Yusuf went back through three months of Slack mentions and found it: an announcement, barely discussed, that had quietly made a core piece of Ember Desk's in-progress feature something any competitor could now build in a fraction of the time. Nobody had flagged it, because nothing about the ad hoc process ever asked whether an announcement touched anything the team was actually building.
We did not miss a headline. We missed the one line inside it that mattered to us specifically, and nothing was built to catch that.
Same surprise announcement, two very different mornings after it lands.
The fix that followed was small: a weekly thirty-minute scan, scored on capability delta and time-to-impact, with Yusuf named as the one owner, and a quick escalation path for anything that scored high on both. Six weeks later, a similar announcement landed. This time it hit the watchlist within the same week, got flagged as high-impact within a day, and the roadmap adjusted before three more weeks of building the wrong version of the same feature.
SPARK, for a process nobody has time to runNot the classic "design the interface" story. This time the anchor is a habit, not a screen.
S
Situation. How it gets done today, without a system.
One person scrolling a couple of feeds most mornings, with no record of what got checked, reacting late to whatever was loud enough to notice.
Naming the real, unglamorous starting point is what makes the fix credible.
P
Payoff. The habit you want to build.
A team that scores every announcement the same way, instead of reacting to whichever one happened to trend that week.
The habit, not the tracking tool, is the actual thing being built here.
A
Anchor. The one structural decision.
A two-axis score, capability delta and time-to-impact, with one named owner and a fixed weekly slot, not an open-ended habit of scrolling.
This is the hardest step, and the one the whole process actually turns on.
R
Risk. What happens the day a provider surprises you.
A high-delta, near-surface announcement gets flagged within the week's scan and escalated the same day, instead of surfacing three weeks late in a planning meeting.
Proves the anchor was built to survive exactly the situation it exists for.
K
Keep out. What this process won't do on day one.
No chasing every rumor, no full benchmark write-up for every minor point release, no daily monitoring that nobody will sustain past a busy month.
Naming what's deliberately left out is what separates a lean process from a wish list.
The recap, one line per letter: situation is naming the ad hoc scroll that only ever caught what was already loud; payoff is teaching the team to score consistently instead of reacting to volume; anchor is the two-axis score with a named owner and fixed cadence; risk is proving a surprise announcement gets triaged within a day instead of three weeks; keep out is refusing to chase every rumor or write a report nobody will read.
And if you want to be sure it really works, try it somewhere elseSame five letters, a photo-editing app instead of a support chatbot. This time the reversal is a removed affordance, not a cut habit.
Snapfen lets users edit photos with natural-language prompts instead of manual sliders. Louisa Petrakis runs product there, competing in a category where a base model update can change what's possible for every competitor overnight. Mapped onto SPARK: situation is that Snapfen's small team had no shared process at all for tracking model releases, each engineer following whatever sources they personally liked. Payoff is a team that reacts to a real capability shift in days, not whenever someone happens to mention it. Anchor is the same two-axis score, capability delta and time-to-impact, applied to image-generation and editing releases specifically, with a named owner. Risk is that a competitor's overnight adoption of a new base model's editing capability gets caught within the week instead of showing up first in a customer's side-by-side comparison post. Keep out is not rebuilding the entire editing pipeline every time a new model ships, only the pieces the score actually flags as both big and near. The old decision here isn't a cut habit, it's a removed affordance: Snapfen once had a shared internal channel where anyone could flag a provider release worth discussing, and quietly let it die when nobody was assigned to check it, so it looked abandoned rather than useful and people stopped posting to it.
The same four branches sort a support-chatbot announcement and a photo-editing one alike.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "score every announcement on capability delta and time-to-impact, name an owner, and run it on a fixed cadence," and stop.
Cost: there's no budget for a dedicated competitive-intelligence hire. Say so honestly, and fold the thirty-minute weekly scan into an existing role instead of waiting until you can afford a new one.
The model gets better, for real: if a provider ships something genuinely useful to your own product, that's still an input worth scoring, sometimes the right move is adopting a capability instead of racing to match it.
Where people run it wrong.
They read everything and score nothing, which feels thorough and catches almost nothing that matters.
They build a process with no named owner, so it quietly dies the first busy month.
They wait for a customer or a new hire to point out a shift the team should have caught internally.
How to use it live. The moment someone asks how you'd track provider announcements, ask yourself out loud: what's the one score that tells me whether to act today or file it and move on? Build the whole process around that one score.
Only the upper-right corner gets a same-week team review. Everything else waits on the list.
A lean team earns the right to skip all three of these, on purpose, not by accident.
Three fixed slots, none of them optional, none of them large.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "design the process you'd use" questions?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It runs forward, building a design decision that has to survive its own risk.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Yusuf Adeyanju, who runs product at Ember Desk, a support chatbot company, and is good at spotting when different complaints are really the same one.
3 · THE HABIT
What did Ember Desk's team rely on before building a real process?
Tap to flip
ANSWER
One person scrolling a couple of feeds most mornings, with no scoring and no record of what had already been checked.
4 · THE ANCHOR
What's the one structural decision this whole process hangs on?
Tap to flip
ANSWER
Scoring every announcement on capability delta and time-to-impact, with one named owner and a fixed weekly slot to run the triage.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Cutting the lightweight weekly "watch" document because it felt like busywork next to a good internal eval suite, which was never built to catch a competitive shift in the first place.
6 · THE NUMBER
Fill in the blank: major provider announcements per quarter rose from about 4 to about ___ over six quarters.
Tap to flip
ANSWER
15, roughly four times the volume, on a process never designed to scale past a quick morning scroll.
7 · THE REPLAY
Same kind of surprise announcement, watch process already in place. What changes?
Tap to flip
ANSWER
It's flagged within the same week's scan, scored high on both axes, escalated the same day, and the roadmap adjusts before weeks of building the wrong version of a feature.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Snapfen, a photo-editing app. The reversal is a removed affordance: a shared channel for flagging provider releases that quietly died once nobody was assigned to check it.
Check yourself Score: 0 / 0
True or false
1. True or false: this answer recommends reading and reacting to every model provider announcement as it happens.
True
False
Show hint
Look at the "keep out" step and "not day one."
Show answer
False. It recommends scoring announcements and routing only the big, near-surface ones to a real review, logging the rest to a watchlist.
Fill in the blank
2. Fill in the blank: the anchor of this process is a score built on capability delta and ___.
Show hint
Look at the direct answer and the "A" step.
Show answer
Time-to-impact. How soon a given capability change could actually touch the team's own product surface.
Short answer, name the reversal
3. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Cutting the weekly watch document because it felt redundant next to a strong internal eval suite. That made sense while the eval suite was catching every quality issue that mattered, and stopped making sense once the real risk was a competitive shift the eval suite was never built to catch.
Multiple choice
4. According to this answer, what should happen to an announcement with a small capability delta and a distant time-to-impact?
A. It should trigger a same-week team review.
B. It should be escalated to leadership immediately.
C. It should be logged to a watchlist that nobody has to act on yet.
D. It should be ignored completely with no record kept.
Show hint
Look at the decision tree diagram and the priority list.
Show answer
C. Small-and-distant items get logged, not ignored outright and not escalated, keeping the team's attention for the ones that actually matter.
Short answer, apply it yourself
5. Think of a fast-moving field you follow personally. What's one simple score you could use to decide what's worth your attention versus what to skip?
Show hint
Think about two things: how big the change actually is, and how soon it could affect something you actually do.
Show answer
Model answer: Something like "how much does this actually change, and how soon could it affect me" works in almost any fast-moving field, not just AI model releases.
Short answer, where it wouldn't matter
6. Name a part of Ember Desk's work where this tracking process genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The internal eval suite catching quality regressions. That job was never about competitive shifts, and it didn't need to change to accommodate this new process.
Before you close the answer
Why this works
Tests whether you can design a sustainable, low-overhead process for a fast-moving input, instead of either ignoring it or building something too heavy for a small team to keep running.
Follow-up traps
"What if the named owner is out sick the week something big drops?" Response: the score and the watchlist are shared artifacts, not something locked in one person's head, so anyone on the team can triage from the same two axes if the owner is out.
"Isn't a two-axis score too simple to capture something this complex?" Response: it's simple on purpose; a complex scoring model is exactly the kind of thing a busy team stops maintaining, and a lean team is better served by a simple score run reliably than a rich one run occasionally.
If pressed
Ember Desk's eventual scoring rubric used a 1 to 3 scale on each axis, with anything scoring 5 or 6 combined routed automatically to a same-day Slack escalation, and anything scoring 2 or below logged with no action required for at least a month.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.