How would you evaluate an AI opportunity in a market where every competitor already has the feature?
Wrenfast reads a job seeker's resume, reads an open posting, and gives both sides a percentage: how well they fit. Yalin Kovacs owns that matching model. Truett Beckman runs partner growth, and watched two rivals, Talenza and Hirevane, ship the same score with a fancier chart around it. He wanted the whole next quarter's engineering time to catch up on looks. Yalin had to work out whether a fancier score was actually worth fighting for, or whether the real fight was somewhere else entirely.
- Build the match score to parity, and no further.Why: this is the whole call, and skipping it means the rest of the plan never gets funded.
- Find the data, workflow, or relationship a rival structurally cannot copy.Why: that's the only thing left that a sprint on the other side can't erase in a quarter.
- Spend the real budget on that thing, not on the score everyone already has.Why: the quarter is the scarce resource, and it can only fund one of these two fights.
- Reject any shortcut that fakes the differentiating data instead of collecting it.Why: a guessed label dressed up as a real signal teaches the model nothing it didn't already half-know.
- Check the new signal for bias by segment before it ever reaches a live score.Why: the data behind a real edge is usually messier and less checked than the data behind a copied feature.
- Leave the parts that are already commodity-grade alone.Why: more polish on a solved problem buys nothing measurable, and it's time stolen from the one thing that would.
How to answer this, stage by stage
Nobody in the room is grading whether you'd build the feature. Everyone already has it. They're grading whether you can find the one thing left to actually fight over, and prove it with a number.
Let's learn
Wrenfast is a tool that reads a resume, reads a job posting, and gives both sides a number for how well they fit.
Three years ago, when Wrenfast launched, a match score was new. It cut a recruiter's screening time on a posting from about forty minutes to twelve, and no other tool in the market showed a number like it at all. That gap alone won Wrenfast its first wave of employer partners.
It doesn't work like that anymore. Two rivals, Talenza and Hirevane, now ship the same score. Checked against the same labeled test set, Wrenfast lands at seventy one percent, Talenza at seventy four, Hirevane at sixty nine. Screening time is still twelve minutes, everywhere. The gap that won those early partners closed a while back, and nobody at Wrenfast marked the date it happened.
Here's the turn. The extra sales objections were never really the problem worth solving. Five of Wrenfast's last nine employer renewal calls were lost, and every one of those five mentioned the score looking plainer than Talenza's. That reads like a feature gap. It's actually a spend-the-quarter problem: fund a full visual and precision chase, and there's nothing left over for the one thing only Wrenfast can build.
What Wrenfast had, and never used, was something none of its rivals could touch: two hundred and ten employer partners sending back a ninety-day survey after every hire, retention and a supervisor rating included. It existed for account-management reports. It had never once been fed back into the matching model.
What I would leave alone: the resume-parsing layer, the part that pulls titles, dates, and skills out of a PDF, doesn't need more investment. Every vendor's is roughly the same quality now. Sharper parsing wouldn't move a single retention number.
The lesson: a model that looks identical to every rival's isn't a differentiator, however well it's tuned. The real edge is a data source nobody else can reach, not a sharper number on the one everyone already has.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for what it cost Wrenfast to see it the slow way.
The Wrenfast office runs quiet most evenings, except the Tuesday before a quarterly planning review, when Yalin Kovacs is still at her desk with three browser tabs open, each one a rival's product page.
She'd owned the matching model for two of Wrenfast's three years, long enough to have rebuilt the resume parser twice and to know, without checking, which employer partners cared most about speed versus which cared most about the number itself.
For most of that time, the match score was simply the thing Wrenfast had and nobody else did. Employer partners signed on because of it. Job seekers trusted the percentage because there was nothing else to compare it to. Then, over about a year, quietly, both Talenza and Hirevane shipped the same idea. Nobody at Wrenfast marked the week it happened. There wasn't a single announcement to react to, just a slow sense that the number wasn't special anymore.
Then Talenza shipped a skill radar chart around its score, all animated bars and color. Two weeks later, Truett Beckman, who ran partner growth, brought Yalin a stack of renewal call notes. Five of the last nine had gone to Talenza or stayed only after a discount, and four of those five mentions came back to the same line: "their score just looks more thorough." Truett's ask was simple. Give him the whole next quarter, twelve weeks of engineering, to rebuild the visualization and push the score's apparent precision higher with more parsed signals: certifications, inferred soft skills, anything that would move the number and the chart around it.
Yalin almost said yes. It was an easy yes to reach for, a clear ask tied to a real, painful number.
What stopped her was running the actual test before signing off on the ask. Position first: this isn't "should we improve the score," it's "is there a real edge left in the score, or is it table stakes now, and the real question is somewhere else." She pulled the accuracy numbers against the shared test set: seventy one, seventy four, sixty nine. A two to five point spread. Nobody renews or leaves an account over that.
An engineer on her team floated a fourth shortcut mid-meeting: skip waiting on real employer survey data and generate synthetic post-hire performance labels with a second model instead, guessing at likely job success straight from resume text. Yalin turned it down on the spot. It doesn't add a new signal, it just recycles the same resume words the original score already used, dressed up as if it were new evidence. Training one model's guess into another model's ground truth compounds the error instead of removing it.
Two hundred and ten employer partners were already sending Wrenfast a quarterly survey: was the hire still there at ninety days, and how did their supervisor rate the fit. It existed purely for account management, a page in a report nobody outside partnerships ever opened. It had never been fed back into the matching model itself. Talenza and Hirevane are resume aggregators. They don't hold employer relationships that reach past the point of hire, so they structurally cannot get this signal, not this quarter, not next year either, without building three years of the same trust Wrenfast already had.
Yalin brought Truett a split instead of his full ask: three weeks to bring the score display to a competent parity, close enough that nobody loses a renewal over a chart, and nine weeks to wire the ninety-day outcome data into the model as a real training signal.
The trade-off she named out loud in that meeting: the outcome signal only exists ninety days after a placement, so any brand-new job category or partner starts with nothing but the old resume-similarity score until enough real outcomes pile up. She set that floor at forty completed placements with survey data before the outcome signal blends in at all. Slower to warm up, but a real lift once it's live, instead of a live number built on no evidence.
There was one more thing to watch, and she said so before anyone asked. A supervisor's rating can carry its own bias, reflecting how well someone fit one manager's style rather than whether they were actually good at the job. Left unchecked, that could quietly skew the retrained score against candidates from less traditional backgrounds who a harsher early rater happened to mark down. Before anything shipped, the retrained model got checked against a held-out set, sliced by education path and prior industry, not just judged on its average.
Here's the replay. Same twelve-week quarter, same rival pressure, same lost renewals on the table. This time three weeks go to the reskin, and the other nine to the outcome pipeline. By the end of the first retraining cycle, on a controlled split of twelve hundred placements, candidates matched using the outcome-informed score were still employed at six months seventy nine percent of the time. The next cycle brought it to eighty one, against sixty eight for the old resume-only score. That gap didn't exist as a number three months earlier. It existed as an unread page in an account report.
What I'd tell myself, watching that survey page get built for reporting and never once for the model: the data that makes a real difference in a crowded market is rarely the shiny thing a rival just shipped. It's usually the boring thing you already collect and never finished using.
PICK, for the feature everyone already ships
Not a rule about always chasing what's new. PICK earns its keep here only when it separates the feature that's already a tie from the one thing left that isn't.
The trade worth saying out loud: the outcome signal is slow on purpose. It only exists ninety days after a placement, so new job categories run on the old resume-only score until forty completed placements pile up behind them. That's a real quality-versus-speed trade, accepted deliberately, because a quarter-old signal that's actually true beats an instant one that's just recycled words from the same resume the score already read.
And if you want to be sure it really works, try it somewhere else
Same four letters, a veterinary practice group instead of a job board, and the missing check is a diagnosis outcome, not a ninety-day survey.
Pallisward Veterinary Group runs an AI intake bot across its clinics: a pet owner describes symptoms, and the bot flags how urgent the visit likely is before a vet ever sees the case. Every major practice-management vendor now ships some version of this triage bot, trained once on public symptom checklists and left alone. Pallisward's leadership wanted a full rebuild to make theirs look sharper: more symptom categories, a friendlier chat voice, a cleaner urgency badge.
Position: the question isn't whether to have a triage bot, everyone already does. It's whether Pallisward has a way to make triage genuinely more accurate than a checklist, or whether triage is now table stakes and the real fight is elsewhere. Impact: pour a quarter into voice and badge design, and you're polishing a bot that already performs about as well as every rival's, on the same public checklist data. Treat the real opportunity as table stakes instead, and you miss it: Pallisward's vets already record, for every flagged case, what the real diagnosis turned out to be. Cost asymmetry: a competent badge-and-voice refresh costs a bounded few weeks. Never wiring real diagnosis outcomes back into the triage model costs the one thing that would let it actually get better at telling a true emergency from a worried owner, quarter after quarter, while every rival's bot stays frozen on the same public checklist forever. Kill criteria: could a rival copy this without years of real vet-confirmed diagnosis records tied to each flagged case? A generic triage vendor selling to a thousand unrelated clinics can't build that. Pallisward already has it, sitting in visit notes nobody had linked back to the bot.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the one test, a data source or relationship a rival structurally can't reach, before anything else.
Cost: no budget for a full rebuild before a real deadline. Fine, but fund the parity fix first and smallest, and only chase the real edge with whatever's left.
The model got better, for real: say a rival's version scores near-identical to yours on every public benchmark. Still don't chase the benchmark further. A tied score on a public benchmark says nothing about who holds the private data that would actually move it.
Where people run it wrong.
They spend the whole budget matching a rival's polish because it's the visible, easy-to-defend ask, and never test whether polish was ever the real fight.
They assume any data they happen to already collect is automatically a moat, without checking whether a rival could get the same thing another way.
They chase the differentiation angle first and skip parity entirely, then lose deals on the boring gap a competent baseline would have closed.
How to use it live. If you're ever asked how to compete on a feature everyone already ships, buy yourself a second with one plain question said out loud: "what do we hold that they structurally can't get to." That question is the whole method, asked instead of stated.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't waiting ninety days for outcome data too slow to ever compete on?" Response: the model doesn't wait idle. New categories run on the old resume-only score until forty completed placements accumulate, so coverage stays instant while the real signal blends in behind it, retrained on a quarter-old cycle instead of never.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Opportunity identification for AI
- #1 What characteristics make a workflow a good candidate for AI? List five.
- #2 Describe a method for finding AI opportunities inside an existing product without starting from the technology.
- #3 How do you distinguish a problem AI solves from a problem AI merely touches?
- #4 Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.
- #5 Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
- #6 Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.