ConceptAdvancedAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #15
When does building become a hiring decision rather than an engineering decision?
LEADthe leading signal was a calendar. the lagging one was a missed roadmap item nobody could ignore
Fernbridge Support builds customer-support software. TicketSense drafts a reply to an incoming support ticket for an agent to review and send. Callum Streeter is the AI PM who owns it, and Reya Kessler is the senior engineer whose week quietly filled up with work nobody had ever assigned to her.
The direct answer
Building becomes a hiring decision the moment a senior engineer spends more time keeping the AI feature accurate, rewriting prompts, adding eval cases, checking for regressions after every model update, than building anything new. Track that share of their week on purpose. Once it crosses about a quarter, sustained for a full quarter, stop calling it overflow work and open a dedicated AI hire, because the organization is already paying for that role in lost velocity, just without budgeting or protecting it.
Do this, in order
Track the share of a senior engineer's week spent on AI upkeep, not just their ticket count.Why: this is the number that moves weeks before anyone notices the roadmap slipping.
Set a real threshold in advance, not a feeling you'll notice eventually.Why: without a number, "too much" only gets named after someone burns out or quits.
Watch for the metric being gamed by simply doing less upkeep.Why: a falling percentage can mean the org fixed the problem, or it can mean eval coverage quietly lapsed instead.
Open the hire once the threshold holds for a full quarter, not after one bad week.Why: a single busy sprint is noise. A quarter of it sustained is a real, structural gap.
Name what the new hire actually owns, not just "help with AI."Why: a vague mandate just recreates the same undefined second job with a new title.
Revisit the threshold as the feature matures.Why: a stable, well-evaluated feature needs less ongoing upkeep than one still finding its shape.
How to answer this, stage by stage
Nobody is scoring whether you can name a headcount number. They're scoring whether you'd notice the shift before it shows up as a missed roadmap item.
Stage 1
Scope it to one real team, with a real name in it
Say it like this
"Let's ground this. Fernbridge Support built TicketSense, a reply-drafting tool for support agents. Callum owns the roadmap. Reya is the senior engineer who ended up doing most of the ongoing AI upkeep, without anyone ever deciding that was her job."
Why this works
Keeps "when does this become a hiring decision" from turning into an abstract org-design lecture.
Stage 2
Say your structure out loud before any numbers
Say it like this
"I'll run this as LEAD. Link, the real business outcome. Early signal, what moves before the problem is obvious. Abuse, how the metric could get gamed. Decision, what I'd actually do at each threshold."
Why this works
Signals a repeatable staffing signal, not a gut feeling about when a team "seems stretched."
Stage 3
Reframe the question: this isn't about workload, it's about ownership
Say it like this
"This isn't really asking 'when is the team too busy.' It's asking when an unofficial job, kept alive by one person's goodwill, needs to become an official one with its own headcount and its own protection from getting deprioritized."
Why this works
This is where the answer separates from a generic "hire when you're understaffed" cliche.
Stage 4
Give the one signal: the leading indicator, and its threshold
Say it like this
"Here's the number I'd track. What share of a senior engineer's week goes to AI upkeep, prompt fixes, eval cases, regression checks, instead of new feature work? Under fifteen percent, that's normal. Between fifteen and thirty, give someone a formal, protected part of their role for it. Past thirty percent, sustained for a full quarter, open a dedicated hire."
Why this works
This is the direct answer, stated as a real, trackable number instead of a vague sense of "too much."
Stage 5
Prove it with the compressed evidence
Say it like this
"Reya's share went from eight percent in month one, to nineteen in month three, nobody noticed, to thirty four percent by month five. Nobody flagged it as a staffing problem until a new hire, during onboarding, asked who actually owned checking TicketSense's replies after a model update, and there wasn't a clean answer."
Why this works
Gives the interviewer a real, climbing number and a real, small moment that finally surfaced it.
Stage 6
Name the AI-specific reasoning and the trade-off being accepted
Say it like this
"The honest reason this keeps happening is that a model isn't a one-time build, it drifts, it gets updated underneath you, and every version bump can quietly change how TicketSense answers. That ongoing evaluation work doesn't show up as a feature ticket, so it's invisible on a roadmap even while it's eating a third of someone's week. Hiring for it means accepting slower short-term feature output in exchange for a team that isn't one person's goodwill away from silently regressing."
Why this works
This is the load-bearing judgment. It only makes sense because a model's behavior can shift after every update, not because of any generic staffing principle.
Stage 7
Say where the threshold doesn't apply, then close on one line
Say it like this
"I wouldn't apply this to a small, stable internal tool nobody depends on daily, the upkeep cost there is genuinely low and occasional. For TicketSense, the answer holds: once a senior engineer's AI upkeep crosses thirty percent of their week for a full quarter, that's not overflow anymore, that's a role the org hasn't hired for yet."
Why this works
Closes with real judgment about where the threshold doesn't apply, and restates the direct answer in one breath.
Let's learn
TicketSense reads an incoming support ticket and drafts a reply, which a support agent reviews, edits if needed, and sends, instead of writing the reply entirely from scratch.
None of these were ever written into anyone's job description. All four quietly became Reya's.
When TicketSense first shipped, Reya spent about 8 percent of her week on this kind of upkeep, small, occasional, and easy to absorb alongside her normal feature work.
Share of Reya's week spent on AI upkeep, month by month
Share of week on AI upkeep
This number crossed the threshold in month five. Nobody was watching it until a new hire asked a question in month five's onboarding.
By month three, it had crept to 19 percent, still under most people's radar, since Reya kept shipping her normal feature work alongside it, just at a slower pace nobody had connected to a specific cause yet.
Knowledge spark: why does a shipped model need ongoing upkeep at all?
A model isn't a finished piece of code, it's a behavior that can shift underneath a feature. A version update, a change in the kind of tickets coming in, or a prompt tweak elsewhere in the system can all quietly change how TicketSense answers. Someone has to keep checking that it still behaves the way it did on launch day.
Fernbridge had been living in the fourth branch for months without ever officially choosing it.
Nobody decided to give Reya a second job. It just accumulated, one reasonable fix at a time.
Here's the turn: the real cost was never the hours themselves. It was that TicketSense's roadmap, three planned features, quietly slipped a full sprint, and leadership initially read that as Reya being slower, not as a structural gap in who owned ongoing AI quality.
Where Reya's engineering hours actually went, month by month
New feature workAI upkeep
Same 40-hour week, every month. The slice going to upkeep nearly quintupled while nobody adjusted what else was expected of her.
The choice I would take back
Treating "ship TicketSense" and "keep TicketSense accurate forever" as one undifferentiated engineering task, assigned by default to whoever built the feature first, with no natural point to reassess whether that second part needed its own headcount. That made sense at launch, when upkeep was genuinely light. It stopped making sense once model updates and ticket volume both grew and nobody ever separated the two jobs.
What I would leave alone: Fernbridge's internal expense-approval tool, a small rules-based helper with no model in it at all, needs none of this thinking. There's no drift to watch, so there's no ongoing AI-upkeep job to staff for.
The lesson: a hiring gap doesn't usually announce itself as a hiring gap. It shows up first as a senior engineer's calendar quietly filling with work nobody assigned, and only later as a missed roadmap item everyone notices and misdiagnoses.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like watching a real job form with no name and no owner.
Callum Streeter had shipped TicketSense to genuine relief from the support team: average reply time down, agents happier, a clean launch by most measures. Reya Kessler, the senior engineer who'd built most of it, kept a loose eye on it afterward, the way anyone keeps an eye on something they built.
Nothing about this looked like a crisis at any single point. It only reads as one laid end to end.
A model version update in month two changed how TicketSense phrased refund explanations, subtly enough that nobody flagged it for a week, until an agent noticed replies sounding oddly clipped. Reya fixed the prompt. Nobody assigned her to; she just did, because she understood the system best.
By month four, this had become a pattern: check after each update, patch the eval set when a new ticket type appeared, quietly handle the "why did it say that" questions that came her way instead of a support agent's. None of it was written down as her job. All of it was, in practice.
The fix was never asking Reya to work harder. It was giving the second job a name, and someone whose actual job it was.
The moment that surfaced it: a new engineer, in her second week, asked in a team meeting who was responsible for checking whether TicketSense's replies were still accurate after a model update. There wasn't a clean answer. Reya said, half-joking, "I guess that's me," and the room went quiet in a way that made it obvious nobody had actually decided that.
The new hire's question wasn't really about a process gap. It was about a job that existed and had no name.
Callum never had a fixed rule for exactly when "someone doing extra AI upkeep" became "a role the team is missing." It came down to a feeling with two settings: either the upkeep was occasional enough that anyone could absorb it without noticing, or it had become a real, recurring share of someone's week that was quietly costing the roadmap elsewhere. Thirty four percent, discovered only because a new hire asked a plain question, was unmistakably the second setting.
Once named, the job turned out to have four real parts, not one vague catch-all.
Back when TicketSense first launched, letting Reya informally own its ongoing quality wasn't an unreasonable call, the upkeep really was light at first, and there was no reason yet to think it would grow. It stopped being reasonable the moment model updates and ticket volume both climbed and nobody ever separated "built the feature" from "keeps it accurate" as two distinct, differently-staffed jobs.
The time cost climbed for five months before the ownership axis ever moved to match it.
Here's the replay: Fernbridge opened a dedicated applied-AI engineer role that quarter, with eval pipeline upkeep, drift monitoring, and version migration explicitly named as the job, not implied. Reya's AI-upkeep share dropped back to a manageable 10 percent within two months, spent mostly reviewing rather than doing. The next model update landed without eating into anyone's roadmap at all.
One version of this story lets a senior engineer's week keep quietly filling until she burns out or the roadmap slips badly enough to force an uncomfortable postmortem. The other tracks the actual number, names the threshold in advance, and turns an invisible second job into a real, protected first one.
What I'd tell myself, hearing that new hire's question land in the room: the org had already been paying for this role in lost velocity for months. The only thing missing was a name for it, and a number that would have made the case months earlier.
LEAD, run on a job that existed before anyone named itNot a script for hiring the moment anyone looks busy. LEAD is what tells you a role is missing before the roadmap slips prove it.
L
Link. What's the real business outcome underneath this question?
Shipping an AI feature that keeps working reliably as models update and ticket patterns shift, without quietly consuming a senior engineer's entire week to do it.
Without naming this, "is the team busy" gets asked instead of "is quality being protected by one person's goodwill."
E
Early signal. What moves weeks before the roadmap visibly slips?
The share of a senior engineer's week spent on AI upkeep, prompt fixes, eval cases, regression checks, tracked on purpose, not left to a feeling. It climbed from 8 to 34 percent over five months before anyone connected it to the roadmap slipping.
This is the answer to the actual question: the leading signal is a calendar, not a headcount request that hasn't been written yet.
A
Abuse. How could this metric get gamed?
The percentage can fall simply because upkeep gets skipped, evals stop getting updated, regressions stop getting checked, which looks healthy on the metric while quality quietly degrades underneath it.
A falling number isn't automatically good news; it has to be paired with a real quality check, not read alone.
D
Decision. What do you actually do at each threshold?
Under 15 percent, no change. 15 to 30 percent, give one engineer a formal, protected part-time role. Past 30 percent for a full quarter, open a dedicated applied-AI hire with named responsibilities.
A metric nobody acts on is decoration. This is what makes the number worth tracking in the first place.
The recap, one line per letter: link is protecting reliable AI quality without burning out one engineer, early signal is the share of a senior engineer's week on AI upkeep, tracked before the roadmap notices, abuse is the risk of that number falling because quality work quietly stopped instead of got resolved, and decision is the concrete threshold, roughly 30 percent for a full quarter, that actually triggers a real hire.
And if you want to be sure it really works, try it somewhere elseSame four letters, a legal-document startup instead of a support-software company. This time the early signal shows up in a different kind of hour.
Idris Fontaine is the AI PM at Caswell Legal Docs, which specifies ClauseCheck, a feature that flags risky clauses in a contract draft for a paralegal to review. Mapped onto LEAD: link is a flagging tool lawyers actually trust, not one they've learned to double-check entirely by hand. Early signal here isn't a senior engineer's calendar, it's the share of flagged clauses a reviewing paralegal overrides, climbing from 6 percent to 28 percent over four months as contract types diversified beyond what ClauseCheck was originally tuned on. Abuse is that the override rate could look artificially low if paralegals simply stopped bothering to override and started rubber-stamping the tool's flags instead, which would hide a real accuracy problem behind a falling number. Decision is similar in shape but different in trigger: past a 25 percent override rate sustained for a full quarter, Caswell opens a dedicated legal-AI engineer role, since a generalist engineer patching individual clause types one at a time can no longer keep pace with how fast contract variety is growing.
Same threshold logic, a different leading number. The shape of the decision travels even when the metric itself changes.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "track the share of a senior engineer's week on AI upkeep, hire once it holds above thirty percent for a full quarter," and stop.
Cost: no budget to open a new headcount line yet. Say so honestly, and commit to formally protecting part of one engineer's role as the interim step, rather than letting the informal arrangement continue unnamed.
The model got better, for real: say a new model version needs almost no prompt adjustment at all. The upkeep share would drop on its own, which is real good news here, not a sign the metric is being gamed, since nothing about the underlying eval or monitoring work stopped happening.
Where people run it wrong.
They wait for a missed roadmap deadline to notice the staffing gap, instead of tracking the leading number that predicted it months earlier.
They read a rising workload as one person needing to manage their time better, instead of as a role the organization hasn't hired for yet.
They hire for "AI support" with no named responsibilities, which just recreates the same undefined second job under a new title.
How to use it live. The moment an interviewer asks when building becomes a hiring decision, don't reach for a headcount number. Ask yourself: what's the one number that would show this happening months before a roadmap slips? That's usually the real answer hiding in the question.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits deciding when building becomes a hiring decision?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. This is a metric-shaped question, applied to a staffing signal instead of a product one, so a FLIPS flip family doesn't apply here.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Callum Streeter, the AI PM who owns TicketSense at Fernbridge Support. Reya Kessler is the senior engineer whose week quietly absorbs its ongoing AI upkeep with no formal role behind it.
3 · THE EARLY SIGNAL
What's the leading indicator that a hire is needed, before the roadmap shows it?
Tap to flip
ANSWER
The share of a senior engineer's week spent on AI upkeep, prompt fixes, eval cases, regression checks, instead of new feature work. It climbed for months before anyone connected it to a slipping roadmap.
4 · THE THRESHOLD
At what point does this stop being overflow work and become a hiring decision?
Tap to flip
ANSWER
Roughly 30 percent of a senior engineer's week, sustained for a full quarter. Below 15 percent, no change. Between 15 and 30, a formal part-time role first.
5 · THE OLD DECISION
What old decision does this answer take back?
Tap to flip
ANSWER
Treating "ship the feature" and "keep it accurate forever" as one undifferentiated job assigned to whoever built it. Reasonable when upkeep was light at launch. Wrong once it grew and nobody ever separated the two jobs.
6 · THE NUMBER
Fill in the blank: Reya's AI-upkeep share went from ___ percent in month 1 to ___ percent by month 5.
Tap to flip
ANSWER
8 percent to 34 percent. It crossed the 30 percent threshold by month five, well past the point a formal role should have existed.
7 · THE REPLAY
Same climb, but the threshold gets acted on at month three instead of month five. What changes?
Tap to flip
ANSWER
A formal part-time role gets created two months earlier, the roadmap never visibly slips, and the new hire's onboarding question in month five has an easy, already-true answer instead of an awkward silence.
8 · CROSS PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the leading signal there instead?
Tap to flip
ANSWER
Caswell Legal Docs' ClauseCheck. There, the leading signal is a paralegal's override rate on flagged clauses, not an engineer's calendar, but the same threshold logic decides when a dedicated hire is needed.
Check yourself Score: 0 / 0
Short answer, name the signal
1. What's the leading indicator this answer tracks, and why does it move before the roadmap does?
Show hint
Look at the line chart in "Let's learn."
Show answer
Model answer: The share of Reya's week spent on AI upkeep. It climbs quietly month over month, well before it becomes visible as a missed feature on the roadmap, which is a lagging, not a leading, sign.
Multiple choice
2. Why is a falling AI-upkeep percentage not automatically good news?
A. It always means the model has stopped needing updates entirely.
B. It can mean upkeep work, like eval checks, is quietly being skipped rather than resolved.
C. It means the feature has been discontinued.
D. It only happens after a dedicated hire joins.
Show hint
Look at the "abuse" step in the LEAD recap.
Show answer
B. The percentage can fall because quality work was genuinely resolved, or because someone quietly stopped doing it, and only a real quality check tells the two apart.
True or false
3. True or false: leadership initially diagnosed the slipping roadmap as Reya working too slowly, rather than as a staffing gap.
True
False
Show hint
Look at "the turn" paragraph in "Let's learn."
Show answer
True. The slipping roadmap was first read as a performance issue, not as evidence that AI upkeep had quietly become a real, unstaffed job.
Short answer, where it wouldn't matter
4. Name a feature at Fernbridge where this hiring threshold would NOT apply, and say why not.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The internal expense-approval tool, a rules-based helper with no model in it. There's no drift to monitor, so there's no ongoing AI-upkeep job to staff for in the first place.
Short answer, apply it yourself
5. Think of a team you've been on where one person quietly absorbed extra, unassigned work. What number would have shown that happening earlier?
Show hint
Think about a role where "helping out" slowly became "doing it every time" with nobody deciding that on purpose.
Show answer
Model answer: On a marketing team, one analyst quietly became the only person who could pull a certain report. Tracking how often others requested it from her, rather than pulling it themselves, would have shown the dependency forming months before she became a bottleneck.
Short answer, work the number
6. If Reya's AI-upkeep share had plateaued at 22 percent instead of continuing to 34, would opening a full-time dedicated hire still make sense?
Show hint
Compare 22 percent against the 15-to-30 percent band in the decision step.
Show answer
Model answer: Not yet, at 22 percent sustained, a formal part-time role for one engineer is the better first move, saving a full dedicated hire for if the number kept climbing past 30 percent afterward.
Before you close the answer
Why this works
Tests whether you'll track a real leading signal for a staffing decision instead of waiting for a missed deadline, and whether you understand that ongoing AI quality is a distinct, recurring job, not a one-time build task.
Follow-up traps
"Couldn't Reya have just said no to the extra work?" Response: each individual request was small and reasonable in isolation, the problem was that nobody was tracking the cumulative total, not that any single ask was unfair.
"Isn't a dedicated hire an overreaction to one person's workload?" Response: it's not about one person, it's about a recurring job that will land on whoever's most senior engineer next, unless it gets a name and a budget of its own.
If pressed
The new applied-AI hire's eval pipeline work was scoped to run automatically after every model version bump, rather than being triggered manually, specifically so the next drift wouldn't depend on someone remembering to check.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.