ConceptIntermediateShipping & Model Lifecycle / Pilot design and POC-to-production / #15

What support model does a pilot need?

The direct answer
Size a pilot's support hours from arithmetic, not from a promise: how many issues each pilot user will flag a week, times how long one actually takes to fix, plus a fixed floor of hours someone has to stay reachable to keep whatever response time you promised. Check that total against how many hours two real people can sustain for the whole pilot, not just the first curious week, and go in knowing the response promise moves the number more than almost anything else. "Someone will handle it" is not a support model. It is a guess wearing a plan's clothes.
Do this, in order
  1. Size support hours from real arithmetic: issues per pilot user each week, times minutes to fix one, plus a coverage floor for the promised response time.Why: "someone will handle it" isn't a plan, it's a hope, and hope runs out around week five.
  2. Set the response-time promise before you set the headcount, because the promise decides the coverage floor.Why: a one-hour promise needs someone reachable nearly all day; a four-hour promise needs a few planned check-ins, same issue count, very different staffing.
  3. Give the estimate a range, not one number, based on how new the feature feels to the pilot users.Why: a familiar AI suggestion gets a shrug and a quick check; a first-time AI call on someone's job gets questioned every time.
  4. Check the total against how many hours two real people can sustain for the whole pilot window, not just week one.Why: a plan that only survives the first curious week isn't a support plan, it's a honeymoon.
  5. Know that tightening the response promise moves the total more than how new the feature feels.Why: cut the wrong assumption and either the promise breaks or the two people covering it burn out.

How to answer this, stage by stage

Nobody is grading whether you land on exactly 20. They're grading whether you name the real build-up before touching a number, whether the range comes from something real, and whether you can say which assumption would move the total most. Six moves get you there.

1
Scope it to one real pilot, not a category of pilots
Say it like this
"Let's ground this. Say I'm running a pilot of an AI return-reason classifier with Windale & Co, an apparel retailer. Fifteen of their returns agents are using it for eight weeks, and it suggests why a package is coming back before they have to type it in by hand."
Why this works
A named retailer and a named headcount give the interviewer something to push on, instead of a shrug about "pilots in general."
2
Reframe the question before naming a headcount
Say it like this
"This isn't asking me to name a support tier off the top of my head. It's asking how many real hours a support plan needs, and who's actually free to cover them, before I promise anyone a response time."
Why this works
Rules out "we'll have someone on Slack" before it becomes the answer. That instinct sounds reassuring and prices nothing.
3
State the build-up before touching a number
Say it like this
"Weekly support hours equal two things added together. One, the time it takes to actually resolve every issue an agent flags. Two, a fixed floor of hours someone has to stay reachable, just to keep whatever response time we promised, even on a quiet week."
Why this works
States the equation before a single figure lands, so what follows reads as arithmetic, not a guess dressed up as one.
4
Own the numbers
Say it like this
"Fifteen agents, about three flagged issues each a week, that's 45 issues. Twenty minutes to actually resolve one, reading the flag, checking the model's reasoning, replying, that's 900 minutes, 15 hours. We promised a reply inside four business hours, so someone stays reachable roughly an hour a day, five hours a week, whether or not anything's actually broken. Fifteen plus five is 20 hours a week."
Why this works
Turns "we'll keep an eye on it" into a number a reviewer can check against a real, named build-up.
5
Give the range, and check it against real hours
Say it like this
"That 20 assumes the feature feels moderately new to them. If they already trusted a similar AI suggestion somewhere else on the same screen, it drops closer to 12 and a half. If this is the first time an AI's made this call on their job at all, it climbs to about 30. Either way, split across two of our support engineers, that's ten to fifteen hours each a week, on top of the other accounts they cover. Sustainable at 20. Tight at 30."
Why this works
Ties the range to something real, how new the feature feels, and checks it against two actual people's actual weeks, not an abstract confidence interval.
6
Name the lever, then close in one breath
Say it like this
"If I had to bet on what moves this number most, it's the response-time promise, not how new the feature feels. Tighten it from four hours to one, and the coverage floor alone jumps from five hours a week to something close to thirty, because now someone has to be reachable almost the whole day, not just checking in a few times. So: 20 hours a week at a four-hour promise, up near 45 at a one-hour promise, and the promise is the first thing I'd interrogate before assuming we just need more people."
Why this works
Answers the hardest follow-up directly and closes in one breath, the way a strong answer actually sounds.
If you remember one thing A support plan that only says "someone will handle it" will always sound fine in a kickoff meeting. Size it as two added parts, resolution hours plus a coverage floor, and check the total against how many hours two real people can actually sustain for the whole pilot.

Let's learn

What actually breaks a pilot's support plan? Not a bad model. Bad math about how many hours answering questions really costs.

The product here is small: a button that reads a returned item's photo and the customer's note, and suggests why it's coming back, wrong size, damaged, not as described, before a returns agent has to type it in themselves.

Before the pilot, Windale & Co's returns agents picked the reason themselves, one by one, from a dropdown that didn't always fit. It worked. It was just slow, and it wasn't consistent from one agent to the next.

During the pilot, fifteen agents got the classifier's suggestion first, and mostly it saved real time. But agents flagged the odd one, a shirt the model called "not as described" that was really just the wrong size, or a plain question about why it picked what it picked.

Here is the turn. Those flagged questions were never the real problem. The real problem is that nobody had worked out how many hours answering them would actually take, or who was supposed to be free to do it.

The build-up: two parts, one weekly total
15 hours resolving 45 flagged issues (15 agents, 3 each, 20 min apiece)15 hrs
+ 5 hour coverage floor, to keep the 4-hour response promise20 hrs
The floor looks like the smaller number, but it's the part a normal support inbox never bothers to size. It exists whether the queue is busy or dead quiet, because the promise doesn't care which one it is.
We didn't understaff a chat channel. We understaffed a promise.

At its worst, that costs the whole pilot's credibility. Eight weeks in, a support plan can look fine on a kickoff slide and still leave agents waiting a full day for an answer, because nobody ever turned "we'll keep an eye on it" into a number of hours.

The decision that mattered Size support hours from real numbers: issues per pilot agent each week, times minutes to fix one, plus a fixed floor for whatever response time you promised. "Someone will handle it" isn't a number.
Knowledge spark: what's a coverage floor? The minimum hours someone has to stay reachable to keep a promised response time, even on a day nothing goes wrong. It doesn't shrink just because the queue happens to be quiet that week.

The choice I would take back. The pilot plan promised Windale a reply within four business hours. Nobody worked out what that promise actually costs to keep, in real hours, before agreeing to it in the kickoff meeting.

What I would leave alone. Not every pilot needs this. A pilot testing a one-line copy change on a button doesn't need a coverage floor at all, nobody's job depends on it being right within four hours. Save the real hours for changes that touch a call an agent gets graded on.

The lesson. A support model isn't a sentence in a kickoff deck. It's a staffing plan with real hours attached, and it has to survive the whole pilot window, not just the first curious week.

Now here is the same thing as a story

Skip this if you already believe a pilot's support plan needs real hours attached to it, not just a name on a Slack channel. Read on if you want to feel why.

Every Monday morning, Nayeli Cortes opened the same shared Slack channel before she'd finished her coffee, checking what had piled up over the weekend.

Nayeli's a product manager at Thornfield AI, and Windale & Co's returns floor was her first pilot with real customers on the other end of every flagged item. Fifteen agents, eight weeks, watching whether her team's return-reason classifier actually held up outside a demo.

For the first two weeks, that channel was the best part of her mornings. She checked it every hour or so. The questions were mostly curious, "why'd it pick this one," "is this a bug or is it right," and she'd answer in minutes. Agents warmed to the tool fast. It was doing exactly what it was supposed to.

So she checked it a little less. Week three, twice a day. Week four, once, usually at the end of hers. By weeks five and six, whenever she had a minute, because nothing bad had happened yet, and she had two other pilots pulling at the same hours.

Nobody told her to slow down. She just found, somewhere in that stretch, that she wasn't the only thing standing between an agent's question and an answer, except that she was, and she'd stopped noticing.

Then, in week six, on their regular call, Harlan Ruiz, who ran Windale's returns floor, mentioned it almost as an aside. "Who's actually covering that channel after three? A few of my folks have been waiting till the next morning."

Nayeli went and actually counted, instead of guessing. Sixty-one questions were sitting unanswered in the channel right then, some of them over a day old.

Hand-sketched number line showing weekly support hours. Low bound 12.5 hours, calm case, feature already feels familiar. Point estimate 20 hours, the plan as sized. High bound 30 hours, anxious case, first time an AI's made this call. A separate mark past the high end for a sanity check, one support engineer's full week, 40 hours.
The range Nayeli should have sized before week one: 12.5 hours on a calm week, 30 on an anxious one, both well under one engineer's own working week, which is the whole point of checking.

Here's the part that actually cost something. It wasn't the awkward pause on the call. It was that a few agents had told Harlan they'd started pulling up the old dropdown to double-check the model's suggestion before trusting it, exactly the extra step the tool was supposed to remove.

We didn't understaff a chat channel. We understaffed a promise, and the promise was the only reason agents were willing to trust the model's suggestion over their own read.

So here's what Nayeli took back. She'd assumed one person, herself, loosely watching a channel between three other accounts, counted as a support plan. It never was. It was just her own spare attention, and that ran out around week three without anyone deciding it should.

She rebuilt it with real numbers. Fifteen agents, three flagged issues each a week, forty-five issues. Twenty minutes to actually resolve one. Fifteen hours. Plus a five-hour coverage floor to keep the four-hour promise. Twenty hours a week, split between her and one other Thornfield engineer, ten hours each.

Two weeks later, every flagged issue got answered inside the four-hour window. Thirty-eight of forty-five landed the same day. By the second week of the new plan, Harlan told her his agents had quietly stopped pulling up the old dropdown to check the model's work.

The thing I'd tell myself, back before I ever opened that channel the first Monday: a pilot doesn't fail because a person forgets to check Slack. It fails because nobody ever worked out how many hours checking it would really take, and called that number a plan.

BOUND, run for fifteen agents and one returns queue

This is a sizing question about how many real hours a pilot's support model needs, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. Weekly support hours equal two things added together: the time it takes to resolve every issue a pilot agent flags, plus a fixed coverage floor, the hours someone has to stay reachable to keep the promised response time, whether or not the queue is actually busy.
O, own the numbers. Fifteen pilot agents at Windale, about three flagged issues each a week, gives 45 issues. Twenty minutes average to resolve one, reading the flag, checking the model's reasoning, correcting it, replying. 45 times 20 is 900 minutes, 15 hours. The four-business-hour promise needs someone reachable roughly an hour a day, five days, five hours a week, regardless of volume. Total: 15 plus 5 is 20 hours a week.
U, use a range. That 20 assumes the feature feels moderately new. If agents already trusted a similar AI suggestion elsewhere on the same screen, issues drop to about 1.5 a week each, 12 and a half hours total. If this is the first time an AI has made this call on their job at all, issues climb to about 5 a week each, 30 hours total. Range: about 12.5 to 30 hours a week, centered on 20.
N, nail the sanity check. Two Thornfield support engineers split 20 hours a week, ten each, on top of the other pilot accounts they already cover, sustainable across an eight-week window. At the high end, 30 hours a week is fifteen each, more than a third of a working week for two straight months, worth flagging before the pilot starts, not in week six when someone's asking who's covering the channel.
D, direction. The response-time promise moves the total more than how new the feature feels does. Holding the issue rate steady and tightening the promise from four hours to one hour turns the coverage floor from five hours a week, a few planned check-ins, into something close to thirty, someone reachable nearly the whole day. That's a bigger jump than moving from the calmest case to the most anxious one on novelty alone.

What moves the total most
Response promise tightens from 4 hours to 1 hour+25
Feature novelty rises: 3 issues/agent/week to 5+10
Feature novelty falls: 3 issues/agent/week to 1.5−7.5
One more agent added to the pilot (15 to 16)+1
Tightening the promise swings the total two and a half times harder than the worst novelty case. That's why the promise, not the feature's newness, is the first thing worth re-checking if this number needs to move.

And if you want to be sure it really works, try it somewhere else

A regional veterinary chain runs the same question on a different floor. Auric Vet Group pilots an AI symptom-triage assistant at five clinics, twelve front-desk staff, sorting incoming calls into urgent and routine before a vet ever hears about them.

B, break it down. Weekly support hours equal the time to resolve every flagged mis-triage, plus a coverage floor for how fast a flagged urgent case has to get checked.
O, own the numbers. Twelve receptionists, about two flagged calls each a week, 24 issues. Fifteen minutes average to resolve one, 24 times 15 is 360 minutes, 6 hours. Because a missed urgent case is a safety question, not a convenience one, the promise is thirty minutes during clinic hours, which needs someone reachable close to the whole ten-hour day, call it 20 hours a week. Total: 6 plus 20 is 26 hours a week.
U, use a range. With one clinic in the pilot instead of five, resolution drops to about 1.5 hours, but the coverage floor barely moves, someone still has to watch the queue the whole day whether one clinic is calling in or five. Range: about 21.5 to 26 hours a week, tighter than Windale's, because the floor already dominates.
N, nail the sanity check. One Auric vet tech and one Thornfield engineer split it, thirteen hours each, workable for an eight-week pilot, though thin if either one takes a week off.
D, direction. Here the lever isn't the promise, it's already fixed at thirty minutes because it's a safety question nobody's going to loosen. What actually swings the total is how many clinics join the pilot. Five clinics needs a dedicated watcher for the whole day. One clinic's call volume would fit inside someone's existing job.

Same shape, different lever At Windale, tightening the response promise was the swing. At Auric, the promise is already fixed for safety, so the swing comes from how many clinics are in the pilot, not from how new the feature feels or how fast anyone promised to answer.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at 90 seconds. Skip straight to the split: two categories, 20 hours a week at a four-hour promise, and the promise is what to interrogate first.
Cost: a manager caps the support budget at fifteen hours a week instead of naming engineers. Work backward: that covers the resolution hours, but not the five-hour floor, so either the promise loosens or a second person has to be found.
The model got better: a newer version of the classifier misses far less. The coverage floor doesn't shrink, a quiet model still needs someone reachable in case the one time it's wrong is also the one time it matters. Only the resolution hours drop.

Where people run it wrong.
They treat "someone will keep an eye on it" as a plan instead of a number.
They price the resolution work and forget the coverage floor entirely, so the promise breaks on a quiet week just as easily as a busy one.
They size the plan off week one, when everyone's still curious and careful, instead of the steady state six weeks in.

How to use it live. Say the equation before any number: "weekly support hours is the time to resolve what gets flagged, plus a coverage floor for whatever response time you promised, sized to who's actually free, not to how calm week one looked." That buys the time to name real hours instead of reaching for "we'll have someone on it" and calling it a plan.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about sizing a pilot's support model, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question, how many real hours a pilot's support plan needs, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Nayeli Cortes, product manager at Thornfield AI, running her first real pilot of its return-reason classifier with fifteen agents on Windale & Co's returns floor.
3 · THE HABIT
What did Nayeli's support coverage quietly shrink to by week five?
Tap to flip
ANSWER
Checking the shared Slack channel whenever she had a spare minute, down from every hour in week one. Nobody had ever sized how many hours it actually needed, so nothing stopped it fading.
4 · THE BUILD-UP, IN THIS STORY
What's the sizing build-up this answer turns on?
Tap to flip
ANSWER
Weekly support hours equal resolution hours (issues per agent per week, times agents, times minutes to fix one) plus a coverage floor, the hours someone must stay reachable to keep the promised response time.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Treating one person loosely watching a Slack channel as a real support plan, instead of pricing the four-hour promise in hours before agreeing to it. It made sense because it cost nothing to say yes in the kickoff meeting, and nothing broke in week one.
6 · THE NUMBER
Fill in the blank: 15 agents at 3 issues each is ___ issues a week, needing ___ hours to resolve, plus a ___ hour coverage floor, for ___ hours a week total.
Tap to flip
ANSWER
45 issues; 15 hours; 5 hours; 20 hours a week total.
7 · THE REPLAY
Same pilot, response promise tightens from four hours to one. What changes?
Tap to flip
ANSWER
The coverage floor alone jumps from 5 hours a week to close to 30, because someone now has to stay reachable nearly the whole day instead of checking in a few times. Total support hours climb from 20 a week to around 45.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which product, and what plays the role the response promise played at Windale?
Tap to flip
ANSWER
Auric Vet Group's AI symptom-triage assistant, piloted at five veterinary clinics. There, the response promise is already fixed for safety, so the number of clinics in the pilot is the lever that swings the total.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Windale's support plan actually fail by week six?
  • A. The classifier's suggestions got less accurate over time.
  • B. Nayeli assumed one person loosely watching a Slack channel counted as a sized support plan, and never priced what the four-hour promise actually cost in hours.
  • C. Windale's agents refused to use the new suggested reasons.
  • D. The pilot ran through Windale's busiest return week of the year.
Show hint
Look for the actual planning decision made before the pilot started, not something that happened during it.
Show answer
B. The plan was never a number. "Keep an eye on the channel" doesn't scale to real hours, and once Nayeli got pulled onto other accounts, coverage quietly dropped to whenever she had a spare minute.
True or false, with why
2. True or false: most of the 20 hours a week goes toward keeping the four-hour response promise, not toward actually resolving flagged issues.
  • True
  • False
Show hint
Check the O step's own split between the two added parts.
Show answer
False. Resolving the 45 flagged issues themselves is 15 of the 20 hours; the coverage floor for the promise is only 5. The bigger cost is doing the actual work, not just staying reachable, until the promise gets tightened.
Fill in the blank
3. Fifteen agents at three flagged issues each a week is ___ issues. At twenty minutes each, that's ___ hours. Add a five-hour coverage floor and the total is ___ hours a week.
Show hint
Check the O step's own arithmetic.
Show answer
45 issues; 15 hours; 20 hours a week. 15 times 3 is 45. 45 times 20 minutes is 900 minutes, 15 hours. 15 plus 5 is 20.
Multiple choice
4. What old decision does this answer actually take back?
  • A. Promising a four-hour response time in the first place.
  • B. Letting fifteen agents into the pilot instead of a smaller group.
  • C. Treating one person loosely watching a Slack channel as a sized support plan, instead of pricing the promise in real hours before agreeing to it.
  • D. Building the classifier to suggest a reason instead of just flagging low-confidence returns.
Show hint
Look for the actual staffing decision made before the pilot started, not a scope or design choice.
Show answer
C. The four-hour promise itself was fine. What broke it was never turning "someone will handle it" into an actual number of hours and a name against the coverage floor.
Short answer, apply it yourself
5. Pick a pilot or a beta feature you've used yourself. What response time did it seem to promise, even if nobody said so out loud, and who did you assume was covering it?
Show hint
Think about a feature with a feedback button, a chat widget, or a "report a problem" link, and what happened the one time you actually used it.
Show answer
Model answer: "A budgeting app had a 'flag this transaction as wrong' button. It implied someone would look at it soon, but nothing said how soon, and I assumed a real person was watching. My flag sat unanswered for eleven days, because nobody had ever sized how many hours answering those would take against how many people were free to do it."
Short answer, the number question
6. If the response promise tightened from four hours to one hour, roughly how many hours a week would the coverage floor take, and why does it jump that much instead of rising a little?
Show hint
Check the D step. Think about the difference between checking in a few times a day and staying reachable continuously.
Show answer
About thirty hours a week. A four-hour promise only needs a few planned check-ins, about an hour a day. A one-hour promise means someone has to be reachable close to the entire working day, which turns a small floor into almost a full-time watch, a much bigger jump than resolution hours alone would ever produce.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more