What support model does a pilot need?
- Size support hours from real arithmetic: issues per pilot user each week, times minutes to fix one, plus a coverage floor for the promised response time.Why: "someone will handle it" isn't a plan, it's a hope, and hope runs out around week five.
- Set the response-time promise before you set the headcount, because the promise decides the coverage floor.Why: a one-hour promise needs someone reachable nearly all day; a four-hour promise needs a few planned check-ins, same issue count, very different staffing.
- Give the estimate a range, not one number, based on how new the feature feels to the pilot users.Why: a familiar AI suggestion gets a shrug and a quick check; a first-time AI call on someone's job gets questioned every time.
- Check the total against how many hours two real people can sustain for the whole pilot window, not just week one.Why: a plan that only survives the first curious week isn't a support plan, it's a honeymoon.
- Know that tightening the response promise moves the total more than how new the feature feels.Why: cut the wrong assumption and either the promise breaks or the two people covering it burn out.
How to answer this, stage by stage
Nobody is grading whether you land on exactly 20. They're grading whether you name the real build-up before touching a number, whether the range comes from something real, and whether you can say which assumption would move the total most. Six moves get you there.
Let's learn
What actually breaks a pilot's support plan? Not a bad model. Bad math about how many hours answering questions really costs.
The product here is small: a button that reads a returned item's photo and the customer's note, and suggests why it's coming back, wrong size, damaged, not as described, before a returns agent has to type it in themselves.
Before the pilot, Windale & Co's returns agents picked the reason themselves, one by one, from a dropdown that didn't always fit. It worked. It was just slow, and it wasn't consistent from one agent to the next.
During the pilot, fifteen agents got the classifier's suggestion first, and mostly it saved real time. But agents flagged the odd one, a shirt the model called "not as described" that was really just the wrong size, or a plain question about why it picked what it picked.
Here is the turn. Those flagged questions were never the real problem. The real problem is that nobody had worked out how many hours answering them would actually take, or who was supposed to be free to do it.
At its worst, that costs the whole pilot's credibility. Eight weeks in, a support plan can look fine on a kickoff slide and still leave agents waiting a full day for an answer, because nobody ever turned "we'll keep an eye on it" into a number of hours.
The choice I would take back. The pilot plan promised Windale a reply within four business hours. Nobody worked out what that promise actually costs to keep, in real hours, before agreeing to it in the kickoff meeting.
What I would leave alone. Not every pilot needs this. A pilot testing a one-line copy change on a button doesn't need a coverage floor at all, nobody's job depends on it being right within four hours. Save the real hours for changes that touch a call an agent gets graded on.
The lesson. A support model isn't a sentence in a kickoff deck. It's a staffing plan with real hours attached, and it has to survive the whole pilot window, not just the first curious week.
Now here is the same thing as a story
Skip this if you already believe a pilot's support plan needs real hours attached to it, not just a name on a Slack channel. Read on if you want to feel why.
Every Monday morning, Nayeli Cortes opened the same shared Slack channel before she'd finished her coffee, checking what had piled up over the weekend.
Nayeli's a product manager at Thornfield AI, and Windale & Co's returns floor was her first pilot with real customers on the other end of every flagged item. Fifteen agents, eight weeks, watching whether her team's return-reason classifier actually held up outside a demo.
For the first two weeks, that channel was the best part of her mornings. She checked it every hour or so. The questions were mostly curious, "why'd it pick this one," "is this a bug or is it right," and she'd answer in minutes. Agents warmed to the tool fast. It was doing exactly what it was supposed to.
So she checked it a little less. Week three, twice a day. Week four, once, usually at the end of hers. By weeks five and six, whenever she had a minute, because nothing bad had happened yet, and she had two other pilots pulling at the same hours.
Nobody told her to slow down. She just found, somewhere in that stretch, that she wasn't the only thing standing between an agent's question and an answer, except that she was, and she'd stopped noticing.
Then, in week six, on their regular call, Harlan Ruiz, who ran Windale's returns floor, mentioned it almost as an aside. "Who's actually covering that channel after three? A few of my folks have been waiting till the next morning."
Nayeli went and actually counted, instead of guessing. Sixty-one questions were sitting unanswered in the channel right then, some of them over a day old.
Here's the part that actually cost something. It wasn't the awkward pause on the call. It was that a few agents had told Harlan they'd started pulling up the old dropdown to double-check the model's suggestion before trusting it, exactly the extra step the tool was supposed to remove.
So here's what Nayeli took back. She'd assumed one person, herself, loosely watching a channel between three other accounts, counted as a support plan. It never was. It was just her own spare attention, and that ran out around week three without anyone deciding it should.
She rebuilt it with real numbers. Fifteen agents, three flagged issues each a week, forty-five issues. Twenty minutes to actually resolve one. Fifteen hours. Plus a five-hour coverage floor to keep the four-hour promise. Twenty hours a week, split between her and one other Thornfield engineer, ten hours each.
Two weeks later, every flagged issue got answered inside the four-hour window. Thirty-eight of forty-five landed the same day. By the second week of the new plan, Harlan told her his agents had quietly stopped pulling up the old dropdown to check the model's work.
The thing I'd tell myself, back before I ever opened that channel the first Monday: a pilot doesn't fail because a person forgets to check Slack. It fails because nobody ever worked out how many hours checking it would really take, and called that number a plan.
BOUND, run for fifteen agents and one returns queue
This is a sizing question about how many real hours a pilot's support model needs, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Weekly support hours equal two things added together: the time it takes to resolve every issue a pilot agent flags, plus a fixed coverage floor, the hours someone has to stay reachable to keep the promised response time, whether or not the queue is actually busy.
O, own the numbers. Fifteen pilot agents at Windale, about three flagged issues each a week, gives 45 issues. Twenty minutes average to resolve one, reading the flag, checking the model's reasoning, correcting it, replying. 45 times 20 is 900 minutes, 15 hours. The four-business-hour promise needs someone reachable roughly an hour a day, five days, five hours a week, regardless of volume. Total: 15 plus 5 is 20 hours a week.
U, use a range. That 20 assumes the feature feels moderately new. If agents already trusted a similar AI suggestion elsewhere on the same screen, issues drop to about 1.5 a week each, 12 and a half hours total. If this is the first time an AI has made this call on their job at all, issues climb to about 5 a week each, 30 hours total. Range: about 12.5 to 30 hours a week, centered on 20.
N, nail the sanity check. Two Thornfield support engineers split 20 hours a week, ten each, on top of the other pilot accounts they already cover, sustainable across an eight-week window. At the high end, 30 hours a week is fifteen each, more than a third of a working week for two straight months, worth flagging before the pilot starts, not in week six when someone's asking who's covering the channel.
D, direction. The response-time promise moves the total more than how new the feature feels does. Holding the issue rate steady and tightening the promise from four hours to one hour turns the coverage floor from five hours a week, a few planned check-ins, into something close to thirty, someone reachable nearly the whole day. That's a bigger jump than moving from the calmest case to the most anxious one on novelty alone.
And if you want to be sure it really works, try it somewhere else
A regional veterinary chain runs the same question on a different floor. Auric Vet Group pilots an AI symptom-triage assistant at five clinics, twelve front-desk staff, sorting incoming calls into urgent and routine before a vet ever hears about them.
B, break it down. Weekly support hours equal the time to resolve every flagged mis-triage, plus a coverage floor for how fast a flagged urgent case has to get checked.
O, own the numbers. Twelve receptionists, about two flagged calls each a week, 24 issues. Fifteen minutes average to resolve one, 24 times 15 is 360 minutes, 6 hours. Because a missed urgent case is a safety question, not a convenience one, the promise is thirty minutes during clinic hours, which needs someone reachable close to the whole ten-hour day, call it 20 hours a week. Total: 6 plus 20 is 26 hours a week.
U, use a range. With one clinic in the pilot instead of five, resolution drops to about 1.5 hours, but the coverage floor barely moves, someone still has to watch the queue the whole day whether one clinic is calling in or five. Range: about 21.5 to 26 hours a week, tighter than Windale's, because the floor already dominates.
N, nail the sanity check. One Auric vet tech and one Thornfield engineer split it, thirteen hours each, workable for an eight-week pilot, though thin if either one takes a week off.
D, direction. Here the lever isn't the promise, it's already fixed at thirty minutes because it's a safety question nobody's going to loosen. What actually swings the total is how many clinics join the pilot. Five clinics needs a dedicated watcher for the whole day. One clinic's call volume would fit inside someone's existing job.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at 90 seconds. Skip straight to the split: two categories, 20 hours a week at a four-hour promise, and the promise is what to interrogate first.
Cost: a manager caps the support budget at fifteen hours a week instead of naming engineers. Work backward: that covers the resolution hours, but not the five-hour floor, so either the promise loosens or a second person has to be found.
The model got better: a newer version of the classifier misses far less. The coverage floor doesn't shrink, a quiet model still needs someone reachable in case the one time it's wrong is also the one time it matters. Only the resolution hours drop.
Where people run it wrong.
They treat "someone will keep an eye on it" as a plan instead of a number.
They price the resolution work and forget the coverage floor entirely, so the promise breaks on a quiet week just as easily as a busy one.
They size the plan off week one, when everyone's still curious and careful, instead of the steady state six weeks in.
How to use it live. Say the equation before any number: "weekly support hours is the time to resolve what gets flagged, plus a coverage floor for whatever response time you promised, sized to who's actually free, not to how calm week one looked." That buys the time to name real hours instead of reaching for "we'll have someone on it" and calling it a plan.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pilot design and POC-to-production
- #1 Design a four-week pilot for an AI feature with one enterprise customer.
- #2 What success criteria should be agreed before a pilot begins?
- #3 Explain the difference between a pilot and a beta.
- #4 How do you choose pilot customers, and what makes a bad one?
- #5 Describe the pilot-to-production gap and the work that lives in it.
- #6 Why do most AI POCs fail to reach production? Give four reasons.