Describe how you would scale a pilot from one customer to twenty.
- Size the rollout on two real numbers, onboarding hours per firm and support hours per firm, both multiplied across all twenty.Why: a rollout with no numbers behind it is a guess wearing a growth target.
- Sample each new firm's own leases against the first customer's before signing them, not after.Why: that similarity is what decides whether one firm costs twenty hours to onboard and another costs fifty.
- Check the weekly total against the team's real free hours, not their job titles.Why: the near miss in this story happened because that real ceiling was never written down anywhere.
- Give the plan a range, low if the new firms look like the first customer, high if they don't, not one number for the board deck.Why: a single number hides a nine-firm swing between the two ends.
- Stagger the hard-to-onboard firms further apart in the schedule instead of signing all nineteen on the same clock.Why: spacing costs nothing and it's what actually keeps the queue from backing up.
- Know which lever to pull first if the hours don't fit: check similarity before reaching for a new hire.Why: a hire only moves the ceiling by a fixed amount; an unfamiliar batch of leases moves the total further than one hire can close.
How to answer this, stage by stage
Nobody is grading whether you land on exactly fifty-five hours a week. They're grading whether you can defend the two numbers behind "twenty customers," whether the range is honest, and whether you close on something the room can check. Seven moves get you there.
Let's learn
The product is a tool that reads a commercial lease, pages of legal language, and pulls out a short summary: the rent, when it goes up, when the firm can walk away, when it has to renew. An analyst used to build that summary by hand, one lease at a time.
Before the tool, a pilot ran with one customer, a commercial real estate firm with about sixty leases. It went well. The model learned that firm's lease templates, the analysts trusted the summaries, and the pilot became the case study everyone wanted to repeat nineteen more times.
So the plan was simple. If the tool worked for one firm, it would work for twenty. Same onboarding hours, same support hours, just multiplied. That math fit neatly inside the small support team's calendar, on paper, and everyone signed off.
Here's the turn. The tool reading a new firm's leases correctly was never really the hard part. The hard part was that not every firm's leases look like the first customer's. Some are clean, standard, easy. Some are full of one-off clauses that take the model, and the team, far longer to get right. A rollout plan built on one firm's numbers doesn't know the difference until it's already signed nineteen contracts.
At its worst, that gap shows up as a support queue nobody's watching. If the team's hours all go toward the one new firm that's fighting them, the firms that are already live and quiet get checked less carefully. And a lease flag that sits unread is exactly the kind of thing that turns into a missed deadline nobody notices until it's almost too late.
The choice I would take back. I sized the whole nineteen-firm rollout off the one firm we'd already onboarded, because that was the only real number the team had. That felt responsible. Real data, not a guess. It also assumed every future firm's leases would look the same, which is the one thing a single data point can never tell you.
What I would leave alone. The model itself doesn't need to change to scale. A firm whose leases take fifty hours to calibrate today will take fifty hours whether you're onboarding your second customer or your twentieth. Scaling doesn't make one firm's onboarding harder. It just means more of them stack up on the same small team's calendar. Don't waste time trying to make the model handle scale. It already does its part fine.
The lesson. Twenty customers was never really a question about the model. It was a staffing and scheduling question wearing a growth number. Price each new firm's onboarding by checking its leases first, the same way a moving company prices a job by counting boxes, not houses.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why checking similarity first is the whole decision, not a scheduling detail before the real work starts.
Zofia Kovarik has run onboarding for Abstraq, a lease-abstraction tool, for two years. She's the one who takes a signed contract and turns it into a working analyst tool. She can tell within the first data pull whether a firm's lease templates are going to be clean or a mess, mostly from how consistent the file names are.
Halgren Commercial Realty was the pilot. Sixty leases, mostly standard retail and office space, one lease template repeated with small changes. Zofia's team connected their document folder, calibrated the model on a sample of fifteen leases, and trained their two analysts on how to review what the model flagged. Twenty hours of work, spread over two weeks.
For three months, that pilot was the best thing on Zofia's roadmap. The model learned Halgren's leases fast. Their analysts stopped double-checking every summary and started spot-checking the ones the model itself flagged as unsure, about two hours a week, and caught a real issue maybe once a month. Everyone who saw the numbers wanted the next nineteen firms signed immediately.
So they did. Sales closed nineteen new firms in six weeks, and the board wanted all twenty live within six months. Zofia wrote the rollout plan the way the pilot had gone: twenty hours to onboard each new firm, two hours a week to support each one once live, a new firm added roughly every ten days. On a spreadsheet, that fit inside her three-person team's fifty-five hours a week with room to spare. Everyone signed off in an afternoon.
The first five firms went almost exactly like Halgren. Then came the sixth: Corvath Realty.
Corvath's portfolio ran heavy on negotiated ground leases and percentage-rent retail deals, nothing like Halgren's clean template. What was supposed to be a twenty-hour onboarding turned into fifty-two, most of it Zofia's team hand-correcting the model's first guesses on nearly every lease in the calibration batch. For those two weeks, Corvath ate almost half the team's entire week, and everyone's attention went where the fire was.
Which meant nobody was watching Halgren's queue as closely. On the Tuesday of Corvath's second onboarding week, the model flagged one of Halgren's own leases as unsure, a rent-escalation date on a hand-annotated amendment page. The flag sat there. Not because anyone decided to ignore it. Because the team spent that week putting out Corvath's fire instead.
It sat unread for four days. On the fifth, one of Halgren's analysts, doing an unrelated check, noticed the lease had ninety days to give notice on a rent renegotiation, and the flagged date meant that window closed in two days, not the six weeks everyone had assumed. They caught it. Two days to spare, one phone call, no missed deadline. But it was close enough that Zofia paused the rollout that same afternoon.
In the original plan, when Zofia proposed "twenty hours a firm, two hours a week to support it," nobody in the room questioned it, because it was the only real number anyone had, and Halgren had proven it. It made sense in that meeting. It stopped making sense the moment Corvath's leases turned out to be nothing like Halgren's.
The second version of the plan kept the same team, the same fifty-five hours, and the same target of twenty firms. What changed: before signing any new firm, Zofia's team now pulls five sample leases and checks them against Halgren's, before the contract goes out, not after. Firms that look similar get onboarded on the original clock. Firms that don't, like Corvath, get spaced six weeks apart from any other hard firm, so the team is never fighting two fires at once.
The nineteenth firm went live in month eight instead of month six, two months later than the original promise. But every support flag across all twenty firms got reviewed within two days, every time, for the rest of the rollout. No second near miss.
The thing I'd tell myself, back in that afternoon when everyone signed off on "twenty hours a firm": one real customer proves the model can do the work. It doesn't prove the next nineteen will hand you the same kind of work. The only way to tell the difference is to look before you sign, not after.
BOUND, priced before the sixth contract goes out
This is a sizing question about how many hours it actually takes to support twenty firms, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Twenty customers isn't one number. It's two: onboarding hours per new firm, times the nineteen firms still being added, plus support hours per firm per week, times all twenty firms once they're live. The two run at the same time during the rollout, onboarding for the newest firms on top of support for the ones already live, not instead of it.
O, own the numbers. For a three-person customer success team: onboarding runs about twenty hours a firm when its leases look like the pilot customer's, connecting the document system, calibrating on a sample batch, training the analysts. Support runs about two hours a week a firm once live, mostly reviewing the leases the model flags as unsure. The team has about fifty-five hours a week free for this, combined.
U, use a range. If the new firms' leases look like the pilot customer's, standard templates, normal terms, twenty hours of onboarding and two hours a week of support holds for all twenty firms. If a firm's leases run heavy on negotiated, one-off terms, onboarding can run to fifty hours and support to five hours a week. Same twenty firms, the honest range for how many the team can actually support is anywhere from about eleven to twenty, depending only on how much of the new nineteen's paperwork looks like the first firm's.
N, nail the sanity check. At the busiest point in the rollout, nineteen firms live and the twentieth still onboarding, that week needs about forty-eight hours: thirty-eight for support across the nineteen, ten for onboarding the last one. Against the team's real fifty-five hours a week, that's seven hours of slack. Tight, and it only holds if none of those nineteen turned out to be the expensive kind of firm.
D, direction. Two assumptions could move this number, and they don't move it the same amount. Whether the new firms' leases actually look like the pilot customer's swings the team's realistic ceiling from twenty firms down to about eleven, because it moves both onboarding hours and weekly support hours at once. Hiring one more support specialist adds about fifteen hours a week of capacity, which helps, but only closes part of that gap. Similarity is the bigger lever. Hiring is the smaller one, and it's the one people reach for first because it feels like doing something.
And if you want to be sure it really works, try it somewhere else
A veterinary group pilots an AI tool that triages after-hours phone calls, deciding which ones need an on-call vet and which can wait until morning, with one clinic as the first customer, scaling to a twelve-clinic regional group.
B, break it down. Same shape, different work. Onboarding hours per clinic, calibrating the triage thresholds to that clinic's own vets and connecting the after-hours phone line, times eleven new clinics. Support hours per clinic per week, reviewing the calls the tool flags as uncertain, times all twelve clinics live.
O, own the numbers. Onboarding runs about fifteen hours a clinic for a small-animal-only practice like the pilot clinic. Support runs about three hours a week a clinic once live. The rollout team, two people, has about forty hours a week free.
U, use a range. If the new clinics are small-animal-only like the pilot, fifteen hours onboarding and three hours a week support holds for all twelve. If a clinic also handles large animals, onboarding can run to thirty-five hours and support to seven hours a week, because a horse with colic needs a different triage judgment than a cat with a limp.
N, nail the sanity check. Twelve clinics at the low end need about thirty-six hours a week once fully live, comfortably under the forty-hour cap. At the high end, they need about eighty-four hours a week, more than double what the team has.
D, direction. Same tension as the lease tool. Whether the new clinics handle the same kind of animals as the pilot clinic swings the realistic ceiling the most. Adding a third person to the rollout team helps, but adding species the tool has never triaged before is what actually breaks the plan.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: check lease similarity before signing, don't schedule scope by the calendar or the customer count alone.
Cost: the company can only afford to hire zero more people this year. Don't cut the check, cut the pace: onboard fewer firms at once so the team's real hours stay the same size relative to the queue.
The model got better: a newer extraction model rarely misses on standard leases anymore. That doesn't remove the need to check similarity, it just moves where the risk sits, from accuracy on common leases to accuracy on the rare, negotiated ones.
Where people run it wrong.
They multiply the pilot customer's numbers straight across twenty without checking whether the next nineteen are actually like the first one.
They count hours available by job title instead of by what the team actually has free once everything else on their plate is subtracted.
They treat a new hire as the fix for a scaling problem before checking whether the real problem is which customers got signed, not how many people review them.
How to use it live. Say the equation before naming a single number: "the twenty customers aren't the real question, the hours behind each one are, so before I size the team, I'd want to know how similar the other nineteen actually look to the one we've already onboarded." That buys the room to ask a real question instead of guessing a headcount that only sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pilot design and POC-to-production
- #1 Design a four-week pilot for an AI feature with one enterprise customer.
- #2 What success criteria should be agreed before a pilot begins?
- #3 Explain the difference between a pilot and a beta.
- #4 How do you choose pilot customers, and what makes a bad one?
- #5 Describe the pilot-to-production gap and the work that lives in it.
- #6 Why do most AI POCs fail to reach production? Give four reasons.