CaseAdvancedShipping & Model Lifecycle / Pilot design and POC-to-production / #13

Describe how you would scale a pilot from one customer to twenty.

The direct answer
Size the rollout on two real numbers: onboarding hours for each new firm, and support hours for each live firm, both multiplied across all twenty, never assumed to match the first customer. Before signing any of the other nineteen, sample a handful of that firm's own leases against the first customer's, because that similarity is what decides whether onboarding takes twenty hours or fifty. If the weekly total doesn't fit inside the team's real hours, slow the rollout or add hours. Don't sign the deal and hope.
Do this, in order
  1. Size the rollout on two real numbers, onboarding hours per firm and support hours per firm, both multiplied across all twenty.Why: a rollout with no numbers behind it is a guess wearing a growth target.
  2. Sample each new firm's own leases against the first customer's before signing them, not after.Why: that similarity is what decides whether one firm costs twenty hours to onboard and another costs fifty.
  3. Check the weekly total against the team's real free hours, not their job titles.Why: the near miss in this story happened because that real ceiling was never written down anywhere.
  4. Give the plan a range, low if the new firms look like the first customer, high if they don't, not one number for the board deck.Why: a single number hides a nine-firm swing between the two ends.
  5. Stagger the hard-to-onboard firms further apart in the schedule instead of signing all nineteen on the same clock.Why: spacing costs nothing and it's what actually keeps the queue from backing up.
  6. Know which lever to pull first if the hours don't fit: check similarity before reaching for a new hire.Why: a hire only moves the ceiling by a fixed amount; an unfamiliar batch of leases moves the total further than one hire can close.

How to answer this, stage by stage

Nobody is grading whether you land on exactly fifty-five hours a week. They're grading whether you can defend the two numbers behind "twenty customers," whether the range is honest, and whether you close on something the room can check. Seven moves get you there.

1
Scope it to one real product and one real starting customer
Say it like this
"Let's ground this. Say we've built a tool that reads a commercial lease and pulls out the numbers that matter, the base rent, the yearly increases, the renewal window, the way out of the contract, into a short summary an analyst used to write by hand. One customer means one real firm, say a commercial real estate firm with about sixty leases in its portfolio, not a logo on a slide. Twenty means that firm plus nineteen more, signed over the next several months."
Why this works
Stops the answer from staying abstract before a single number gets attached.
2
Say the structure out loud before naming any numbers
Say it like this
"The number that matters here isn't twenty customers. It's two numbers per customer: how many hours it takes to onboard one firm, and how many hours a week it takes to support one firm once it's live. Multiply both across twenty and you get the real size of this job. A plan that only counts logos signed isn't a plan. It's a scoreboard."
Why this works
Shows the equation before the arithmetic, so the numbers that follow read as a plan, not a headcount wish.
3
Break down the equation and what has to add up
Say it like this
"Here's the shape. Total onboarding hours equals hours per firm, times the nineteen new firms. Total weekly support hours equals hours per firm, times all twenty firms once they're live. Add the two, because onboarding the firms still ramping up happens on top of supporting the ones already live, not instead of it."
Why this works
This is the B step of BOUND: the equation, stated before a single number gets attached to it.
4
Own the real numbers behind each term
Say it like this
"For a lease-abstraction tool with a three-person customer success team, that's: onboarding runs about twenty hours a firm when its leases look like the first customer's, connecting their document system, calibrating the model on a sample batch, training their analysts. Support runs about two hours a week a firm once it's live, mostly reviewing the handful of leases the model flags as unsure. The team has about fifty-five hours a week free for this, combined."
Why this works
This is the O step: a real proposed capacity plan, not "we'll figure it out as we go."
5
Give the range, not one number
Say it like this
"If the new firms' leases look like the first customer's, standard retail and office leases with a normal template, twenty hours of onboarding and two hours a week of support holds for all twenty. If a firm's leases are heavy on negotiated ground leases, percentage rent, one-off clauses, onboarding can run to fifty hours, and support can run to five hours a week. Same twenty firms, more than double the hours, and the only thing that changed is how much their paperwork looks like the first customer's."
Why this works
A single number here would claim a confidence the plan doesn't have yet.
6
Sanity check the plan against the team's real hours
Say it like this
"At the busiest point in the rollout, say nineteen firms are already live and the twentieth is still being onboarded, that week needs about forty-eight hours: thirty-eight for support on the nineteen, ten for onboarding the last one. Against a team with fifty-five real hours a week, that's seven hours of slack. Tight, but it holds, as long as none of those nineteen turned out to be the expensive kind."
Why this works
This is the N step, and it's the step most rollout plans skip entirely.
7
Name what moves it most, then close on the one line
Say it like this
"If I had to bet on what changes this number most, it's not how many people we hire. It's whether the new firms' leases actually look like the first customer's. So here's the line I'd say out loud: I'd check a sample of each new firm's leases before we ever sign the contract, because that one check decides whether we're supporting twenty firms comfortably or drowning at eleven, and no amount of hiring fixes a batch of leases nobody looked at first."
Why this works
Closes on the literal ask, a claim someone could actually check, not a vibe about "scaling carefully."
If you remember one thing Check how close a new firm's leases are to the first customer's before you sign it, not after. That's the one number that decides whether the team's fifty-five hours a week covers twenty firms or eleven.

Let's learn

The product is a tool that reads a commercial lease, pages of legal language, and pulls out a short summary: the rent, when it goes up, when the firm can walk away, when it has to renew. An analyst used to build that summary by hand, one lease at a time.

Before the tool, a pilot ran with one customer, a commercial real estate firm with about sixty leases. It went well. The model learned that firm's lease templates, the analysts trusted the summaries, and the pilot became the case study everyone wanted to repeat nineteen more times.

Knowledge spark: what's a lease abstract? A short summary of a long lease. It pulls out the parts that matter, the rent, the dates, the way out, so nobody has to reread forty pages to answer one question.

So the plan was simple. If the tool worked for one firm, it would work for twenty. Same onboarding hours, same support hours, just multiplied. That math fit neatly inside the small support team's calendar, on paper, and everyone signed off.

The extra hours didn't come from a new budget. They came out of the first customer's own queue.

Here's the turn. The tool reading a new firm's leases correctly was never really the hard part. The hard part was that not every firm's leases look like the first customer's. Some are clean, standard, easy. Some are full of one-off clauses that take the model, and the team, far longer to get right. A rollout plan built on one firm's numbers doesn't know the difference until it's already signed nineteen contracts.

At its worst, that gap shows up as a support queue nobody's watching. If the team's hours all go toward the one new firm that's fighting them, the firms that are already live and quiet get checked less carefully. And a lease flag that sits unread is exactly the kind of thing that turns into a missed deadline nobody notices until it's almost too late.

The decision that mattered Check how similar a new firm's leases are to the first customer's before signing it, not after. That's the one number that decides whether onboarding takes twenty hours or fifty.

The choice I would take back. I sized the whole nineteen-firm rollout off the one firm we'd already onboarded, because that was the only real number the team had. That felt responsible. Real data, not a guess. It also assumed every future firm's leases would look the same, which is the one thing a single data point can never tell you.

What I would leave alone. The model itself doesn't need to change to scale. A firm whose leases take fifty hours to calibrate today will take fifty hours whether you're onboarding your second customer or your twentieth. Scaling doesn't make one firm's onboarding harder. It just means more of them stack up on the same small team's calendar. Don't waste time trying to make the model handle scale. It already does its part fine.

The lesson. Twenty customers was never really a question about the model. It was a staffing and scheduling question wearing a growth number. Price each new firm's onboarding by checking its leases first, the same way a moving company prices a job by counting boxes, not houses.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why checking similarity first is the whole decision, not a scheduling detail before the real work starts.

Zofia Kovarik has run onboarding for Abstraq, a lease-abstraction tool, for two years. She's the one who takes a signed contract and turns it into a working analyst tool. She can tell within the first data pull whether a firm's lease templates are going to be clean or a mess, mostly from how consistent the file names are.

Halgren Commercial Realty was the pilot. Sixty leases, mostly standard retail and office space, one lease template repeated with small changes. Zofia's team connected their document folder, calibrated the model on a sample of fifteen leases, and trained their two analysts on how to review what the model flagged. Twenty hours of work, spread over two weeks.

For three months, that pilot was the best thing on Zofia's roadmap. The model learned Halgren's leases fast. Their analysts stopped double-checking every summary and started spot-checking the ones the model itself flagged as unsure, about two hours a week, and caught a real issue maybe once a month. Everyone who saw the numbers wanted the next nineteen firms signed immediately.

So they did. Sales closed nineteen new firms in six weeks, and the board wanted all twenty live within six months. Zofia wrote the rollout plan the way the pilot had gone: twenty hours to onboard each new firm, two hours a week to support each one once live, a new firm added roughly every ten days. On a spreadsheet, that fit inside her three-person team's fifty-five hours a week with room to spare. Everyone signed off in an afternoon.

The first five firms went almost exactly like Halgren. Then came the sixth: Corvath Realty.

Corvath's portfolio ran heavy on negotiated ground leases and percentage-rent retail deals, nothing like Halgren's clean template. What was supposed to be a twenty-hour onboarding turned into fifty-two, most of it Zofia's team hand-correcting the model's first guesses on nearly every lease in the calibration batch. For those two weeks, Corvath ate almost half the team's entire week, and everyone's attention went where the fire was.

Hand-sketched comparison of two desks. Left, labeled One profile assumed, a document icon with the caption New firm onboarding, 52 hours this week, and a red-orange marked sheet reading First customer's flag, unread, day 4. Right, labeled Similarity checked first, a document icon with the caption New firm spaced 6 weeks out, and a green marked sheet reading First customer's flag, reviewed, same day.
Same team, same fifty-five hours. One plan spends all of it on the loudest firm. The other leaves the first customer's queue covered.

Which meant nobody was watching Halgren's queue as closely. On the Tuesday of Corvath's second onboarding week, the model flagged one of Halgren's own leases as unsure, a rent-escalation date on a hand-annotated amendment page. The flag sat there. Not because anyone decided to ignore it. Because the team spent that week putting out Corvath's fire instead.

We didn't spend Corvath's fifty-two hours from a new budget line. We spent them out of Halgren's own queue.

It sat unread for four days. On the fifth, one of Halgren's analysts, doing an unrelated check, noticed the lease had ninety days to give notice on a rent renegotiation, and the flagged date meant that window closed in two days, not the six weeks everyone had assumed. They caught it. Two days to spare, one phone call, no missed deadline. But it was close enough that Zofia paused the rollout that same afternoon.

In the original plan, when Zofia proposed "twenty hours a firm, two hours a week to support it," nobody in the room questioned it, because it was the only real number anyone had, and Halgren had proven it. It made sense in that meeting. It stopped making sense the moment Corvath's leases turned out to be nothing like Halgren's.

The second version of the plan kept the same team, the same fifty-five hours, and the same target of twenty firms. What changed: before signing any new firm, Zofia's team now pulls five sample leases and checks them against Halgren's, before the contract goes out, not after. Firms that look similar get onboarded on the original clock. Firms that don't, like Corvath, get spaced six weeks apart from any other hard firm, so the team is never fighting two fires at once.

The nineteenth firm went live in month eight instead of month six, two months later than the original promise. But every support flag across all twenty firms got reviewed within two days, every time, for the rest of the rollout. No second near miss.

The thing I'd tell myself, back in that afternoon when everyone signed off on "twenty hours a firm": one real customer proves the model can do the work. It doesn't prove the next nineteen will hand you the same kind of work. The only way to tell the difference is to look before you sign, not after.

BOUND, priced before the sixth contract goes out

This is a sizing question about how many hours it actually takes to support twenty firms, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. Twenty customers isn't one number. It's two: onboarding hours per new firm, times the nineteen firms still being added, plus support hours per firm per week, times all twenty firms once they're live. The two run at the same time during the rollout, onboarding for the newest firms on top of support for the ones already live, not instead of it.
O, own the numbers. For a three-person customer success team: onboarding runs about twenty hours a firm when its leases look like the pilot customer's, connecting the document system, calibrating on a sample batch, training the analysts. Support runs about two hours a week a firm once live, mostly reviewing the leases the model flags as unsure. The team has about fifty-five hours a week free for this, combined.
U, use a range. If the new firms' leases look like the pilot customer's, standard templates, normal terms, twenty hours of onboarding and two hours a week of support holds for all twenty firms. If a firm's leases run heavy on negotiated, one-off terms, onboarding can run to fifty hours and support to five hours a week. Same twenty firms, the honest range for how many the team can actually support is anywhere from about eleven to twenty, depending only on how much of the new nineteen's paperwork looks like the first firm's.
N, nail the sanity check. At the busiest point in the rollout, nineteen firms live and the twentieth still onboarding, that week needs about forty-eight hours: thirty-eight for support across the nineteen, ten for onboarding the last one. Against the team's real fifty-five hours a week, that's seven hours of slack. Tight, and it only holds if none of those nineteen turned out to be the expensive kind of firm.
D, direction. Two assumptions could move this number, and they don't move it the same amount. Whether the new firms' leases actually look like the pilot customer's swings the team's realistic ceiling from twenty firms down to about eleven, because it moves both onboarding hours and weekly support hours at once. Hiring one more support specialist adds about fifteen hours a week of capacity, which helps, but only closes part of that gap. Similarity is the bigger lever. Hiring is the smaller one, and it's the one people reach for first because it feels like doing something.

The build-up: weekly team hours as the rollout climbs to twenty firms
5 firms live20 hrs
10 firms live30 hrs
15 firms live40 hrs
19 live + 20th onboarding48 hrs
Onboarding load stays a flat ten hours a week throughout, whichever firm is going through it. It's the growing support hours underneath that push the total toward the team's real ceiling of fifty-five.
What moves the team's real ceiling (swing from a baseline of 20 firms)
New firms run heavy on negotiated, one-off leases−9
Hire one more support specialist (+15 hrs/week)+6
Slow the rollout to one new firm a month+3
Average new firm's portfolio bigger than the pilot's−2
Lease similarity swings the ceiling the most, because it moves both onboarding hours and weekly support hours at once. Hiring only moves the ceiling by a fixed amount, which is why it helps but never fully closes the gap on its own.

And if you want to be sure it really works, try it somewhere else

A veterinary group pilots an AI tool that triages after-hours phone calls, deciding which ones need an on-call vet and which can wait until morning, with one clinic as the first customer, scaling to a twelve-clinic regional group.

B, break it down. Same shape, different work. Onboarding hours per clinic, calibrating the triage thresholds to that clinic's own vets and connecting the after-hours phone line, times eleven new clinics. Support hours per clinic per week, reviewing the calls the tool flags as uncertain, times all twelve clinics live.
O, own the numbers. Onboarding runs about fifteen hours a clinic for a small-animal-only practice like the pilot clinic. Support runs about three hours a week a clinic once live. The rollout team, two people, has about forty hours a week free.
U, use a range. If the new clinics are small-animal-only like the pilot, fifteen hours onboarding and three hours a week support holds for all twelve. If a clinic also handles large animals, onboarding can run to thirty-five hours and support to seven hours a week, because a horse with colic needs a different triage judgment than a cat with a limp.
N, nail the sanity check. Twelve clinics at the low end need about thirty-six hours a week once fully live, comfortably under the forty-hour cap. At the high end, they need about eighty-four hours a week, more than double what the team has.
D, direction. Same tension as the lease tool. Whether the new clinics handle the same kind of animals as the pilot clinic swings the realistic ceiling the most. Adding a third person to the rollout team helps, but adding species the tool has never triaged before is what actually breaks the plan.

Weekly hours once all twelve clinics are live: low scenario vs high scenario
All small-animal clinics36 hrs
Some mixed-species clinics84 hrs
The team's real ceiling is forty hours a week. The low scenario clears it with room left. The high scenario blows past it before a single support call from week one even gets a second read.
Same shape, different lever At the lease tool, negotiated one-off clauses were the swing lever, and hiring closed only part of the gap. At the vet group, species mix plays the same role, mixed-species clinics need a different, harder triage judgment than the pilot clinic ever tested. The math is identical. Which assumption to check before signing isn't.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: check lease similarity before signing, don't schedule scope by the calendar or the customer count alone.
Cost: the company can only afford to hire zero more people this year. Don't cut the check, cut the pace: onboard fewer firms at once so the team's real hours stay the same size relative to the queue.
The model got better: a newer extraction model rarely misses on standard leases anymore. That doesn't remove the need to check similarity, it just moves where the risk sits, from accuracy on common leases to accuracy on the rare, negotiated ones.

Where people run it wrong.
They multiply the pilot customer's numbers straight across twenty without checking whether the next nineteen are actually like the first one.
They count hours available by job title instead of by what the team actually has free once everything else on their plate is subtracted.
They treat a new hire as the fix for a scaling problem before checking whether the real problem is which customers got signed, not how many people review them.

How to use it live. Say the equation before naming a single number: "the twenty customers aren't the real question, the hours behind each one are, so before I size the team, I'd want to know how similar the other nineteen actually look to the one we've already onboarded." That buys the room to ask a real question instead of guessing a headcount that only sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a "scale this pilot" sizing question, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question: how many onboarding and support hours it actually takes to cover twenty customers, not a person's trust switching between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Zofia Kovarik, who runs onboarding for Abstraq, a lease-abstraction AI tool. Two years in the role, Halgren Commercial Realty was the pilot customer.
3 · WHAT THE FIRST PLAN GOT WRONG
What did the first rollout plan assume that turned out to be the wrong basis for sizing the team?
Tap to flip
ANSWER
It assumed the other nineteen firms' leases would look like Halgren's clean, standard templates, so it applied Halgren's twenty-hour onboarding number to all of them instead of checking each firm first.
4 · THE STRUCTURE IN THIS STORY
What's the difference between the plan that assumed one customer profile and the plan that built in a range?
Tap to flip
ANSWER
The first plan applied one onboarding and support number to all twenty firms, no matter how different their leases actually were. The second checked each new firm's leases against Halgren's before signing, and spaced the hard ones further apart in the schedule.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Sizing the whole nineteen-firm rollout off Halgren's one data point, because Halgren was the only real number the team had, and building a plan around real data feels like the responsible move when it's all you've got.
6 · THE NUMBER
Fill in the blank: the busiest week of the rollout needs about ___ hours, against the team's real capacity of ___ hours, leaving only ___ hours of slack.
Tap to flip
ANSWER
48 hours. 55 hours. 7 hours.
7 · THE REPLAY
Same twenty firms, same team, second design. What changes?
Tap to flip
ANSWER
The same twenty firms go live, but two months later than the original promise, because the hard firms get spaced out instead of rushed. Every support flag across all twenty gets reviewed within two days, and there's no second near miss.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same sizing question for a different product. Which product, and which lever swings it most?
Tap to flip
ANSWER
A veterinary group scaling an after-hours call-triage tool from one clinic to twelve. Whether the new clinics handle the same animals as the pilot clinic swings the team's real ceiling more than adding another person does.

Check yourself Score: 0 / 0

Short answer
1. Zofia's manager says: "Just hire two more support people and we're covered for all twenty." Why doesn't hiring alone settle whether the rollout can actually reach twenty?
Show hint
Compare what hiring one more specialist changes in the D step against what lease similarity changes.
Show answer
Model answer: Hiring only moves the team's hour ceiling by a fixed amount, about fifteen hours a week per person. It doesn't touch the bigger lever, whether the new firms' leases actually look like Halgren's. That alone swings the realistic number of firms the team can support from twenty down to about eleven. Hiring narrows a gap. It doesn't close a gap caused by leases nobody checked first.
Multiple choice
2. Why does the B step use two separate numbers, onboarding hours and support hours, instead of one combined "team hours per customer" figure?
  • A. Onboarding is a one-time cost that happens once per new firm, while support is an ongoing cost that keeps happening for every live firm, and they don't grow the same way.
  • B. The model needs different settings for onboarding than for live support.
  • C. Halgren's contract required the two numbers to be reported separately.
  • D. Onboarding hours only apply to the first five firms signed.
Show hint
Think about what happens to each number as more firms go live over time.
Show answer
A. Onboarding fades once a firm is live. Support keeps running every week for as long as the firm is a customer. Collapsing them into one number hides that they behave completely differently as the rollout grows.
True or false
3. True or false: the pilot proved the model could read Halgren's leases correctly, so the same twenty-hour onboarding number should hold for all nineteen new firms.
  • True
  • False
Show hint
Ask what the pilot actually tested, and what it never got the chance to test.
Show answer
False. The pilot proved the model works on Halgren's own lease templates. It never proved anything about firms whose leases look different, like Corvath's negotiated ground leases. That's exactly the gap the range in the U step exists to cover.
Fill in the blank
4. At the busiest point in the rollout, ___ firms are live and the ___ is still onboarding. That week needs about ___ hours against the team's real ___-hour capacity.
Show hint
Check the N step's numbers directly.
Show answer
Nineteen firms; twentieth firm; 48 hours; 55-hour. That leaves seven hours of slack, which is what a plan without this check never sees coming.
Short answer, apply it yourself
5. Think of a pilot or rollout you could run at your own job, scaling from one customer or site to many. What two numbers would you check per unit, and whose real free hours would you check before promising the last one's timeline?
Show hint
Name a real onboarding-style cost and a real ongoing cost, not "we'll figure it out as we scale."
Show answer
Model answer: A regional bakery chain piloting an AI tool that forecasts daily bread orders for one store, scaling to fifteen. I'd check hours to calibrate the model on each store's own sales history, and hours a week the store manager actually has free to review flagged forecasts, not the hours listed on their job description, before promising all fifteen stores are live by a fixed date.
Short answer, the number question
6. If the average new firm's support load comes in at three and a half hours a week instead of the base case's two, how many hours a week does the team need once all twenty firms are live, and does that still fit inside the fifty-five hour cap?
Show hint
Multiply the new per-firm rate across all twenty firms, then compare to the cap.
Show answer
No, it wouldn't fit. Twenty firms times three and a half hours is seventy hours a week, fifteen hours over the fifty-five hour cap. The team would need to hire, slow the rollout so fewer firms are live and being supported at once, or find out sooner which firms are driving the higher support load.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more