ConceptIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #10

What is the right pilot length and success criteria for an AI vendor evaluation?

LEAD a junior coordinator's question about what happens for a whole year before anyone finds out

Alderwick Stage Collective is a small nonprofit consortium of five theaters. Corwin Delacroix manages the box office and set pricing by hand for years. Seatwise is the vendor pitching AI dynamic pricing to recommend seat prices show by show. Priyasha Nayar, a marketing coordinator three weeks into the job, asked the question that changed the pilot plan.

The direct answer
Run a short, bounded pilot, one venue, six to eight weeks, not twelve months across every venue. Success isn't "revenue went up," checked once at the end. It's a weekly leading signal, here the rate of season-ticket holders switching to single-ticket purchases, with a pause threshold set before the pilot starts. Revenue lift alone can look great for months while quietly pricing out the subscribers you can't win back.
Do this, in order
  1. Bound the pilot to one slice, not the whole business, for six to eight weeks.Why: a problem found on one venue in weeks costs far less than the same problem found everywhere a year later.
  2. Name a weekly leading signal, not a lagging number checked once at the end.Why: the lagging outcome can look perfectly healthy for months right up until real damage is already done.
  3. Set the pause and rollback thresholds before the pilot starts, not after something looks off.Why: a threshold decided in the moment always gets talked down by whoever wants the pilot to keep running.
  4. Check for how the vendor's own headline metric could be gamed or misleading.Why: a metric nobody's tried to game yet is just a metric nobody's checked closely.
  5. Don't apply this same caution to a tool with no real relationship at stake.Why: a simple, reversible tool doesn't need a six-week pilot with weekly thresholds; that effort should go where loyalty is actually on the line.

How to answer this, stage by stage

Nobody is scoring whether you can say "run a pilot first." They're scoring whether you can name the one number that would have warned you early, and the one that wouldn't.

Stage 1
Scope it to one real pilot decision
Say it like this
"I'll answer this for Alderwick Stage Collective, piloting Seatwise's dynamic pricing, not vendor pilots in general."
Why this works
Keeps "pilot length and success criteria" from turning into a generic project-management answer.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link to the real business outcome, find the early signal, check how it could be abused, and name the decision at each threshold."
Why this works
Shows this is a metric question, and that you know which framework actually fits it.
Stage 3
Reframe what the question is really testing
Say it like this
"This isn't really asking how long a pilot should run. It's asking whether you'll trust a number that looks good for months, or find the one that would have warned you weeks earlier."
Why this works
Separates a real answer from a generic "always pilot before you scale" line.
Stage 4
Give the one decision that matters most
Say it like this
"Run it on one venue for six to eight weeks, and watch the season-holder switch-to-single-ticket rate every week, not revenue lift checked once at the end."
Why this works
This is the direct answer, said the way you'd actually say it out loud.
Stage 5
Prove it with the near miss
Say it like this
"Revenue was up eleven percent by week six. The switch rate had also climbed to fourteen percent, and that's the one that would have hurt us for a whole season."
Why this works
Turns "watch your leading indicators" into one specific, checkable Tuesday.
Stage 6
Say what you'd still leave alone
Say it like this
"For a tool with nothing loyalty-related at stake, like a scheduling assistant, a short trial and one end number is plenty."
Why this works
Shows judgment instead of demanding a six-week weekly-tracked pilot for every tool.
Stage 7
Close on the one line
Say it like this
"Bound the pilot, name the number that moves first, and set the threshold before you need it. A great-looking revenue chart doesn't mean nothing's wrong."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

Say a nonprofit theater consortium wants an AI tool that sets ticket prices show by show, instead of a box office manager guessing at a fair tier from memory.

Before Seatwise, Corwin Delacroix set prices by hand, checking last season's attendance and picking a tier that felt fair, about three hours of work per show. Later analysis showed that manual guess left about twelve percent of possible revenue unrealized most seasons, mostly on shows that sold out faster than expected.

Hand sketched flow diagram titled Setting show prices, before Seatwise. Five boxes in sequence: Check past attendance, Guess a fair tier highlighted, Post the prices, Watch the box office, Adjust next season.
This is the whole job Seatwise is meant to sharpen. Nothing about it names a real subscriber yet.

Seatwise proposed a twelve month pilot, run across all five venues at once, with one success measure: at least ten percent revenue lift and no reported season-holder cancellations by the end of the year.

The real question was never whether revenue would rise. It was whether "no reported cancellations, checked once a year" could tell Corwin anything before the damage was already done.
Season-holder to single-ticket switch rate, week by week during the actual pilot run
20% 10% 0 pause threshold, 10% Wk 1 Wk 3 Wk 6 3% 14%
This crossed the pause threshold in week four. The vendor's own success check wasn't due for another eleven months.

At its worst, trusting the vendor's twelve-month, revenue-only plan risks running a subscriber-alienating pricing pattern across all five venues for nearly a full year before anyone found out, by which point a whole season of loyal subscribers could have quietly moved on. Watching a weekly leading signal instead cost Corwin about two extra hours a week of tracking. Against a year of unnoticed damage across the whole consortium, that trade is an easy one to accept.

The choice I would take back Alderwick's old pilot template only ever asked "did revenue go up, checked once at the end." That was fine for simple tools with nothing relational at stake. It stopped making sense the moment the tool being piloted could quietly reprice the exact seats season holders count on, with the only planned check a full year away.

What I would leave alone: a low-stakes tool, like an AI assistant that drafts internal show-schedule emails, doesn't need this level of pilot design. A short trial and one end-of-month read is genuinely enough there, since nothing loyalty-related is on the line.

The lesson: a good pilot isn't measured by how long it runs. It's measured by whether its success number would have warned you weeks before the number the vendor wanted you to wait for.

Now here is the same thing as a story

The short version above is what you'd say defending this pilot design to Alderwick's board. Read this one for how the real signal actually got found.

Corwin Delacroix had priced Alderwick's five stages by hand for eleven years, and could guess a fair tier for any show within a few dollars just from remembering how fast last year's version sold.

Hand sketched comparison titled The vendor's plan versus the real one. Left, a red-orange question mark box labeled Vendor's plan, caption 12 months, all venues, one late number. Right, a green document icon labeled The real plan, caption 8 weeks, one venue, checked weekly.
Seatwise's proposal was the left side of this picture. Alderwick shipped the right side instead.

Seatwise's sales team pitched the twelve-month, all-venue rollout as the fastest way to see a real number. Corwin nearly signed it as written. It read like due diligence: a whole year of data, a clean pass or fail at the end.

Knowledge spark: what's a leading versus a lagging signal? A lagging signal tells you what already happened, like a season-holder cancellation. A leading signal moves first, before the real damage shows up, the way a subscriber quietly switching to single tickets usually comes weeks before they decide not to renew at all.

Priyasha Nayar, three weeks into her marketing coordinator job, asked one question in the planning meeting: "What happens to our season holders if this is wrong for a whole year before we find out?" Nobody in the room had an answer.

Hand sketched quadrant titled Which signal to trust during a pilot, axes How soon it moves and How easy to game. Annual renewal rate sits low on both. Revenue lift sits low on soon, high on easy to game. Season to single ticket switch sits high on soon, low on easy to game.
The vendor's proposed metric sat in exactly the wrong corner of this map.

Corwin had considered simply asking Seatwise for weekly revenue reports instead of building anything new. He dropped that idea. Revenue reports come from the vendor's own system, and they wouldn't catch a subscriber behavior shift the vendor's own model wasn't tracking as a metric at all.

Instead, Alderwick bounded the pilot to its smallest venue and eight weeks, and Corwin added one number of his own: how many season-ticket holders bought a single ticket instead of using their season seat, checked every week, against a normal baseline of about three percent a month.

Hand sketched decision tree titled What the leading signal decides. Root, season to single ticket switch rate, checked weekly. Three branches: stays under 5 percent leads to continue the pilot, climbs past 10 percent leads to pause and inspect pricing, climbs past 15 percent leads to roll back now.
These thresholds were set before the pilot started, not argued over once the number looked bad.

By week three, the switch rate had already climbed to nine percent. By week four it crossed the ten percent pause line Corwin had set in advance. Seatwise's model, it turned out, was quietly raising prices fastest on exactly the shows season holders always claimed early, since to the model that early claiming looked like strong demand from new buyers, not loyalty from people who already had a seat.

Revenue was up eleven percent by week six. The number that mattered wasn't the one going up. It was the one climbing quietly underneath it.

Alderwick paused the pilot at week six, added a season-holder carve-out to Seatwise's pricing rules, and re-ran the same eight weeks. The switch rate settled back near its normal three percent, and revenue still rose, just nine percent instead of eleven, a smaller gain the consortium was glad to take in exchange for keeping its subscribers.

LEAD, in one screenNot a lecture on picking north star metrics. LEAD is what tells you which number would have warned you first.

L
Link. The real business outcome.
Not "did revenue rise," but "did revenue rise without costing us subscribers we can't win back."
Grounds the whole pilot in what the theater actually can't afford to lose.
E
Early signal. What moves first.
The season-holder to single-ticket switch rate, which climbed weeks before any cancellation would have shown up at renewal.
This is the hardest step, and the whole answer to the question turns on it.
A
Abuse. How the metric gets gamed.
Revenue lift alone can look great for months while a pricing pattern quietly prices out the exact subscribers the theater depends on.
Every metric has a way to be hit without doing the real work, and revenue lift was this pilot's easiest one.
D
Decision. What you'd actually do at each threshold.
Under five percent, continue. Past ten, pause and inspect. Past fifteen, roll back, decided before the pilot ever started.
A metric nobody's committed to acting on ahead of time is a dashboard decoration.
Hand sketched labeled parts diagram titled What a real pilot success criterion needs. A document icon at center labeled Pilot Plan, with four callouts: one bounded slice, a weekly leading signal, a pause threshold, a named decision date.
Seatwise's original plan had none of these four. That absence was the whole warning.

The recap, one line per letter: link is naming what the theater actually can't afford to lose, early signal is the season-holder switch rate moving weeks before renewal data ever could, abuse is revenue lift looking healthy while that switch rate climbs underneath it, and decision is the pause and rollback thresholds set before the pilot ever started.

Hand sketched timeline titled A bounded pilot, on paper, week 6 emphasized. Week 1, caption pilot starts, one venue. Week 3, caption leading signal checked. Week 6, caption threshold crossed, pause. Week 8, caption go, no-go decision.
Eight weeks, one venue, a real decision at the end. Not twelve months across everything Alderwick runs.

And if you want to be sure it really works, try it somewhere elseSame four letters, a funeral home network instead of a theater consortium. A different early signal breaks the second pilot.

Grayfield Funeral Alliance piloted Rememberly, a vendor offering AI-drafted tribute page content for families to personalize. The vendor's proposed success measure was "fewer than five formal complaints over six months." Mapped onto LEAD: link is the real outcome, families feeling the tribute page truly represented their person, not merely the absence of a complaint. Early signal is the rate of families requesting a full manual rewrite instead of lightly editing the AI draft, which climbed weeks before any family would file a formal complaint, since most grieving families quietly redo the page rather than complain about it. Abuse is that "complaints" as a metric rewards families staying silent, not the tool actually working. Decision is treating a manual-rewrite rate above twenty percent in any two-week window as an automatic pause, set before the pilot began.

Hand sketched icon list titled Signs a tribute page needs a full rewrite. Details feel generic. Family requests a full redo. Tone doesn't match the person. No personal story survived.
Grayfield's early signal wasn't a number on a dashboard. It was families quietly asking for a redo.
Revenue lift versus the leading signal, both measured at the same week six
20% 10% 0 +11% Revenue lift (headline) +14 pts Switch rate (leading)
The headline number looked better than the pilot actually was. The leading signal was already flashing red at the exact same moment.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "bound the pilot, name the number that moves before the damage does, and set the pause line up front," and stop.
Cost: there's no staff time to track a new weekly number by hand. Say so honestly, and pull the smallest possible proxy from data you already have, rather than skipping the leading signal entirely.
The model gets better, for real: if Seatwise's next version genuinely stops over-pricing season-holder seats, that's still worth a fresh short pilot, since a fixed model deserves the same scrutiny as a first one.

Where people run it wrong.
They accept a vendor's proposed pilot length and success metric without asking who wrote it and what it's convenient for them to measure.
They pick a single lagging number and wait the whole pilot to check it, instead of watching something weekly.
They never set the pause threshold in advance, so it gets negotiated away in the moment by whoever wants the pilot to keep running.

How to use it live. The moment someone proposes a pilot's length and success metric, ask: what would move first if this goes wrong, and can we watch that instead of waiting for the number that only shows up at the end? Let that answer set the real plan.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a metric question like designing pilot success criteria?
Tap to flip
ANSWER
LEAD: link to the real outcome, find the early signal, check for abuse, name the decision at each threshold.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Corwin Delacroix, Alderwick Stage Collective's box office manager, who priced five stages by hand for eleven years.
3 · THE SHARP QUESTION
What question reframed the whole pilot design?
Tap to flip
ANSWER
"What happens to our season holders if this is wrong for a whole year before we find out?" Nobody had an answer.
4 · THE EARLY SIGNAL
What was the actual leading indicator in this story?
Tap to flip
ANSWER
The rate of season-ticket holders buying single tickets instead of using their season seat, which climbed weeks before any cancellation could show up.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Accepting a twelve-month, all-venue pilot with revenue lift as the only success check, a plan that suited the vendor more than the subscribers.
6 · THE NUMBER
Fill in the blank: by week four, the season-holder switch rate crossed the ___ percent pause threshold Corwin had set in advance.
Tap to flip
ANSWER
10 percent, set before the pilot started, so nobody had to argue about it in the moment.
7 · THE REPLAY
Same eight-week pilot, but with a season-holder pricing carve-out from the start. What changes?
Tap to flip
ANSWER
The switch rate stays near its normal 3 percent, and revenue still rises, 9 percent instead of 11, a smaller gain worth keeping the subscribers for.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different organization. Which one, and what's its early signal?
Tap to flip
ANSWER
Grayfield Funeral Alliance, piloting Rememberly. Its early signal is the rate of families requesting a full manual rewrite instead of a formal complaint.

Check yourself Score: 0 / 0

Short answer, name the position
1. In one sentence, what is the actual right pilot length and success criterion this answer lands on?
Show hint
Look at the direct answer.
Show answer
Model answer: A short, bounded pilot, one venue for six to eight weeks, judged by a weekly leading signal with pre-set pause thresholds, not a year-long revenue-only check.
Multiple choice
2. Why was "revenue lift, checked once at the end of a year" a weak success criterion here?
  • A. Revenue is never a real business outcome.
  • B. It could look healthy for months while a leading problem, subscriber pricing risk, went unnoticed.
  • C. Nonprofits aren't allowed to track revenue.
  • D. Seatwise refused to report revenue at all.
Show hint
Look at the A step, abuse, in the LEAD recap.
Show answer
B. Revenue lift alone rewards the wrong thing early, hiding the real risk until it's already done damage.
True or false
3. True or false: by the time the pilot was paused, revenue had already dropped, which is what caught the team's attention.
  • True
  • False
Show hint
Look at the highlight block about revenue being up eleven percent.
Show answer
False. Revenue was up 11 percent, looking great. The switch rate, a separate leading signal, is what actually caught the problem.
Short answer, where it wouldn't matter
4. Name a kind of tool where this whole six-week, weekly-tracked pilot design would be overkill.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A low-stakes tool like an AI assistant drafting internal emails, where nothing relational or loyalty-based is on the line.
Short answer, apply it yourself
5. Think of a product or service you've used that changed slowly for the worse. What would have been the early signal, versus the lagging one everyone actually noticed?
Show hint
Ask what small behavior of yours changed weeks before you actually complained or left.
Show answer
Model answer: Usually a small workaround or reduced usage shows up first, well before an actual complaint or cancellation.
Short answer, the number question
6. If the switch rate had stayed at 6 percent through week eight instead of climbing to 14, would the pilot still have been worth pausing to inspect? Why or why not?
Show hint
Look at the decision tree's threshold levels.
Show answer
Model answer: No, not automatically. Six percent stays under the ten percent pause line, so the pilot would continue, with the number still watched weekly.
Before you close the answer
Why this works
Tests whether you'll accept a vendor's proposed pilot length and metric at face value, or find the number that would actually warn you before real damage is done.
Follow-up traps
"Isn't a six-week pilot too short to trust for a whole year of pricing decisions?" Response: the pilot length isn't meant to prove a full year of performance, it's meant to catch the fastest-moving risk early enough to fix before scaling further.

"What if a leading signal like this doesn't exist for every pilot?" Response: then that's the first thing to find before the pilot starts, not something to skip; if nothing moves before the lagging outcome does, that's worth saying plainly too.
If pressed
The season-holder carve-out Alderwick added wasn't a blanket price freeze, it capped how much Seatwise's model could move price on any seat tied to a season pass within a single week, leaving the model free to price everything else normally.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more