ConceptIntermediateQuality, Cost & Token Economics / Pricing AI products: seat, usage, outcome / #2

Why does seat-based pricing break when AI reduces the number of seats needed?

Scoutwire got better at writing and sending outreach on its own. The invoice did not celebrate that. It shrank, because the invoice was built around counting the humans the tool was busy making unnecessary.

The direct answer
Seat-based pricing breaks because it charges for the exact resource an AI outreach tool is built to shrink: the human reps needed to hit a target number of meetings. As the AI gets better, customers correctly need fewer seats to reach the same result, so the number on the invoice falls even while the value delivered holds steady or grows. Move the metric you charge on off headcount and onto something that scales with what the AI actually produces, like meetings booked or verified replies, so revenue and value move in the same direction instead of opposite ones.
Do this, in order
  1. Move the pricing metric off seats and onto what Scoutwire actually produces.Why: seats is the exact resource the product is built to shrink, so the harder it works, the smaller the bill, unless the metric tracks output instead of headcount.
  2. Trace the seat drop with an evidence test before rebuilding anything.Why: confirm the cut tracks with automation adoption, not a company-wide hiring freeze or real dissatisfaction, or the fix targets the wrong cause entirely.
  3. Rule out the boring explanations first, a billing error or a broad layoff wave.Why: both can shrink a seat count exactly the way automation does, and chasing the wrong one wastes a quarter.
  4. Recut retention by adoption depth, not one blended number for the whole book.Why: a healthy-looking company-wide average can sit directly on top of one heavy-adopter segment collapsing underneath it.
  5. Keep a seat-priced option alive for accounts still running human review on every send.Why: not every customer runs the tool in autonomous mode, so the fix shouldn't force everyone onto the same new metric on day one.
  6. Track the pairing between automation-adoption dates and renewal seat counts going forward.Why: the same collapse can happen again the next time a feature quietly does more of a person's job.

How to answer this, stage by stage

Nobody is grading whether you know the phrase "seat-based pricing doesn't scale with AI." They're grading whether you can find the actual mechanism, not just point at the symptom.

1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one product. Scoutwire is the AI outreach agent inside Farrowmede, an outbound sales platform: it finds prospects, writes the outreach, and sends the sequence, so a sales team needs fewer human reps to hit the same number of booked meetings. Rafi Vantrell is the finance partner who supports Farrowmede's sales org and watches its renewal numbers every quarter."
Why this works
Keeps the answer attached to a real invoice and a real number instead of a general claim that AI pricing is hard.
2
Reframe what "seat-based pricing breaks" actually means
Say it like this
"This isn't a normal churn question. Nobody is unhappy and nobody is leaving. The problem is that Farrowmede charges per SDR seat, and Scoutwire's whole job is to need fewer SDR seats. So the better the product gets, the smaller the number on the invoice gets, on purpose, by design."
Why this works
Separates a candidate who blames the customer from one who blames the pricing metric, which is where the real judgment lives.
3
Give the one decision, plainly
Say it like this
"Stop billing on seats and start billing on something that grows when Scoutwire works: meetings booked, verified replies, or contacted accounts. Whatever it is, pick a number that goes up when the product succeeds, not a number the product's whole job is to bring down."
Why this works
This is the actual fix, said as a decision, not a category like "rethink the pricing model."
4
Rule out the boring explanations before blaming the product
Say it like this
"Before I trust that automation caused the seat cuts, I check two things. One, did these same accounts cut seats in unrelated software too, which would mean a company-wide hiring freeze, not something about Scoutwire. Two, did meetings booked at those accounts actually fall along with the seats, which would mean real dissatisfaction instead of automation doing the job."
Why this works
Shows you know a falling seat count can look identical whether the cause is automation, a layoff wave, or a customer quietly checking out.
5
Prove it with the finding, cut to four sentences
Say it like this
"Here's what Farrowmede actually found. Forty one heavy-adopter accounts cut Scoutwire seats by 34 percent at renewal. Meetings booked across those same accounts rose 6 percent over the same stretch, and only 3 of the 41 also cut seats in their separate CRM contract. That rules out a broad freeze and rules out dissatisfaction in the same test."
Why this works
One test, two hypotheses eliminated, and a number that shows the product working, not failing.
6
Say what you'd measure and act on going forward
Say it like this
"I'd stop watching one blended retention number and split it by how deep an account runs in autonomous mode, every quarter, not just once. I'd keep the meetings-booked and cross-vendor seat checks as standing signals, so the next time a feature quietly does more of someone's job, we catch it at the segment, not two quarters later inside a blended average."
Why this works
Shows you're building a way to catch the next version of this, not just explaining the one that already happened.
7
Close on the decision, not the story
Say it like this
"So: don't bill on the thing your AI is built to make smaller. Find what the AI actually produces, and charge on that instead. Otherwise your best product quarter and your worst revenue quarter end up being the same quarter."
Why this works
Ends on a rule a candidate can reuse any time a question involves AI cutting into headcount, not just this one story.

Let's learn

Hand sketched comparison titled Charging for the thing built to disappear. Left figure, a person icon in red, labeled Seats, caption the thing on the invoice. Right icon, a document, labeled Meetings booked, caption the thing Scoutwire grows.
One side of this invoice was never supposed to grow forever. The other side was.

For three years, a sales team's price with Farrowmede grew exactly the way its headcount did. Add a fifth SDR, pay for a fifth seat. That was the whole model, and it made sense, because a seat was a person logging in every day to send outreach by hand.

Scoutwire is the part of Farrowmede that finds prospects, writes a personal opening line for each one, and sends the outreach sequence. For its first two years it worked like a very fast assistant: it drafted the message, and a human still had to click send. A rep with Scoutwire could carry a bigger list, but the team still needed a rep for every few hundred accounts.

Under that model, adding Scoutwire licenses tracked headcount closely. A twelve person SDR team paid for twelve seats, sent about 4,800 emails a month between them, and booked around 190 meetings. Revenue and team size moved together, quarter after quarter.

Hand sketched comparison titled Seats before and after Auto-send. Left panel, a person icon in blue, labeled Before, six SDR seats, human approves every send. Right panel, a person icon in red, labeled After, two SDR seats, same meetings booked.
Same team, same result on the board. Four fewer humans needed to get there.

Then Scoutwire shipped Auto-send: for sequences the model is confident about, it writes and sends without waiting on a human click. A single rep running Auto-send oversees far more accounts than one who has to approve every message by hand. Teams running mostly on Auto-send started hitting the same number of booked meetings with a third of the reps.

Net revenue retention: blended vs by adoption depth, week 20
100% = flat 130% 0 Blended (97%) Light adopters (106%) Heavy adopters (81%)
Blended, whole bookLight adoptersHeavy adopters
Blended retention looked only mildly soft at 97 percent. Split by how deep an account ran on Auto-send, light adopters barely moved while heavy adopters alone had fallen to 81 percent, the segment hiding inside the average.

Here's the part that isn't a mistake. The dropped seat count isn't a sign the product broke. It's a sign the product is finally doing the job it was built to do: replacing hours of manual outreach with a system that runs itself. The problem was never the product working. It was that Farrowmede was charging for the exact thing Scoutwire was designed to make smaller.

The tool did not get worse. It got better, and that is exactly what shrank the invoice.

What that costs at its worst: if nobody catches this, a team's best quarter for the product becomes its worst quarter for revenue. A company that doesn't trace the seat drop back to automation, and instead assumes customers are unhappy, can panic-throttle Auto-send to protect the old pricing, which means slowing down the exact feature that makes Scoutwire worth buying, to prop up a metric that should never have been the thing on the invoice.

The decision that mattered Stop billing on seats. Bill on meetings booked or verified replies instead, something that grows when Scoutwire works, not something the product is built to shrink.

What I'd leave alone: accounts in regulated industries, like healthcare or financial outreach, that keep a human reviewing every send by policy, not preference. Seat pricing still tracks their real usage fine, because a person still has to be in the loop for every message. Don't rebuild pricing for accounts where the flip never happened.

The lesson: a pricing metric is a bet about what will stay scarce. Seats were scarce back when a person had to read and approve every message. The bet stopped paying off the moment the product's whole purpose became making that scarcity go away. Before pricing anything on a resource, ask whether the product's own roadmap is trying to make that resource disappear.

Now here is the same thing as a story

Read the short version above for the two minute answer. Read this for why a shrinking seat count looked exactly like trouble, right up until someone actually tested it.

Rafi Vantrell had supported Farrowmede's sales org for four years, the kind of finance partner who could recite a segment's net revenue retention from memory before anyone else in the room found the spreadsheet. For most of that time the story was simple: sales teams that liked Scoutwire added seats, teams that didn't shrank, and NRR sat comfortably above 108 percent for two straight years.

Knowledge spark: what is net revenue retention? A number that tracks how an existing group of customers' spending changes over a year, with no new customers counted. Above 100 percent means the same customers are paying more. Below 100 means they're paying less, even if none of them left.

Then Auto-send shipped, quietly, as an opt-in setting for sequences the model was confident about. Adoption crept up for six weeks before anyone thought to watch it as its own number. By week ten, blended NRR had drifted from 108 percent to 104, a soft dip nobody flagged, the kind of wobble a metric shows most quarters.

Hand sketched timeline titled When the invoice stopped matching the product. Four points: Auto send ships week zero opt in only. Heavy use crosses seventy percent week six nobody splits NRR by this. Blended NRR slips week ten one oh eight to one oh four still looks fine. Segment found cratering week twenty eighty one hiding inside ninety seven, marked in red.
The exec view only ever showed one blended number. It stayed calm the whole time one segment was quietly cratering underneath it.

It kept drifting. By week twenty, blended NRR sat at 97 percent, the first time in two years the number had dropped under 100. Rafi's first instinct, and the room's first instinct, was the obvious one: something's wrong, customers are leaving, find out who.

Hand sketched flow diagram titled How the drop hid inside one blended number. Four boxes connected by arrows: Auto send ships. Blended NRR looks fine. Heavy adopters cut seats. No split by adoption depth, outlined in red.
Every one of these four steps looked fine on its own. Nobody was watching the one split that would have caught it.

The obvious read didn't survive five minutes with the account list. Almost nobody had cancelled. The accounts driving the drop were still customers, still renewing, just renewing with fewer Scoutwire seats than the quarter before. That ruled out plain churn, and it left three real suspects on the table: a company-wide hiring freeze that happened to hit sales hardest, real dissatisfaction with a product that was quietly getting worse, or Scoutwire simply doing more of the job a rep used to do.

Hand sketched icon list titled Three suspects behind the seat cuts. One, headwinds, ruled out, CRM seats held steady, grey icon. Two, dissatisfaction, ruled out, meetings booked rose, grey icon. Three, automation did the job, confirmed, red icon.
All three looked equally plausible before the evidence test. Only one survived it.

Rafi ran the test that told them apart. For the same forty one heavy-adopter accounts, she pulled two more numbers: meetings booked, and whether those accounts had also cut seats in their unrelated CRM contract during the same renewal window. A company-wide freeze would show up in both tools, not just one. Real dissatisfaction would show meetings booked falling right alongside the seats.

It was never a shrinking team. It was a team that finally didn't need as many hands on the keyboard to hit the same number.

Neither showed up. Only 3 of the 41 accounts had touched their CRM seat count at all. Meetings booked across the group rose 6 percent, even as SDR seats fell 34 percent. The accounts weren't struggling and they weren't leaving. Auto-send was doing outreach that used to need a person, and the invoice was the only thing in the building that hadn't caught up.

Seats vs meetings booked, heavy adopters, indexed to week 0 (twenty weeks)
110 60 week 0: 100 (both) week 20: seats 66 week 20: meetings 106 Week 0 Week 10 Week 20
Seats, indexed to week 0 = 100Meetings booked, indexed to week 0 = 100
Seats fell 34 percent over twenty weeks. Meetings booked rose 6 percent over the same twenty weeks, on the same accounts. Same window, opposite directions.

The old decision that opened the door went back three years, to the pricing meeting where Farrowmede decided the simplest number to bill on was the number of people logged into the tool. Nobody in that room was wrong. A seat was a real, scarce, countable thing back then, because a human had to sit in it to get any outreach sent at all.

Run the story forward with the fix already in place. Farrowmede re-prices the heavy-adopter tier on meetings booked instead of seats, with a small base platform fee underneath it. Over the next two quarters, blended NRR climbs back to 109 percent, not because Auto-send slowed down, but because the number on the invoice finally tracks the number that was already going up the whole time.

What I'd tell myself, back in that pricing meeting three years ago: a metric that's easy to bill on isn't the same as a metric that's safe to bill on forever. Seats were easy. They were never going to stay safe once the product's job became doing without them.

TRACE, and the one test that told automation from attrition

Not a checklist for a renewal review. Five moves that build toward the one that actually separates a real cause from a coincidence: the evidence test.

TTimeline. When did the seat drop actually start, and what shipped near that date?
Auto-send shipped in week zero as an opt-in setting. Adoption crossed 70 percent of send volume at the heaviest accounts by week six. Blended NRR held near 108 percent through week eight, then drifted quietly down to 97 percent by week twenty, the point where anyone actually noticed.
A timeline that starts at week twenty starts fourteen weeks too late. The real start is the release that read as a pure quality upgrade.
RRecut. Slice retention by adoption depth, not the whole book at once.
Blended NRR across Farrowmede's book sat at 97 percent, soft but not alarming. Recut by how much of an account's send volume ran through Auto-send, light adopters held at 106 percent, almost unchanged. Heavy adopters, over 70 percent of sends on Auto-send, had cratered to 81 percent.
A blended number sitting at 97 can hide one real segment sitting at 81 underneath a much larger, healthy one.
AAssume nothing. Rule out a billing artifact or an unrelated cost freeze before trusting the correlation.
Before trusting that Auto-send explained the seat cuts, Rafi checked two things: whether the accounts cutting Scoutwire seats had also cut seats in their separate, unrelated CRM contract during the same renewal window, and whether billing had actually processed correctly rather than defaulting to a lower tier by mistake. The billing system checked out clean. Only 3 of the 41 heavy-adopter accounts had touched their CRM seat count at all.
A seat count falling for the wrong reason looks identical to a seat count falling for the right one, on a chart alone.
CCause candidates. Name three, not everything possible.
One, a company-wide hiring freeze cut sales headcount for reasons that have nothing to do with Scoutwire. Two, real dissatisfaction, accounts quietly unhappy with output quality and scaling back before churning outright. Three, Auto-send doing outreach that used to require a person, so the same results need fewer humans logged in.
Three named suspects, not a shrug. Most pricing post-mortems jump straight to the first one someone guesses.
EEvidence test. The one check that told the suspects apart.
Across the 41 heavy-adopter accounts, meetings booked rose 6 percent over the same twenty weeks that Scoutwire seats fell 34 percent, and only 3 of those accounts had cut seats anywhere else in their software stack. A hiring freeze would have shown up broadly across vendors. Dissatisfaction would have dragged meetings booked down along with the seats. Neither did. Output went up while headcount went down.
This is the single strongest move in the whole method. It turned a scary-looking seat cut into proof the product was working exactly as designed.

Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was raising the price per seat at the next renewal to make up the shortfall. It lost because heavy-adopter accounts had already cut down to two or three seats, so a higher per-seat price barely touched their bill, while it punished light adopters still ramping up who hadn't earned the automation discount yet, and it did nothing to fix the mechanism causing the drop. The AI-specific failure worth naming is silent quality drift inside Auto-send itself: because nobody reviews a send before it goes out, a bad batch of personalization or a wrong contact match can hurt reply rates for weeks with no error message anywhere, since a sent email that gets ignored looks identical to a sent email that landed well. The guardrail is a confidence floor on which sequences qualify for Auto-send, plus a weekly spot-check sample of auto-sent messages graded for quality, so a quiet drop in message quality shows up before reply rates do. And the trade being accepted on purpose is real: pricing on meetings booked ties Farrowmede's revenue to Scoutwire's own performance, so a bad model quarter now costs twice, once in product quality and again in revenue that automatically falls with it, which is the price of a metric that finally moves with adoption instead of against it.

The five, in one line each:
T: the real start is the Auto-send release that shipped ten weeks before anyone noticed the blended number slip.
R: a flat blended average can sit right on top of one adoption-depth segment losing a real quarter of its seats.
A: rule out a cross-vendor freeze or a billing artifact before trusting that automation explains the drop.
C: name three real suspects, headwinds, dissatisfaction, and automation, never jump to the first guess.
E: one clean check, meetings booked and cross-vendor seats, tells you what's real and what's a coincidence.

Same five moves, a claims desk instead of a sales floor

Not every seat-collapse question is about a sales team. The same test works anywhere a company bills per person and then builds an AI that needs fewer of them.

Nettlecombe is a claims processing firm that handles overflow claims work for mid-size insurers. Ledgerlark is the AI copilot built into its claims desk: for straightforward claims, it drafts the full adjustment decision and cites the policy language behind it, leaving a licensed adjuster to spot-check rather than write the decision from scratch. Sefton Quilling is the pricing analyst who noticed Nettlecombe's per-seat adjuster licenses sliding at renewal.

Nettlecombe bills per licensed adjuster seat, the same shape as Farrowmede's model. Once Ledgerlark's auto-draft feature crossed 70 percent of low-complexity claims, adjuster seat counts at the heaviest-using client accounts fell 31 percent at renewal. The naive read looked like churn risk. It wasn't.

Hand sketched labeled parts diagram titled What Nettlecombe's evidence check found. Center, a document icon labeled Adjuster seats cut thirty one percent. Four labeled callouts: Claims closed up eight percent. CRM seats unchanged. Complexity mix unchanged. Ledgerlark drafts seventy four percent of claims.
Same test, a different desk. Every callout points the same direction: automation, not attrition.
The evidence test that settled it Sefton's team checked the same two things Rafi did. Claims closed per month for the affected accounts rose 8 percent, not fell, and the mix of claim complexity hadn't shifted, so this wasn't just easier claims arriving. Ledgerlark was auto-drafting 74 percent of claims for those accounts, and fewer adjusters were producing more finished claims, not less.

Same method, different shape: a sales outreach agent and a claims-drafting copilot look nothing alike, but both are the same TRACE move: don't trust the naive seat drop, recut by adoption depth, then run one test that separates a real cause from a coincidence.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the one line: don't bill on the thing your AI is built to shrink, find what it actually produces, and charge on that instead.
Cost: there's no budget this quarter to rebuild the whole pricing system. Re-price only the heaviest-adopter tier first, where the mismatch is worst, and hold the rest of the book on seats until the evidence test says otherwise.
The model got better, for real: say the next model version books even more meetings per rep. That's exactly when the seat count falls fastest, right as the metric on the invoice looks like the healthiest thing about the whole account.

Where people run it wrong.
They see seats falling and ship a win-back campaign aimed at unhappy customers who were never actually unhappy.
They redesign the whole pricing model company-wide before checking whether the drop is real or a billing artifact.
They raise the seat price to cover the shortfall, which punishes the accounts still ramping up while barely touching the ones that already automated past it.

How to use it live. Say the two things a seat-pricing question has to get right before saying anything else: that a falling seat count on an AI product can mean the product is working, not failing, and that the fix is to bill on what the AI produces instead of what it replaces. That buys a beat of thinking time, and it tells the interviewer you know most candidates stop at "we're losing customers."

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Built for diagnosis questions, telling a real cause apart from a coincidence sitting on top of it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Rafi Vantrell, the finance partner who supports Farrowmede's sales org and watches its renewal numbers every quarter.
3 · THE HABIT
What did Farrowmede stop watching that let the drop hide for ten weeks?
Tap to flip
ANSWER
They watched one blended net revenue retention number for the whole book, never split by how deep an account ran on Auto-send, so a heavy-adopter segment could crater while the average still looked only slightly soft.
4 · THE CONFOUND
What's the confound this whole story turns on?
Tap to flip
ANSWER
A company-wide hiring freeze would shrink Scoutwire seats and look exactly like automation on a simple seat-count chart. Checking unrelated CRM seats and meetings booked is what tells the two apart.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Pricing Scoutwire per SDR seat at launch. That made sense when a human had to operate the tool by hand. It stopped being safe the day Auto-send let the product operate itself.
6 · THE NUMBER
Fill in the blank: Scoutwire seats fell 34 percent at the 41 heavy-adopter accounts, while meetings booked across those same accounts rose ___ percent.
Tap to flip
ANSWER
6 percent. That's the number that rules out dissatisfaction, since a shrinking, unhappy account would show meetings booked falling too, not rising.
7 · THE REPLAY
Same drift, new pricing, what changes?
Tap to flip
ANSWER
Farrowmede re-prices the heavy-adopter tier on meetings booked instead of seats. Blended NRR climbs from 97 percent back to 109 percent over the next two quarters, because the invoice finally tracks the number that was already going up.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what did the evidence test find there?
Tap to flip
ANSWER
Nettlecombe, a claims processing firm, and its AI claims copilot Ledgerlark. Adjuster seats fell 31 percent while claims closed rose 8 percent, the same TRACE test confirming automation, not churn.

Check yourself Score: 0 / 0

True or false
1. True or false: because Farrowmede's heavy-adopter accounts cut Scoutwire seats by 34 percent, that alone proves those customers were unhappy with the product.
  • True
  • False
Show hint
Look at what happened to meetings booked at those same accounts over the same stretch.
Show answer
False. Meetings booked across those accounts rose 6 percent over the same twenty weeks, and only 3 of 41 also cut seats in an unrelated tool. A shrinking, unhappy account looks different from this.
Multiple choice
2. Farrowmede's first instinct was to raise the price per seat at the next renewal to cover the shortfall. Why does that not fix the real problem?
  • A. Heavy adopters had already cut down to two or three seats, so a higher price barely touches their bill, while it punishes light adopters still ramping up.
  • B. Farrowmede's contracts legally prevent any price increase.
  • C. Scoutwire cannot track how many seats an account is using.
  • D. Raising software prices is against most states' commerce laws.
Show hint
Think about how many seats a heavy-adopter account actually has left to raise the price on.
Show answer
A. A price increase on a handful of remaining seats raises little revenue where the mismatch is worst, and it lands hardest on light adopters who haven't automated yet, doing nothing to fix the actual mechanism.
Fill in the blank
3. Fill in the blank: of the 41 heavy-adopter accounts that cut Scoutwire seats, only ___ also cut seats in their unrelated CRM contract during the same renewal window.
Show hint
The number appears in the "assume nothing" step of the TRACE recap.
Show answer
3 (about 7 percent). If a company-wide hiring freeze were the real cause, that number would be much closer to 41, since a freeze hits every vendor's seat count, not just one.
Short answer, name the rejected alternative
4. What alternative did Rafi's team consider and reject once the heavy-adopter segment's NRR cratered to 81 percent?
Show hint
Look at the paragraph right after the evidence test step in the TRACE recap.
Show answer
Model answer: They considered raising the per-seat price at renewal to cover the shortfall. It lost because the heaviest-cutting accounts had already shrunk to a handful of seats, so a price increase would barely touch their bill while punishing light adopters who hadn't automated yet, and it did nothing to fix the mismatch between what's billed and what the product actually shrinks.
Multiple choice
5. At Nettlecombe, adjuster seats fell 31 percent at the heaviest Ledgerlark-using accounts, and claims closed per month rose 8 percent over the same window. What does that combination tell you?
  • A. Ledgerlark is producing lower-quality claim decisions that adjusters have to redo.
  • B. The auto-drafting feature is doing work that used to need a licensed adjuster, the same pattern found at Farrowmede.
  • C. Nettlecombe's clients are sending fewer claims overall.
  • D. The seat drop is unrelated to Ledgerlark and should be ignored.
Show hint
Compare this to what the evidence test found at Farrowmede: output rising while headcount falls.
Show answer
B. Fewer adjusters produced more finished claims, with complexity mix unchanged, which is the same automation signature the evidence test found at Farrowmede, not a quality problem or a demand drop.
Short answer, apply it yourself
6. Pick a product you use, or a team you know, that's priced or staffed by headcount. If an AI feature there got good enough to need fewer people, what evidence would tell you that was a healthy sign instead of a warning sign?
Show hint
Think about what you'd check that would look different if people were leaving unhappy versus if the AI were simply doing more of the work.
Show answer
Model answer: A software agency that bills clients by the number of consultants staffed on an account. If an AI research tool let the team use two consultants instead of four, I'd check whether project output, deliverables shipped, client satisfaction scores, held steady or improved, and whether the client's spend on other parts of the firm stayed flat. Output holding up while headcount falls points to automation working, not the client losing confidence.
Before you close the answer
Why this works
Tests whether you'll blame the customer, or the pricing metric, when the number on an invoice starts falling right as a feature gets good. Most candidates jump straight to "we're losing accounts" and never check whether the accounts are actually fine.
Follow-up traps
"Couldn't you just cap how much Auto-send is allowed to reduce seats, to protect revenue?" Response: that throttles the exact feature that makes Scoutwire worth buying, on purpose, to protect a metric that was never the right thing to bill on. It buys time, it doesn't fix the mismatch.

"Isn't billing on meetings booked risky, since Farrowmede doesn't fully control whether a meeting happens?" Response: yes, and that's a real trade being accepted, not ignored. It ties revenue to the model's own performance, so a bad model quarter now costs twice. That's still better than a metric that falls automatically every time the product improves.
If pressed
The evidence test only worked because Rafi picked a control the confound itself couldn't explain away, an unrelated CRM contract with a completely different vendor's billing, so a hiring freeze would have shown up there too. Running the check only inside Farrowmede's own data would have left the freeze hypothesis unfalsifiable.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more