Explain how you would set a target for a metric with no historical baseline.
A number nobody has ever measured before does not get a target from the best footage you own. It gets one built from three real numbers and an honest gap between them.
- Triangulate the target from three real numbers instead of one demo reel.Why: one confident figure is a guess wearing a suit. Three real inputs is an estimate.
- Measure the manual baseline fresh, from the customer's own records, not from memory.Why: it is the only number in the mix nobody can argue with, because it already happened.
- Pull a real industry benchmark from full scale field data, not a curated pilot.Why: a curated pilot only ever shows the cases someone already knew were bad.
- Set a floor from the smallest improvement that would actually change what a person does next.Why: below that floor the number is real but nobody can act on it.
- Report a range, not a point, and put a working number in the middle of it.Why: false precision fools nobody who checks it twice.
- Test the range against what a skeptical exec already expects before it lands in a deck.Why: a target that only survives inside the building is not ready to leave it.
How to answer this, stage by stage
Nobody is grading whether you can say the word "estimate." They are grading whether you can turn a great demo number into an honest one, with the arithmetic shown. Six moves get you there.
Let's learn
What happens when a company has to promise a number for something nobody has ever measured before, on this farm, with this camera, on this cow?
Grazeline sells Sentry, a system of barn cameras and a few wearable sensors that watches a herd around the clock and flags a cow that might be sick or hurt, days before a person walking the barn would catch it by eye. Farriday Dairy, a hundred and twenty cows, is about to sign as Grazeline's very first paying customer.
The board's question sounds simple: how many days early will Sentry catch a lame cow, compared to a farmhand? Simple, except nobody has ever run Sentry on a real, paying customer's barn before. Every number Grazeline has came from somewhere else.
Here is the turn. The most tempting number in the whole company was the nine days from Grazeline's own research pilot. It was real, it was measured, and it made for a great slide. Say this plainly: nine days was never a lie, it was just a lie of context. That pilot ran on a different farm, and the vet team had already picked the clearest, most obvious cases for the camera to watch. The number was true. The situation it described does not exist at Farriday.
At its worst, this costs more than an awkward board slide. If Grazeline had promised nine days and Sentry only delivered two or three in Farriday's actual barn, the ninety day renewal review would open with a broken promise instead of a working product, and the very first paying customer Grazeline ever had would be the one who found out the hard way.
The choice I would take back is that Grazeline almost let that nine day number travel from a pilot slide into a real contract, before anyone asked where it came from. It made sense in the moment. It was the only number the company had, and it made for a strong pitch. Once Selke pulled the vet log and the industry trials, it became clear the honest number was less than half of that.
What I would leave alone: Sentry's injury alert, a separate feature that texts a farmhand when a cow stops moving normally after a fall. That one already has a clean baseline, because Grazeline's early customers have been rating those alerts by hand for a year. No triangulating needed there. Only the lameness lead time, the metric nobody has ever measured in a live paying barn, needs this whole exercise.
The lesson: a brand new number does not get to borrow confidence from an old one just because they sound similar. Build the target from three real inputs, then let a range, not a lucky guess, do the talking.
Now here is the same thing as a story
The short version sits above. Read on for the Tuesday this almost went out with the wrong number attached.
Selke Nkemelu has run product analytics at Grazeline for two years, most of it spent on a system that, until three weeks ago, had never watched a single cow that belonged to a paying customer. Every number in the company's own slide deck came from a research pilot on a university teaching herd, forty miles from Farriday.
The Farriday contract closed on a Tuesday morning. By Tuesday afternoon, Grazeline's head of sales, riding the win, forwarded the pilot's numbers to Farriday's owner as "what to expect": nine days of early warning, on average. Nobody on the sales side had built that number. It was just the best one sitting in a slide from eight months earlier.
Selke caught the email in a shared inbox before it went any further. She didn't argue about the number being wrong exactly. She argued about what it was actually describing.
She spent the next two days doing the thing nobody had done yet: pulling Farriday's own vet log, forty confirmed lameness cases over the past year, and asking the vet to walk back through each one and estimate, as best she could, when the injury actually started versus when a farmhand first noticed it. Median gap: three and a half days. That was the first real number anyone at Grazeline had that belonged to this specific barn.
Then she pulled three published dairy science field trials, the kind that ran cameras on whole herds, not curated ones. Four to nine days of lead time, clustering around six. A real range, from real production conditions, just not Farriday's conditions specifically.
Then she called the vet directly and asked one more question: how much warning would actually change what you do? The vet didn't hesitate. Anything under a day, she said, and there's no time to book a hoof trim before the cow is already limping badly enough that you'd have caught it yourself. One day was the real floor. Below it, the number could be true and still be worthless.
Two years earlier, when Grazeline first built its pilot demo for investors, the team had made a small, sensible choice: show the system's best work. They picked the clearest cases, the ones a vet had already flagged as textbook lameness, because that's what makes a compelling five minute pitch. Nobody in that room was picturing a sales rep forwarding that same number to a farmer as a delivery promise. Why would they. The demo wasn't built to be a contract.
By Thursday, Selke had a different number ready for the board, and a different email for Farriday's owner: a target range of two to five days, three days as the working number, with a plain sentence explaining the gap from the pilot's nine. She sent the vet log alongside it, so nobody had to take her word for it.
Ninety days later, at the renewal review, Sentry's actual median lead time at Farriday came in at three and a half days, comfortably inside the range, nowhere near the number that almost got promised on a Tuesday afternoon by someone who'd never seen the vet log at all.
What I would tell myself, before any of this: the day a great number shows up with no source attached is the day to go find out what it's actually describing.
BOUND, five moves for a number nobody has measured yet
This is a target setting question with zero history to lean on, so BOUND does the work here, not a habit with two settings and not a ranked list of features to build.
Two things worth naming directly, since this is where the real judgment sits. First, the easy shortcut on offer was reporting the nine day demo number, since it was already real, already measured, and already in a slide. That got rejected on purpose: it came from cases a vet had already flagged as obvious, not from a random slice of an actual herd, so it overstates what a brand new deployment can promise. Second, the AI specific risk worth naming is distribution shift between curated demo footage and a messy real barn, different camera angles, dust on the lens, cows standing in a crowd. The guardrail is a golden set: thirty days of fresh, vet confirmed cases collected from Farriday's own barn, used to check and, if needed, adjust the target before the ninety day review, instead of trusting the pilot footage or the outside papers alone. The trade accepted here is real too: catching a case earlier means flagging fainter signs, and fainter signs mean more healthy cows get flagged by mistake, so a lower cut off only ships once it clears that fresh golden set at under one false alarm in twenty, most weeks, not on the one lucky day it was first tested.
And if you want to be sure it really works, try it somewhere else
Same five letters, rooftop HVAC compressors instead of dairy cows, so the method proves itself instead of repeating a story I happened to prepare.
Barrow Mechanical runs Foresight, a vibration and temperature monitor that predicts a rooftop compressor failure on a commercial building days before it actually breaks. Briar Hallstrom leads product for it, and a mid size office building is about to become Foresight's first paying customer.
B, break it down. The target sits between how fast the building's own maintenance crew already notices trouble, what published vibration monitoring case studies show, and the smallest warning that leaves enough time to actually get a part.
O, own the numbers. The building's informal repair log shows technicians usually notice a compressor "acting up" about a day and a half before it fails. Published vibration analysis case studies report five to twelve days of lead time, median eight. Foresight's own pilot, run on equipment already near end of life with obvious vibration signatures, showed ten days early. Procurement says a replacement compressor part takes at least three business days to arrive, no matter what.
U, use a range. Eight days trimmed hard for a first real building, checked against the three day parts floor, which sits far above the building's own one and a half day baseline, lands on a working range of three to six days, four as the committed number.
N, nail the sanity check. The facilities director only ever saw the ten day pilot result. Four survives the question of why it's less than half that, because the pilot ran on equipment already failing loudly, not a live building's full mix of old and new units.
D, direction. How representative the pilot building's aging equipment is of Barrow's actual customer base, mostly newer, better maintained systems that may show weaker warning signs.
Same shape, different stakes. At Grazeline the floor came from a vet's schedule. At Barrow it comes from how fast a warehouse can ship a part. The method doesn't change: measure a fresh baseline, pull a real outside benchmark, find the floor where a number stops mattering, then report a range.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the range: never report a pilot's best case number as a production target, always triangulate from a fresh manual baseline, a real benchmark, and a floor.
Cost: the customer wants the target committed before the vet log even exists. Don't invent a fresh baseline out of nothing, use the industry benchmark alone, trimmed harder, and say plainly that the range will tighten once real barn data exists.
The model got better, for real: say Sentry's detection model genuinely improves next quarter. That still doesn't mean the target jumps to match the old demo, a better model chasing a still unproven barn deployment needs its own fresh triangulation, not a borrowed one.
Where people run it wrong.
They report the best number they already have, because building a fresh one feels like extra work nobody asked for.
They pick a single point estimate because a range feels like it shows weakness, when a range is what an honest first read actually looks like.
They skip the sanity check, so the first time anyone compares the target to the flashy old number is in front of the customer, not before.
How to use it live. Open with the refusal before any numbers: "I wouldn't use the number the demo gave us, because it was built on the easiest cases, not the real distribution the product will actually see." That buys you room to walk through real arithmetic instead of reciting "we'd set an ambitious goal" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the demo's nine days is actually representative and you're leaving value on the table with a lower target?" Response: that's exactly what the thirty day golden set is for. If real barn data clears six or more, the target moves up on evidence, not on hope.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?