How do you write acceptance criteria for cost per interaction?
- Write the ceiling as a range tied to how long the interaction runs, never one flat number for every case.Why: a single number either overpays on the short ones or gets blown through by the long ones.
- Break the cost into its real parts before setting any number: transcription, then the model's input tokens, then its output tokens.Why: a target with no build-up under it is a guess wearing a number's clothes.
- State every assumption behind the estimate out loud: minutes of input, tokens per minute, price per token.Why: an assumption nobody can see is one nobody can challenge, or defend, later.
- Check the ceiling against what the interaction actually earns, not against a number that merely sounds cheap.Why: fifteen cents of revenue can't quietly absorb sixty cents of cost.
- Name which assumption, usage length or model choice, would blow the ceiling first.Why: that's the one an interviewer, or a budget review, will push on hardest.
- Set a real trigger to re-check the criterion, tied to a shift in usage length or a price change, not a calendar date.Why: the criterion goes stale the moment usage shifts, and nobody notices until finance does.
How to answer this, stage by stage
Nobody's grading you on landing exactly fourteen cents. They're grading whether the ceiling comes from real arithmetic, checked against what the thing earns, instead of a number that just felt safe. Seven moves get you there.
Let's learn
Reelframe is a video hosting site. One of its features, AutoChapters, watches a video after someone uploads it and writes out timestamps and titles, so a viewer can jump straight to the part they want instead of scrubbing through the whole thing.
Before anyone set a real ceiling, the number that showed up in the planning doc was twenty cents a video. That's what a typical twenty-minute upload costs to process, plus a bit of room. It felt safe. It matched the actual number for the video everyone was picturing when they wrote it down.
Then someone ran the same math on a ninety-minute recording. Same three-part equation, same prices, just a longer video. Sixty cents. Three times the ceiling, on a video that hadn't done anything wrong.
At its worst, the check gets set to log a warning instead of fail the run, because nobody wants to block a customer's webinar from getting chapters. Then it just doesn't get looked at again. Three months later, finance finds the compute line for AutoChapters ran forty percent over budget, and the reason turns out to be six podcasts and a conference recap that all ran past ninety minutes.
The choice I would take back. We wrote one flat ceiling for every video length, instead of building it from the equation and checking it against what each length actually earns. That made sense at the time, because almost every upload really was around twenty minutes when the number got written down.
What I would leave alone. The model's output cost, the part that writes the chapter list itself, barely moves the total. Doubling the number of chapters adds less than half a cent. No ceiling logic needs to fuss over which exact model writes the titles; the money is almost entirely in transcription and the transcript's own length.
The lesson. A cost ceiling built from one example video is a ceiling built for one length of video. The equation doesn't change when the input runs long. I'd size the check to the equation's actual range next time, not to whichever video happened to be open on my screen when I wrote the number down.
Now here is the same thing as a story
Skip this part if you already believe a flat cost ceiling can't hold both a four-cent clip and a sixty-cent recording. Read on if you don't.
Solveig Marchetti reviews the AutoChapters cost dashboard every Friday afternoon, right before she signs off for the week. She's run the feature for nine months, long enough that she can glance at a spend alert in Slack and know, within a cent, whether it's real or noise.
AutoChapters launched to almost nobody's surprise. The first few months, uploads clustered where they always had: fifteen to twenty-five minute tutorials, product walkthroughs, the odd conference talk. The twenty-cent ceiling held. The dashboard stayed boring, which, for a cost check, is the sign of a spec doing its job.
Reelframe's sales team had a good spring. They landed a run of course creators and event hosts, people who host ninety-minute workshops and three-hour panel recordings, not fifteen-minute demos. The uploads got longer. Nobody rewrote the ceiling to match, because the check only logged a warning, never blocked a run, and a warning nobody reads might as well not exist.
Then, on an ordinary Tuesday, a customer success rep dropped one line into the AutoChapters Slack channel: "cost on the Highline Summit recording came back as $1.20, is that right?"
It was right. Highline Summit had uploaded a three-hour recap of their conference, and the pipeline had run exactly as designed, transcription plus input tokens plus output tokens, on a video six times longer than the one the ceiling was ever priced against.
Solveig pulled the compute logs, expecting one strange video. She found a pattern instead. Every long-form upload since March had been quietly running the same way, past the ceiling, past the warning, past anyone actually reading it.
It wasn't the one three-hour video that hurt Reelframe. It was every long video since spring, doing the same quiet thing, one Friday dashboard at a time.
She remembered the meeting where the twenty-cent number got written into the spec. Someone had pulled up a demo video, twenty-one minutes, ran the math once, and the number felt right, because it was right, for that one video. Nobody in the room asked what a webinar would cost.
She didn't propose a bigger flat number. A single number, however large, would still be wrong for something. Instead she tiered the ceiling by length: eight cents for anything under ten minutes, twenty cents from ten to forty, and seventy cents above that, each one checked against what a video that length actually earns the plan. The check moved from a warning nobody read to a hard block with a clear reason attached.
The next long recording that came in, a ninety-four-minute product summit, got flagged the moment it finished processing, not three months later in a finance review. The following month's AutoChapters compute line came in within three percent of budget, the first time in two quarters it had landed that close.
The thing I'd tell myself, back in the room where we wrote down twenty cents: I priced the video sitting in the demo. I never priced the equation.
What each letter buys you, priced out
This is an estimation question with a build-up hiding inside it, so BOUND fits, not FLIPS. Nobody's trust is flipping here. It's an argument about what a ceiling is made of, and what it costs to get that wrong.
B, break it down. The cost isn't one number, it's a sum: transcription cost, plus the model's input cost, plus its output cost.
O, own the numbers. Twenty minutes of video: twelve cents of transcription, a cent and a half of input tokens, under half a cent of output tokens. About fourteen cents, ceiling set at twenty.
U, use a range. Four cents for a five-minute clip. Sixty cents for a ninety-minute recording. One flat ceiling can't sit inside that range and mean anything.
N, nail the sanity check. The plan this feature ships in earns about fourteen and a half cents of revenue per video. A twenty-cent ceiling already spends more than that, before any other cost is paid.
D, direction. Video length swings the total far more than which model writes the chapter titles. That's the assumption worth watching, not the one that's easiest to tune.
And if you want to be sure it really works, try it somewhere else
LinguaLoop is an AI conversation partner for people learning a new language. A learner opens the app and talks out loud, back and forth, practicing a real conversation.
B, break it down. Cost per interaction is cost per practice session: speech-to-text on what the learner says, plus the model's input and output tokens for its replies, plus text-to-speech to speak those replies back.
O, own the numbers. An average eight-minute session: about five cents of speech-to-text, six tenths of a cent of input tokens, just over a cent of output tokens, and about two cents of text-to-speech. Roughly eight and a half cents, ceiling set at twelve.
U, use a range. A quick three-minute drill costs about three cents. A twenty-minute immersive session costs about twenty-one cents, seven times the short one.
N, nail the sanity check. The subscription this sits in earns about thirty cents of revenue per session. Even the twenty-minute high end still leaves real margin, unlike the video case, where the ceiling ran straight past what a long video earned.
D, direction. Same lever as AutoChapters. Session length, quick drills versus long immersive sessions, swings the total more than which text-to-speech voice tier gets used.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the ceiling and the range, fourteen cents for a twenty-minute video, four to sixty across what actually gets uploaded, checked against fourteen and a half cents of revenue. The build-up backs it up if they ask.
Cost: Reelframe caps the whole AutoChapters compute budget at a fixed monthly figure instead of asking what one video should cost. Same equation, solved backward, divide the budget by expected upload volume and length mix to find what ceiling the equation can actually afford.
The model got better: a cheaper transcription vendor ships and cuts the per-minute rate in half. The ceiling doesn't need a rewrite, just a re-run of the same three-part sum with the new price in the transcription term.
Where people run it wrong.
They set the ceiling from whichever video happens to be open in the demo, instead of the range of lengths actually shipping.
They build the equation once at launch and never rerun it when the model, the price, or the customer mix changes.
They set a ceiling with zero margin, so ordinary variation in a normal-length video trips it constantly and everyone starts ignoring the warning.
How to use it live. Say the equation before any number: "Cost per interaction is transcription, plus the model's input, plus its output, three terms, and I want to see where those land for a short case and a long one before I commit to a ceiling." That sentence buys the time to actually do the arithmetic instead of guessing a number that sounds safe.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Acceptance criteria for non-deterministic output
- #1 Rewrite this criterion to be testable: the model should not hallucinate.
- #2 How do you express an acceptance criterion as a rate rather than an absolute?
- #3 What is the difference between a threshold criterion and a distributional criterion?
- #4 Write acceptance criteria for an AI feature that extracts fields from an invoice.
- #5 How do you set a pass bar when human performance on the same task is 92 percent?
- #6 Describe acceptance criteria that account for the severity of different error types.