ConceptFoundationalQuality, Cost & Token Economics / Cost modeling and unit economics / #2

What cost drivers exist for an AI feature beyond model tokens?

ORDER · cost & budget

The model bill never moved. Five other things were stacked on top of it, and only one of them kept climbing long after everyone had stopped watching it.

The direct answer
Rank the non-token cost drivers by how hard each one is to undo, not by which one looks biggest this month. Storage and retention of the raw video, audio, and transcripts goes first, because it compounds every month under a legal hold that takes a legal review to undo, not a settings change, followed by speech-to-text and diarization compute, human standards review, bandwidth, and retries, in that order. On a newsroom transcription tool processing 3,200 videos a month, tokens held flat at $1,216 while storage climbed from $36 a month to $9,360 within three years, and by the time anyone noticed, no setting could pull it back down.
Rank these, in order
  1. Rank storage and retention of raw video, audio, and transcripts first.Why: it's the one cost that compounds every month under a legal hold, and undoing it later takes a legal review, not a settings change.
  2. Budget speech-to-text and diarization compute as its own line, not folded into "model cost."Why: it runs on every upload before the summarizer ever sees a token, and here it already ran more than twice the token line.
  3. Route only flagged, risk-tier quotes to human standards review, never every summary.Why: reviewing all 3,200 uploads by hand would cost more than running the tool at all; reviewing the one in eight that got flagged kept the line honest.
  4. Fix what triggers the retention hold before touching storage architecture.Why: a cheaper storage tier doesn't help if the hold still fires on every raw upload instead of only on published stories.
  5. Treat bandwidth and retries as flow costs to tune later, not rank first.Why: they reset the moment routing or retry logic gets fixed; they don't keep growing on their own the way storage does.
  6. Run a short pilot on real upload volume before signing any storage contract.Why: a two-week check would have shown the archive on pace to pass three years of footage inside fourteen months, not a slow surprise a year later.

How to answer this, stage by stage

Nobody's grading whether you can list five costs. They're grading whether you can rank them by which one you can still undo tomorrow, not by which one looks worst on this month's invoice.

1
Scope it to one product before ranking anything in the abstract
Say it like this
"Let's ground this in one tool. ReelScribe is Stavemark Media's AI feature that transcribes and summarizes raw video producers upload, so a producer can find a usable quote in minutes instead of scrubbing through footage by hand. Ruth Achieng owns it, and Diego Salaverry is the finance partner who reforecasts what it costs to run."
Why this works
A "what else costs money here" question turns into a checklist fast. One product turns it into a real ranking problem.
2
Say your structure out loud before naming a single driver
Say it like this
"I'm going to name what these costs are all actually competing to protect, rank them by which is hardest to undo rather than which is biggest today, say what has to be decided before what, name a cheap way to check my ranking, then give the order."
Why this works
Tells the interviewer you have a method for ranking, not a list you're inventing out loud.
3
Name the outcome every cost driver is actually competing to protect
Say it like this
"Every one of these costs is competing to protect the same thing: keeping the cost per video small enough that nobody ever has to ask whether the tool is still worth running. Rank without naming that first, and you're ranking by gut."
Why this works
Without a named outcome, "rank these" is just an opinion with a number stuck on it.
4
Rank by what's hardest to undo, not what's biggest this month
Say it like this
"If I ranked by dollars this month, I'd chase human standards review first, since it already costs more than storage does. But storage is the one that compounds under a legal hold and can't be undone with a settings change. That's why it goes first, even before it's the biggest number on the page."
Why this works
This is the whole test of the framework: a ranking that only follows this month's dollar figure gets the order wrong.
5
Say what has to be decided before what
Say it like this
"You can't fix the storage line by picking a cheaper storage tier. You have to fix what triggers the three-year hold first, because right now it fires on every upload instead of only on a published story. Architecture is downstream of that policy call, not the other way round."
Why this works
Naming the dependency stops a team from spending a quarter re-architecting storage before the one decision that would have made most of that work pointless.
6
Name the cheap evidence you'd gather before committing a budget
Say it like this
"Before signing any storage contract, I'd run a two-week pilot on real upload volume and watch how fast the archive actually grows. That alone would have shown the archive on pace to pass three years of footage in fourteen months, not a slow surprise a year later."
Why this works
Cheap evidence beats a confident guess dressed up as a signed contract.
7
Close by stating the order and defending the top pick
Say it like this
"So, in order: storage and retention first, speech-to-text and diarization compute next, human standards review after that, then bandwidth, then retries. Storage goes first because it's the only one still costing money three years after the video that caused it was uploaded."
Why this works
Ending on the stated order, defended in one line, is what makes this sound like a ranked decision instead of a list read aloud.

Let's learn

ReelScribe is an AI tool Stavemark Media built so a producer can upload raw video and get a transcript and a short summary back, quotes and all, in under two minutes.

Before ReelScribe, only footage that actually made it into a published story got logged and kept. An assistant producer spent about fifty minutes transcribing and cataloguing each eighteen-minute tape by hand, and that was only worth doing for the roughly 640 stories a month that ran.

ReelScribe processes raw video the moment it's uploaded, published or not. Turnaround runs under 90 seconds. Producers liked it enough that upload volume went from a 400-video pilot to 3,200 videos a month, newsroom wide, in about four months.

Knowledge spark: what's diarization? The step that works out who's speaking, and when they switch, before any summary gets written. It happens on the raw audio, before the model reads a single word of transcript.

The turn: the extra dollars were never about the model getting worse or costing more to run. Tokens held flat at $1,216 a month the whole time. The turn is that Stavemark's old storage rule said anything logged into the archive gets held three years, a rule written back when logging something in was a producer's deliberate choice. ReelScribe logged everything in automatically, so a rule built for 640 choices a month quietly started applying to 3,200 uploads a month, most of them raw footage nobody would ever watch again.

The model didn't do more work. The archive just stopped asking anyone's permission before it grew.

Here's the build-up behind that. At month 14, when the real cost finally got checked, five things beyond the token line were adding up every month.

The build-up: a month of ReelScribe's real cost, beyond tokens, at month 14
$12k $6k 0 Tokens $1,216 Speech $2,720 Review $3,500 Storage $3,024 Egress $960 Retries $448 Total $11,868
Tokens (reference line)Speech-to-text and diarizationHuman standards reviewStorage and retentionEgress and bandwidthRetries and reprocessing
At month 14, human standards review is actually the biggest single line, ahead of storage. Ranking by this month's dollars would put review first. That's the wrong call, and the next chart shows why.

Human standards review isn't automatic. Only about one in eight uploads gets a quote flagged worth a second look, 400 videos a month, 15 minutes each, at a blended reviewer rate of $35 an hour. That's real money, $3,500 a month, and it's the biggest bar in that chart. But it's a flow cost. Fix the flagging threshold or add a reviewer, and next month's number moves with it.

Storage doesn't work that way. Every video ReelScribe touches gets logged into the same archive, and Stavemark's standards rule holds anything logged in for three years. That's not a cost tied to this month's uploads. It's a cost tied to every upload since the rule started applying to everyone.

Two lines over three years: tokens, flat, against storage, climbing
$10k $5k 0 the hallway remark Mo. 1 Mo. 6 Mo. 14 Mo. 24 Mo. 36 $9,360 $3,024
Storage and retention, climbingTokens, flat at $1,216
Storage: $36, then $720, then $3,024 the month it got noticed, then $5,904, then $9,360 by month 36. Tokens never moved. This is why the ranking can't come from this month's invoice.
Hand sketched comparison diagram titled which cost driver can you still undo. Left panel, a gauge icon, labeled flow cost, captioned tokens, speech to text compute, review, bandwidth, retries, fix the routing and next month resets. Right panel, a balance scale icon, labeled stock under hold, captioned raw video, audio, transcripts, locked three years, only a legal review moves it.
Five of these six costs reset the moment you fix them. One doesn't. That's the whole reason it gets ranked first even on the month it isn't the biggest number.
The choice that mattered Stavemark's standards desk wrote a one-line rule years before ReelScribe existed: anything logged into the archive gets held three years, in case a story gets challenged after it airs. That made sense when logging something in was a producer's rare, deliberate choice. It stopped making sense the moment a tool started logging in everything automatically, chosen or not.

At its worst, a cost line nobody's assigned to watch keeps growing every month whether it does anything useful or not, until a routine budget review turns up a transcription tool that costs more to store than it does to run.

What I'd leave alone: the token line genuinely didn't need this treatment. It's a flow cost tied to how many videos get summarized, it's cheap to watch, and 38 cents a video barely moves no matter how much ReelScribe gets used. Spending a quarter renegotiating the model contract to save fractions of a cent would have chased the wrong line entirely.

The lesson: a cost line can be completely honest about what it charges for and still be the wrong thing to leave unranked. Thirty-eight cents a video was a real, true number. It was never the number that was going to blow the budget.

Now here is the same thing as a story

Read the long version below when you want to feel why the biggest line on the invoice wasn't the one that mattered, not just be told that it wasn't.

Ruth Achieng could read a raw camera tape and know inside the first ninety seconds whether there was a usable line in it. She'd worked the assignment desk at Stavemark Media for six years before she started owning ReelScribe, and she knew exactly how the old system worked, because she used to be the one running it by hand.

For the first four months, ReelScribe was a pilot on two desks, politics and courts. About 400 videos a month, and it was good, genuinely good. A producer could pull a Tuesday afternoon news conference and have three usable quotes flagged before the bureau chief finished her coffee. By month five, every desk wanted in, and engineering flipped the switch newsroom wide.

It faded in three beats, and none of them looked like a mistake at the time. Beat one: years before ReelScribe existed, Standards wrote a one-line rule, anything logged into the archive gets held three years, back when logging something in meant a producer chose to enter it, on purpose, because the story ran. Beat two: when ReelScribe went newsroom wide, engineering wired every single upload straight into that same archive, so the transcript would be searchable. Nobody in that build meeting asked whether "logged in" still meant "chosen" once the choosing was gone. Beat three: raw takes that never made a single published story piled up in the same three-year hold as a story that actually ran, month after month, and the one number anyone was actually watching, the token line, stayed exactly the same.

It never announced itself. Diego Salaverry, running Stavemark's Q3 reforecast, noticed the AI infrastructure line under engineering had quietly become the second-biggest line in the whole department, bigger than the newsroom's travel and stringer budget combined. He mentioned it to Ruth in the kitchen, half a question, not even a complaint: "Does a transcription tool really cost that much to run?"

Ruth pulled the token report first, out of habit. Flat, thirty-eight cents a video, exactly what it had always been. She pulled the storage report next, and found 33,600 raw clips sitting under a three-year hold, most of them footage from a story that never aired, that nobody had opened once since the day it uploaded.

The extra dollars were never really about the model. They were about an archive that had quietly stopped asking anyone's permission before it grew.

It was never really about whether thirty-eight cents a video was the right number. Tokens were fine the whole time. It was about a rule written for 640 choices a month that nobody rewrote the day it started firing on 3,200 uploads instead.

The decision that opened the door went back two years, to a Standards meeting that had nothing to do with ReelScribe. The rule was simple and sensible: whatever gets logged into the archive is held three years, in case a story gets challenged after it airs. It made sense when logging in was rare and deliberate. Nobody in that room was picturing a tool that would log in everything, chosen or not.

Run that hallway conversation again with one change: the hold attaches at publish, not at upload. Unpublished raw tape defaults to a 90-day working window instead, long enough for Standards to pull a clip if a complaint comes in on a story that's still fresh, short enough that it doesn't sit there for three years doing nothing. Same 3,200 uploads a month. The archive under permanent hold drops from 33,600 clips to about 9,000, published stories plus whatever Standards actually flagged by hand. The storage bill drops from just over three thousand a month to about eight hundred, inside two 90-day cycles.

One design let a rule written for a human choosing quietly reinterpret itself the moment a machine started choosing everything instead. The other ties the hold to the actual reason it exists, a story going out the door, not a file landing in a folder.

What I'd tell myself, back in the room where we wired every upload straight into that archive: ask what a rule assumed about who was doing the choosing, before you hand the choosing to something that never says no.

ORDER, the five letters behind the ranking

Not a story wearing a framework's clothes. This is a ranking problem, and ORDER is what stops "biggest this month" from quietly standing in for "hardest to undo."

OOutcome. What are all the candidates competing to move?
Every cost driver beyond tokens is competing to protect the same thing: keeping ReelScribe's cost per video small enough that nobody has to ask if it's still worth running. Rank without that named first, and the order is just gut feeling with a spreadsheet attached.
Say the outcome before naming a single driver, or the ranking is opinion wearing numbers.
RReversibility. Which decision is hardest to undo?
Storage and retention wins this, not because it costs the most at month 14, human standards review actually did, but because it's a stock cost that compounds under a legal hold. Undoing it means a legal review, not a settings change. Tokens, speech-to-text compute, egress, and retries are all flow costs: fix the routing or the retry logic, and next month resets.
This is the hardest step, and the one a gut ranking skips. The biggest line today and the hardest one to undo were not the same line.
DDependency. What unblocks what?
You can't fix the storage line with a cheaper storage tier. The trigger has to change first, from upload to publish, or a cheaper tier just holds the same growing pile of footage more cheaply. Architecture is downstream of the policy call, not the other way round.
Naming the dependency stops a team from spending a quarter re-architecting storage before the one decision that would have made most of that work pointless.
EEvidence. What could you learn cheaply before committing?
A two-week pilot on real upload volume, watching how fast the archive actually grows, would have shown it on pace to pass three years of footage in fourteen months, not a slow surprise a year later.
Cheap evidence beats a confident guess dressed up as a signed contract.
RRank. State the order, defend the top pick.
In order: storage and retention first, speech-to-text and diarization compute second, human standards review third, bandwidth fourth, retries last. Storage goes first because it's the only one still costing money three years after the video that caused it was uploaded.
If the ranking would look the same with a different outcome in step one, it was ranked by gut and the outcome got written afterward.

Three things worth stating directly, since this is where the real judgment sits. The alternative Ruth's team considered first, and dropped, was moving the archive to a cheaper storage tier without touching the retention trigger, since a colder tier runs about a third of the price. It lost because a cheaper tier still holds the same growing pile of footage; the archive would have kept compounding, just more slowly, and Stavemark would have hit the same wall a year later instead of avoiding it. The AI-specific failure worth naming by name is scope creep through automation: a policy written for a human's occasional, deliberate choice quietly starts applying to everything the moment a model removes the friction that used to make the choice rare. The guardrail is tying the retention trigger to the actual legal reason it exists, a story going out the door, not a technical event like a file landing in a folder, and checking archive volume against the published-story count every quarter so a gap like this shows up in a report instead of a hallway remark. That guardrail isn't free. Standards accepted paying for a 90-day window on raw footage that will almost certainly never be watched again, because the alternative, deleting it immediately, would have taken away the one tool Standards has for handling a complaint that surfaces after a piece airs. And the bar for the retention system was never zero stored footage, a newsroom has to keep some. It's a bar tied to the actual reason for keeping it, published or flagged, checked every quarter against real archive growth, not a rule that quietly expanded to cover everything a machine happened to touch.

And if you want to be sure it really works, try it somewhere else

Same five letters, a fishing cooperative's dock cameras instead of a newsroom's raw tape, and this time the fix that looked obvious made the bill worse for nine months before anyone noticed.

Tallynet is an AI tool Quaybridge Fishing Cooperative built to watch dock camera footage and count the catch by species as boats unload, so nobody has to review hours of tape by hand to check a quota. Mara Talvik runs operations on it.

The build-up: Tallynet processes about 900 dock videos a month, one per unloading run. Token cost for the count-and-flag summary runs about 22 cents a video, $66 a month, and it's stayed exactly that since launch. The line that actually grew is video storage: every unloading run gets kept for two years under a fisheries compliance rule, in case a catch gets audited against the boat's logged quota. At month 18, with roughly 14,400 videos sitting in that compliance archive, storage alone cost about $2,700 a month, more than forty times the token line.

The decision Mara would take back Assuming a colder, cheaper storage tier would fix the growing bill, without first checking whether every dock video actually needed the full two-year hold, or only the runs flagged for a quota mismatch.

The cheaper tier shaved the bill by about a third and nothing else. Nine months later it had grown right back past where it started, because the tier was cheaper per gigabyte and the pile of gigabytes never stopped growing. Only about one in twenty dock videos ever gets a flagged mismatch worth an audit. The other nineteen didn't need two years. They needed ninety days, the same shape of fix as Stavemark's.

Same rank, different lever: storage goes first here too, not because it's the biggest line at month one, but because it's the one that compounds under a rule nobody revisits. For Tallynet the lever wasn't an automated upload replacing a human's choice, it was a flat rule, hold everything two years, written before anyone had a way to tell a clean unloading run from a flagged one. The fix is the same shape: tie the hold to the actual audit trigger, not to the camera simply having recorded something.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: rank by what's hardest to undo, not what's biggest this month. Storage under a standing hold beats any flow cost every time.
Cost: there's no budget this quarter for both a cheaper storage tier and fixing the retention trigger. Fix the trigger. A cheap tier on a pile that keeps growing forever loses to a smaller pile on any tier.
The model got better, for real: say Tallynet's species classifier gets more accurate and needs fewer retries. That's real, and it shrinks the smallest line on the page. It does nothing to the archive, which keeps growing at the same rate no matter how good the model gets.

Where people run it wrong.
They rank cost drivers by this month's invoice and chase whichever line is biggest today, instead of asking which one they can still undo tomorrow.
They fix a growing storage bill with a cheaper tier and call it solved, without asking whether the retention rule itself still makes sense.
They let a rule built for a rare, human choice quietly cover everything the moment a machine starts making that choice automatically, and nobody revisits the rule.

How to use it live. Say the real question out loud before naming an order: "is this a cost that resets once I fix it, or one that keeps compounding until someone changes a policy?" That buys a beat to actually rank instead of reciting whichever line looks worst on the last invoice.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
ORDER: rank by what's hardest to undo. Built for prioritization questions, not a single number to estimate.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ruth Achieng, product manager for ReelScribe at Stavemark Media. Worked the assignment desk for six years before owning the tool.
3 · THE OLD HABIT
What did the archive quietly start doing that nobody decided on purpose?
Tap to flip
ANSWER
It auto-logged every raw upload, published or not, into the same three-year hold that used to apply only to stories a producer chose to log by hand.
4 · THE RANKING LOGIC
Why does storage rank first when it isn't the biggest line yet?
Tap to flip
ANSWER
Because it's a stock cost that compounds under a legal hold and can't be undone with a settings change, while every other line is a flow cost that resets the moment it's fixed.
5 · THE OLD DECISION
What decision would Ruth take back?
Tap to flip
ANSWER
Wiring every raw upload straight into the archive that triggers the three-year hold, instead of only logging in what a producer chose to publish.
6 · THE NUMBER
Fill in the blank: storage climbed from $36 a month to $___, while tokens held flat at $___.
Tap to flip
ANSWER
$9,360, and $1,216. Storage grew about 260 times over while tokens never moved.
7 · THE REPLAY
Same hallway remark, new design, what changes?
Tap to flip
ANSWER
The hold attaches at publish, not upload. The archive under permanent hold drops from 33,600 clips to about 9,000, and the storage bill drops from just over three thousand a month to about eight hundred.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the same-shaped trap?
Tap to flip
ANSWER
Tallynet, a dock-camera catch counter at Quaybridge Fishing Cooperative. Same trap: fixing storage with a cheaper tier instead of fixing the retention rule that was letting it grow forever.

Check yourself Score: 0 / 0

Fill in the blank
1. Storage climbed from $36 a month to $___ over three years, while the token line held flat at $___ a month the whole time.
Show hint
Check the two endpoints marked on the line chart in Section 1.
Show answer
$9,360, then $1,216. Storage grew roughly 260 times over across three years. Tokens never moved from the day the tool launched.
Multiple choice
2. Why did ranking cost drivers by this month's dollar amount get the order wrong at month 14?
  • A. It ranked speech-to-text compute too low compared to tokens.
  • B. It would have put human standards review ahead of storage, missing that storage is the one cost that keeps compounding under a hold nobody can undo with a setting.
  • C. Dollar amounts change too often to be a useful ranking signal at all.
  • D. Tokens should always rank first regardless of size.
Show hint
Look at the build-up chart in Section 1. Which bar is tallest at month 14, and which one keeps growing in the chart right after it?
Show answer
B. At month 14, human standards review was actually the bigger dollar figure, $3,500 against storage's $3,024. Ranking by that snapshot alone would chase review first and miss that storage is the one line still growing three years later.
True or false
3. True or false: the token cost line needed the same ranking treatment as storage, because both costs technically grow as ReelScribe's usage grows.
  • True
  • False
Show hint
Ask whether tokens are a flow cost or a stock cost, the same question that separated storage from everything else.
Show answer
False. Tokens are a flow cost tied to this month's volume, and stayed flat at $1,216 the whole time. Storage is a stock cost that keeps compounding under a legal hold long after the video that caused it stops being processed.
Short answer, name the rejected alternative
4. What did Ruth's team try first to fix the storage bill, and why did it fail?
Show hint
Look at the paragraph right after the ORDER framework recap, where it names what got tried and dropped.
Show answer
Model answer: Moving the archive to a cheaper storage tier without changing the retention trigger. It failed because a cheaper tier still holds the same growing pile of footage, so the bill kept compounding, just more slowly, and would have hit the same wall a year later.
Short answer, apply it yourself
5. Pick an AI product you use that stores something related to what it processes. Name one cost that might be quietly compounding under a policy nobody has revisited, and how you'd check.
Show hint
Think of a product with a default "keep everything" setting that made sense back when only a person, choosing carefully, decided what got saved.
Show answer
Model answer: A home security camera app might keep every motion-triggered clip for a year, a default written back when a person reviewed and kept only the clips that mattered. Once the camera auto-saves everything the AI flags as motion, that "keep a year" default could be storing thousands of clips nobody will ever open. I'd check by counting how many stored clips are older than 30 days and have never been reopened, and watching whether that count keeps climbing month over month.
Short answer, work the number
6. By month 36 the archive holds a full three years of uploads, and the earliest cohort starts aging off the hold. If upload volume stays flat at 3,200 a month from then on, what happens to the storage bill from month 37 onward, and why?
Show hint
Think about what "held for three years" means the month after the archive has been running for exactly three years.
Show answer
Model answer: It levels off, instead of continuing to climb. For every new month's uploads added to the archive, an equal-sized month from three years earlier ages off the hold and drops out. The archive's size, and its bill, stops growing and settles around $9,360 a month rather than climbing further, as long as volume stays flat.
Before you close the answer
Why this works
Tests whether you'll rank cost drivers by this month's invoice or by which one you can still undo tomorrow. Most candidates list costs. They don't rank them.
Follow-up traps
"Human standards review cost more than storage at month 14, so shouldn't that rank first?" Response: it's a flow cost that resets the moment the flagging threshold or the reviewer headcount changes. Storage compounds under a hold no settings change can touch, that's why it outranks a bigger number today.

"Couldn't you just delete old footage to fix storage?" Response: not the unpublished tape Standards flagged, and not published stories, both sit under a real legal reason to keep them. The fix is changing what triggers the hold, not deleting what's already inside it.
If pressed
The fix never touched the speech-to-text or diarization pipeline at all, those models ran exactly the same before and after. The entire $9,000-plus swing came from one field: a flag set on upload that decided whether a clip's hold clock started the moment it landed, or the moment a story built from it actually went out.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more