What cost drivers exist for an AI feature beyond model tokens?
The model bill never moved. Five other things were stacked on top of it, and only one of them kept climbing long after everyone had stopped watching it.
- Rank storage and retention of raw video, audio, and transcripts first.Why: it's the one cost that compounds every month under a legal hold, and undoing it later takes a legal review, not a settings change.
- Budget speech-to-text and diarization compute as its own line, not folded into "model cost."Why: it runs on every upload before the summarizer ever sees a token, and here it already ran more than twice the token line.
- Route only flagged, risk-tier quotes to human standards review, never every summary.Why: reviewing all 3,200 uploads by hand would cost more than running the tool at all; reviewing the one in eight that got flagged kept the line honest.
- Fix what triggers the retention hold before touching storage architecture.Why: a cheaper storage tier doesn't help if the hold still fires on every raw upload instead of only on published stories.
- Treat bandwidth and retries as flow costs to tune later, not rank first.Why: they reset the moment routing or retry logic gets fixed; they don't keep growing on their own the way storage does.
- Run a short pilot on real upload volume before signing any storage contract.Why: a two-week check would have shown the archive on pace to pass three years of footage inside fourteen months, not a slow surprise a year later.
How to answer this, stage by stage
Nobody's grading whether you can list five costs. They're grading whether you can rank them by which one you can still undo tomorrow, not by which one looks worst on this month's invoice.
Let's learn
ReelScribe is an AI tool Stavemark Media built so a producer can upload raw video and get a transcript and a short summary back, quotes and all, in under two minutes.
Before ReelScribe, only footage that actually made it into a published story got logged and kept. An assistant producer spent about fifty minutes transcribing and cataloguing each eighteen-minute tape by hand, and that was only worth doing for the roughly 640 stories a month that ran.
ReelScribe processes raw video the moment it's uploaded, published or not. Turnaround runs under 90 seconds. Producers liked it enough that upload volume went from a 400-video pilot to 3,200 videos a month, newsroom wide, in about four months.
The turn: the extra dollars were never about the model getting worse or costing more to run. Tokens held flat at $1,216 a month the whole time. The turn is that Stavemark's old storage rule said anything logged into the archive gets held three years, a rule written back when logging something in was a producer's deliberate choice. ReelScribe logged everything in automatically, so a rule built for 640 choices a month quietly started applying to 3,200 uploads a month, most of them raw footage nobody would ever watch again.
Here's the build-up behind that. At month 14, when the real cost finally got checked, five things beyond the token line were adding up every month.
Human standards review isn't automatic. Only about one in eight uploads gets a quote flagged worth a second look, 400 videos a month, 15 minutes each, at a blended reviewer rate of $35 an hour. That's real money, $3,500 a month, and it's the biggest bar in that chart. But it's a flow cost. Fix the flagging threshold or add a reviewer, and next month's number moves with it.
Storage doesn't work that way. Every video ReelScribe touches gets logged into the same archive, and Stavemark's standards rule holds anything logged in for three years. That's not a cost tied to this month's uploads. It's a cost tied to every upload since the rule started applying to everyone.
At its worst, a cost line nobody's assigned to watch keeps growing every month whether it does anything useful or not, until a routine budget review turns up a transcription tool that costs more to store than it does to run.
What I'd leave alone: the token line genuinely didn't need this treatment. It's a flow cost tied to how many videos get summarized, it's cheap to watch, and 38 cents a video barely moves no matter how much ReelScribe gets used. Spending a quarter renegotiating the model contract to save fractions of a cent would have chased the wrong line entirely.
The lesson: a cost line can be completely honest about what it charges for and still be the wrong thing to leave unranked. Thirty-eight cents a video was a real, true number. It was never the number that was going to blow the budget.
Now here is the same thing as a story
Read the long version below when you want to feel why the biggest line on the invoice wasn't the one that mattered, not just be told that it wasn't.
Ruth Achieng could read a raw camera tape and know inside the first ninety seconds whether there was a usable line in it. She'd worked the assignment desk at Stavemark Media for six years before she started owning ReelScribe, and she knew exactly how the old system worked, because she used to be the one running it by hand.
For the first four months, ReelScribe was a pilot on two desks, politics and courts. About 400 videos a month, and it was good, genuinely good. A producer could pull a Tuesday afternoon news conference and have three usable quotes flagged before the bureau chief finished her coffee. By month five, every desk wanted in, and engineering flipped the switch newsroom wide.
It faded in three beats, and none of them looked like a mistake at the time. Beat one: years before ReelScribe existed, Standards wrote a one-line rule, anything logged into the archive gets held three years, back when logging something in meant a producer chose to enter it, on purpose, because the story ran. Beat two: when ReelScribe went newsroom wide, engineering wired every single upload straight into that same archive, so the transcript would be searchable. Nobody in that build meeting asked whether "logged in" still meant "chosen" once the choosing was gone. Beat three: raw takes that never made a single published story piled up in the same three-year hold as a story that actually ran, month after month, and the one number anyone was actually watching, the token line, stayed exactly the same.
It never announced itself. Diego Salaverry, running Stavemark's Q3 reforecast, noticed the AI infrastructure line under engineering had quietly become the second-biggest line in the whole department, bigger than the newsroom's travel and stringer budget combined. He mentioned it to Ruth in the kitchen, half a question, not even a complaint: "Does a transcription tool really cost that much to run?"
Ruth pulled the token report first, out of habit. Flat, thirty-eight cents a video, exactly what it had always been. She pulled the storage report next, and found 33,600 raw clips sitting under a three-year hold, most of them footage from a story that never aired, that nobody had opened once since the day it uploaded.
It was never really about whether thirty-eight cents a video was the right number. Tokens were fine the whole time. It was about a rule written for 640 choices a month that nobody rewrote the day it started firing on 3,200 uploads instead.
The decision that opened the door went back two years, to a Standards meeting that had nothing to do with ReelScribe. The rule was simple and sensible: whatever gets logged into the archive is held three years, in case a story gets challenged after it airs. It made sense when logging in was rare and deliberate. Nobody in that room was picturing a tool that would log in everything, chosen or not.
Run that hallway conversation again with one change: the hold attaches at publish, not at upload. Unpublished raw tape defaults to a 90-day working window instead, long enough for Standards to pull a clip if a complaint comes in on a story that's still fresh, short enough that it doesn't sit there for three years doing nothing. Same 3,200 uploads a month. The archive under permanent hold drops from 33,600 clips to about 9,000, published stories plus whatever Standards actually flagged by hand. The storage bill drops from just over three thousand a month to about eight hundred, inside two 90-day cycles.
One design let a rule written for a human choosing quietly reinterpret itself the moment a machine started choosing everything instead. The other ties the hold to the actual reason it exists, a story going out the door, not a file landing in a folder.
What I'd tell myself, back in the room where we wired every upload straight into that archive: ask what a rule assumed about who was doing the choosing, before you hand the choosing to something that never says no.
ORDER, the five letters behind the ranking
Not a story wearing a framework's clothes. This is a ranking problem, and ORDER is what stops "biggest this month" from quietly standing in for "hardest to undo."
Three things worth stating directly, since this is where the real judgment sits. The alternative Ruth's team considered first, and dropped, was moving the archive to a cheaper storage tier without touching the retention trigger, since a colder tier runs about a third of the price. It lost because a cheaper tier still holds the same growing pile of footage; the archive would have kept compounding, just more slowly, and Stavemark would have hit the same wall a year later instead of avoiding it. The AI-specific failure worth naming by name is scope creep through automation: a policy written for a human's occasional, deliberate choice quietly starts applying to everything the moment a model removes the friction that used to make the choice rare. The guardrail is tying the retention trigger to the actual legal reason it exists, a story going out the door, not a technical event like a file landing in a folder, and checking archive volume against the published-story count every quarter so a gap like this shows up in a report instead of a hallway remark. That guardrail isn't free. Standards accepted paying for a 90-day window on raw footage that will almost certainly never be watched again, because the alternative, deleting it immediately, would have taken away the one tool Standards has for handling a complaint that surfaces after a piece airs. And the bar for the retention system was never zero stored footage, a newsroom has to keep some. It's a bar tied to the actual reason for keeping it, published or flagged, checked every quarter against real archive growth, not a rule that quietly expanded to cover everything a machine happened to touch.
And if you want to be sure it really works, try it somewhere else
Same five letters, a fishing cooperative's dock cameras instead of a newsroom's raw tape, and this time the fix that looked obvious made the bill worse for nine months before anyone noticed.
Tallynet is an AI tool Quaybridge Fishing Cooperative built to watch dock camera footage and count the catch by species as boats unload, so nobody has to review hours of tape by hand to check a quota. Mara Talvik runs operations on it.
The build-up: Tallynet processes about 900 dock videos a month, one per unloading run. Token cost for the count-and-flag summary runs about 22 cents a video, $66 a month, and it's stayed exactly that since launch. The line that actually grew is video storage: every unloading run gets kept for two years under a fisheries compliance rule, in case a catch gets audited against the boat's logged quota. At month 18, with roughly 14,400 videos sitting in that compliance archive, storage alone cost about $2,700 a month, more than forty times the token line.
The cheaper tier shaved the bill by about a third and nothing else. Nine months later it had grown right back past where it started, because the tier was cheaper per gigabyte and the pile of gigabytes never stopped growing. Only about one in twenty dock videos ever gets a flagged mismatch worth an audit. The other nineteen didn't need two years. They needed ninety days, the same shape of fix as Stavemark's.
Same rank, different lever: storage goes first here too, not because it's the biggest line at month one, but because it's the one that compounds under a rule nobody revisits. For Tallynet the lever wasn't an automated upload replacing a human's choice, it was a flat rule, hold everything two years, written before anyone had a way to tell a clean unloading run from a flagged one. The fix is the same shape: tie the hold to the actual audit trigger, not to the camera simply having recorded something.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: rank by what's hardest to undo, not what's biggest this month. Storage under a standing hold beats any flow cost every time.
Cost: there's no budget this quarter for both a cheaper storage tier and fixing the retention trigger. Fix the trigger. A cheap tier on a pile that keeps growing forever loses to a smaller pile on any tier.
The model got better, for real: say Tallynet's species classifier gets more accurate and needs fewer retries. That's real, and it shrinks the smallest line on the page. It does nothing to the archive, which keeps growing at the same rate no matter how good the model gets.
Where people run it wrong.
They rank cost drivers by this month's invoice and chase whichever line is biggest today, instead of asking which one they can still undo tomorrow.
They fix a growing storage bill with a cheaper tier and call it solved, without asking whether the retention rule itself still makes sense.
They let a rule built for a rare, human choice quietly cover everything the moment a machine starts making that choice automatically, and nobody revisits the rule.
How to use it live. Say the real question out loud before naming an order: "is this a cost that resets once I fix it, or one that keeps compounding until someone changes a policy?" That buys a beat to actually rank instead of reciting whichever line looks worst on the last invoice.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't you just delete old footage to fix storage?" Response: not the unpublished tape Standards flagged, and not published stories, both sit under a real legal reason to keep them. The fix is changing what triggers the hold, not deleting what's already inside it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Cost modeling and unit economics
- #1 Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- #3 Explain how a RAG pipeline's cost structure differs from a single model call.
- #4 How does prompt caching change your unit economics, and when does it not help?
- #5 Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- #6 What is the cost impact of moving from a single call to a five-step agent?
- #7 Describe how you would find the most expensive one percent of your traffic.