CaseAdvancedQuality, Cost & Token Economics / Cost modeling and unit economics / #16
How would you set an internal cost ceiling per user and enforce it?
GUARD · cost ceilings and who they cut off
Fernglass capped every writer at the same number of continuations a month. It never once asked whether keeping book three of a saga going costs the same as finishing a five page short story.
The direct answer
Set the per-user cost ceiling in real dollars of what a writer actually costs to serve, tracked continuously, never as a flat count of how many times they hit continue. Warn a writer well before they reach it with a usage meter they can actually see, and when they cross it, finish the chapter they're in and quietly move later calls to a cheaper, trimmed-context mode instead of going silent. Keep a hard stop in reserve, priced far enough out that it almost never fires, and always leave a way to ask for more.
Do this, in order
Set the ceiling in real dollars of inference cost, never a flat count of continuations.Why: a count treats a five page short story and a hundred and thirty thousand word saga as the same cost, and they were never close.
Track cost per active writer by percentile every week, and flag any manuscript that crosses eighty thousand words as elevated risk.Why: the blended number across all of Fernglass looked fine for over a year while a small slice of writers were already running Passerine at a loss.
Build the usage meter and staged warnings before you ship any ceiling at all.Why: no writer could see their own number, so nobody who got capped had any way to see it coming.
At the ceiling, downgrade to a cheaper, trimmed context mode, checked against a continuity eval set, instead of a silent hard stop.Why: a silent stop mid chapter costs a writer their place in the story. A slightly leaner chapter does not.
Reject a bigger flat number as the fix.Why: raising six hundred to nine hundred still cuts someone off with no warning eventually. It just moves the night it happens.
Leave short manuscript writers alone completely.Why: ninety one percent of Fernglass writers never come near any version of this ceiling. Watching them protects a cost that was never at risk.
How to answer this, stage by stage
Nobody is grading whether you can say the phrase "cost ceiling." They're grading whether you'd count what's easy to count, or go find what actually drives the bill.
1
Scope it to one product before answering in the abstract
Say it like this
"Let's ground this in one product. Fernglass is a co-writing tool from Passerine, it helps hobbyist novelists keep drafting a long manuscript, chapter by chapter, in their own voice. Reeta Aldous is the platform engineer who owns the cost ceiling."
Why this works
A cost ceiling question turns into a policy lecture fast. One real product with a real writer on the other end of it keeps it a design problem.
2
Say your structure out loud before naming a single number
Say it like this
"I'd use GUARD here, because setting and enforcing a ceiling is a risk question first. Who holds the lever and who doesn't, where the cost actually lands, who can't see it coming or push back, the real design fix, and how you'd catch it drifting."
Why this works
Two seconds naming the plan is the difference between a method and a guess with a dollar sign on it.
3
Name what actually breaks about a flat generation count
Say it like this
"A flat cap of six hundred continuations a month assumes every continuation costs about the same. On Fernglass it doesn't. Keeping a story straight means sending the model the outline, the last few chapters, and a character sheet every single time, so the longer a manuscript gets, the more every single continue costs. A five page short story and book three of a saga were never the same six hundred."
Why this works
This is the reframe. A weak answer jumps straight to a bigger number. A real one says why counting calls was the wrong unit from the start.
4
Give the one decision
Say it like this
"So: price the ceiling in real dollars per writer, not a count of calls. Show them a meter. Warn them at seventy and ninety percent. And when they cross it, don't go silent, drop later calls to a lighter, cheaper mode and keep the chapter moving."
Why this works
This matches the direct answer almost word for word, which is exactly what an interviewer is listening for.
5
Prove it with the ticket that started it
Say it like this
"Here's what actually surfaced it. Support forwarded Reeta a message from a five year subscriber, deep into book three of a saga: 'It just stopped. I don't even know if I can keep going.' She almost closed it. Then she checked the queue and found close to the same sentence eleven more times that week. She pulled the accounts behind all of it: seventy one percent of everyone who hit the cap that month had a manuscript over eighty thousand words, and those writers were only nine percent of the subscriber base."
Why this works
One sentence, one number, and the harm lands on a real person. That's the compressed version of the story below.
6
Say what you'd measure, and what you'd leave alone
Say it like this
"I'd track cost per active writer by percentile every week, and flag any manuscript that crosses eighty thousand words as elevated risk before it ever costs Passerine money nobody planned for. I'd leave the shorter manuscripts, ninety one percent of writers, completely alone. There's no real risk sitting there to manage."
Why this works
Shows judgment, not blanket caution. Watching everyone the same amount is the same as watching no one closely.
7
Close on the decision, not the arithmetic
Say it like this
"So: a flat count of continuations hides where the real cost lives, and it cuts people off with no warning. Price the ceiling in real dollars, show the meter, degrade before you ever go silent, and never ship an enforcement rule you haven't checked against who it actually lands on."
Why this works
Ending on the rule instead of the last number crunched keeps this sounding like judgment, not a spreadsheet read out loud.
Let's learn
Here is what happens when a limit is generous for months, and then, for one specific writer on one specific night, it quietly isn't.
Fernglass is a co-writing tool from Passerine. A hobbyist writer pastes in what they have so far, or just asks Fernglass to keep going, and it drafts the next few paragraphs in the voice of the story so far: same characters, same plot threads, same tone. Forty two thousand people pay fifteen dollars a month for the Continuing plan: six hundred continuations, unlimited manuscript length.
Knowledge spark: why a continuation costs more the longer a story gets
Every time Fernglass keeps a story going, it has to remember what already happened: the outline, the last few chapters, who everyone is. A short story needs a page of that memory. A hundred and thirty thousand word novel needs a lot more, and every extra page sent to the model costs real money, every single time it's asked to continue.
At launch, that didn't matter. Three hundred beta writers were mostly drafting short stories and opening chapters, a page or two of memory was plenty, and six hundred continuations a month cost Passerine about three dollars and sixty cents a writer. Against fifteen dollars, that's healthy.
Cost to run one continuation, short story vs. long manuscript
Cheap to keep goingNine times the cost
Same continue button, same six hundred a month cap. One version costs Passerine a fraction of a cent. The other costs nine times as much, every single time.
Fourteen months in, Fernglass had forty two thousand subscribers. Nine percent of them, about thirty eight hundred writers, had manuscripts over eighty thousand words, deep into a novel or partway through a series. Every one of their continuations now needed the full outline, the last three chapters, and a running character sheet, about nine thousand four hundred tokens of memory instead of eight hundred.
Fernglass didn't get more expensive because the model changed. It got more expensive because the writers who stuck with it the longest wrote the most story for it to remember.
The choice that mattered
Fernglass launched with no usage meter at all. No writer could see how many continuations they'd used, or how close to six hundred they were. That was a fine call at launch: three hundred writers, short manuscripts, nobody was ever going to get close. It stopped being fine once a slice of writers were burning through six hundred expensive continuations a month with no way of knowing it.
At its worst, this doesn't end in one dramatic collapse. It ends in Fernglass going quiet in the middle of somebody's best chapter, with no warning and no explanation, during the one hour a week they had to write at all.
What I would leave alone: the ninety one percent of writers with shorter manuscripts need none of this. Their continuations already cost a fraction of a cent. Building meters and warnings around them spends engineering time defending a budget that was never actually at risk.
The lesson: a cap can be completely fair by count and still be wildly unfair by cost. Six hundred was the same number for everyone. It was never close to the same bill.
Now here is the same thing as a story
Read the short version above if you're using this to answer out loud. Read the story below for how a quiet launch day decision nearly cost forty two thousand writers a wall they never saw coming.
Reeta Aldous had built usage caps before Fernglass, at a photo editing app where every edit cost roughly the same no matter what photo you fed it. A flat count of two hundred edits a month worked fine there, because two hundred really was two hundred, every time. Reeta carried that same instinct into Passerine without questioning it, because it had never once failed her.
For the first year at Fernglass, it didn't fail her here either. Three hundred beta writers signed up, mostly drafting short stories and opening chapters, and six hundred continuations a month was more than plenty. Reeta shipped the cap, shipped no meter to go with it, and moved on to the next thing. There was no reason yet to look twice.
Word spread in a specific place: fan fiction forums and long form writing groups, where a chapter a week for a year was normal, not rare. By month ten, Fernglass had picked up a real population of long haul novelists, the kind who open the same document every Sunday night for a year and keep going. Usage stopped looking like the beta.
Fourteen months in, Fernglass had forty two thousand subscribers, and Reeta's dashboard still only showed one number: total spend, company wide, trending fine. Nobody was watching cost by how long a manuscript actually was, because nobody had ever built that view.
Then a support message got forwarded to Reeta's queue, the kind she usually skimmed and closed. A five year subscriber, deep into book three of a saga she'd been writing since before Fernglass existed, had written back to a canned "need more help?" reply with one line: "It just stopped. I don't even know if I can keep going."
Reeta almost closed it. Then she checked the rest of that week's queue and found close to the same sentence, eleven more times, from eleven different writers.
Seventy one percent of everyone who hit the hard cap that month had a manuscript over eighty thousand words, though those writers made up only nine percent of the subscriber base, close to eight times overrepresented.
Checking timestamps, forty four percent of them hit the wall in the middle of an active session, a continuation requested within ten minutes of the last one. Fernglass hadn't just run out on them between visits. It went quiet mid scene.
The decision that opened the door went back to launch, to shipping a cap with no meter attached, because with three hundred short story writers nobody was ever going to get close enough to need one. Nobody ever sat down and decided that should still be true at forty two thousand subscribers and a nine percent slice writing full novels. It just kept being true because nobody had reason to check.
Reeta's first instinct was the fast fix: raise six hundred to nine hundred, ship it that week, watch the ticket volume drop. She wrote most of the plan before killing it herself. Nine hundred still counts calls, not cost. It still goes silent with no warning. It just moves the night it happens from week three of the month to week four.
What she asked engineering for instead: price the ceiling in real dollars of what a writer actually costs Passerine, tracked daily, with a visible meter in the app. Warn at seventy percent, warn again at ninety. At the ceiling, don't cut anyone off mid chapter, finish what they're writing and quietly route later continuations to a lighter mode, a trimmed memory and a cheaper model for routine continuing, checked against a continuity eval set first so a reader wouldn't notice the difference in a scene that isn't carrying the plot. Keep a hard stop in reserve, priced with enough room that it almost never fires, and always leave a button that says ask for more.
Run that month again with the redesigned ceiling live: the same subscriber, same pace, crosses seventy percent of her real budget around day nineteen and gets a quiet nudge in the app. She crosses ninety around day twenty four. On day twenty seven, instead of going silent, Fernglass moves her onto the lighter mode. She finishes chapter forty the same night she started it. Passerine's cost on her account for the month settles around twenty one dollars instead of running past thirty one uncapped, still not cheap, but never a wall she didn't see coming.
What I'd tell myself, back at launch: we asked whether six hundred continuations was generous. We never asked whether a continuation cost the same thing for everyone who'd use it, or what happens to someone the day it quietly doesn't.
GUARD, the five checks a call count never asked
This isn't a pricing question wearing a safety word. It's a risk question, and GUARD is what stops a number that's easy to count from standing in for a number that actually matters.
GGroups. Who holds the lever, and who doesn't?
Reeta and Passerine's finance team hold the lever. They can change the ceiling, the price, or the model whenever the dashboard asks them to. Every writer sits underneath one flat number with no lever of their own, and the ones actually at risk are the writers deep enough into a long manuscript that a continuation costs Passerine far more than fifteen dollars a month was ever priced for.
Naming both sides before touching a number is what keeps a cost ceiling a design decision instead of a spreadsheet exercise.
Reeta held the dashboard. A writer at 11pm, deep in chapter forty, held nothing, and was about to get the same outcome anyway.
UUnequal. Where does the cost actually land?
Cost follows manuscript length, not writer count. A continuation from a manuscript over eighty thousand words runs close to nine times the cost of a short story's, and that nine percent slice of subscribers accounted for the bulk of the real overage, even though the blended number across everyone looked completely fine.
The unevenness isn't noise. It's concentrated on exactly the writers who'd stuck with Fernglass the longest.
Who actually hit the hard cap that month, against their real share of Fernglass
Long manuscript writers, cap hitsLong manuscript writers, of everyone
A group that's one in eleven writers accounted for close to three in four of every hard stop that month.
AAbility to contest. Who never gets to push back?
No writer could see their own continuation count against six hundred, because Fernglass never built a meter. When someone hit the cap, there was no warning it was coming, no message explaining why the screen had gone quiet, and no way to ask for one more chapter before the next billing cycle. The people the cap would hit hardest had zero visibility into it, and the people with the dashboard weren't the ones who'd feel it.
This is the hardest step, and the one most cost answers skip entirely. An enforcement rule nobody can see coming isn't a safety rail, it's a wall.
Three working steps and a fourth that was never built at all.
RReduce. The specific product decision.
Price the ceiling in real dollars of inference cost per writer, tracked daily, not a flat count of calls. Show a usage meter in the app. At the ceiling, finish the current chapter, then route later continuations to a trimmed context, cheaper model mode, checked against a continuity eval set so the lighter mode doesn't quietly drop a plot detail a reader would notice. Reserve a true hard lockout for far past the ceiling, priced with enough headroom that it essentially never fires.
A real design decision, not a policy memo. The bar isn't zero cost on a heavy writer. It's a calibrated threshold, checked against an eval set so the cheaper fallback doesn't quietly get worse.
DDetect. How you'd know before the next queue of tickets.
Track cost per active writer by percentile every week, P fifty, P ninety, P ninety nine, not one blended number a month. Flag any manuscript crossing eighty thousand words as elevated risk before it ever costs Passerine money nobody expected. Watch "hard lockout fired mid session" as its own count, and hold it near zero.
The blended number is exactly what let this run for over a year. Detection has to live under it, not inside it.
Three things worth stating directly, since this is where the real judgment sits. The alternative Reeta wrote up first and killed herself was raising the flat cap from six hundred continuations to nine hundred. It lost because it never touches the actual mechanism, a call count blind to real cost that goes silent with no warning, it just delays the night it happens. The AI specific failure worth naming by name is continuity context creep: keeping a long story coherent means sending the model more of it every time, so cost isn't flat per call, it grows with how much of the story the writer has actually written, and a flat cap hides that growth completely. The guardrail is percentile based real cost monitoring paired with a graduated, eval gated degrade path, so a cheaper fallback mode never ships without being checked against whether it still remembers the story. And the trade-off is real: the trimmed context mode is measurably worse at catching a small detail from ten chapters back than the full mode is. That's the cost being traded for keeping the story moving instead of going silent, and the bar was never zero risk of a missed detail. It's a recall threshold on the continuity eval set, set high enough that a reader wouldn't notice, checked before the fallback ever ships to a real writer.
And if you want to be sure it really works, try it somewhere else
Same five letters, a livestock health app instead of a writing tool, and this time the thing driving cost isn't story length, it's how many photos it takes to see what's actually wrong.
Hoofnote is a livestock health app from Grazeworth. A farmer photographs a sick or injured animal, an eye, a hoof, a patch of skin, and Hoofnote flags how urgent it looks and what it might be, before a vet visit gets booked. Cato Rusk runs product for it.
The build-up: Grazeworth capped every farm account at fifty diagnoses a month, flat, the same shape of mistake as Fernglass's flat continuation count. A single clean eye photo is one image and costs almost nothing to read. A hoof crack under matted hair, or a lesion that needs three angles to judge properly, can take four to six photos to diagnose with any confidence, and Hoofnote has to read every one of them.
The decision Cato would take back
Pricing Hoofnote as one flat "fifty diagnoses" tier, with no distinction between a single clear photo and a six photo hoof case, because at launch almost every submitted photo set was one image.
G, groups. Grazeworth's ops team holds the lever. Smallholder farms running mixed species, goats, sheep, poultry together, hold none, especially the ones without a handling chute who can't get one clean photo of a spooked animal and end up submitting several blurry ones instead. U, unequal. Photo count, not farm size, drives the real cost. A small mixed species farm during lambing season can burn through its ceiling in cost terms while a large single species dairy operation with a proper chute never comes close. A, ability to contest. A farm has no visibility into its own diagnosis cost, only a raw count of fifty, so an account can be cost heavy and still look fine on the one number it can see, right up until a lockout message during calving season. R, reduce. Route photo sets past a per diagnosis threshold to a lightweight single best photo triage mode, still returning an urgency flag, instead of raising the flat count for every farm regardless of what it submits. D, detect. Track cost per farm by average photos per diagnosis weekly, not by farm size or plan tier, so a photo heavy mixed species account doesn't hide inside an average built from mostly single photo dairy checks.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: never enforce a cost ceiling as a flat count of calls, price it in real dollars, warn early, and degrade before you ever go silent on someone mid task.
Cost: no budget this quarter for both the meter and the eval gated fallback. Build the meter first. You can't safely degrade a cost you can't see broken out by segment yet.
The model got better, for real: say Fernglass's underlying model gets cheaper per token across the board. That helps the blended number, but it doesn't undo the shape of the problem. A hundred and thirty thousand word manuscript still needs far more context than a five page short story, just at a lower price per token.
Where people run it wrong.
They set the ceiling in whatever's easy to count, calls, photos, messages, instead of what actually costs money.
They enforce it with a hard, silent stop, and never build the meter that would have let anyone see it coming.
They fix a cost problem by raising the number for everyone, without checking who the ceiling actually lands on first.
How to use it live. Ask the real question out loud before answering: "is this ceiling measuring what something actually costs, or just what's easy to count?" That buys a beat to think, and it's almost always where the real design flaw is hiding.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question about setting a cost ceiling and enforcing it, and why?
Tap to flip
ANSWER
GUARD, for risk. The real question is who sets the lever, who doesn't, and who an enforcement rule actually lands on, exactly what GUARD is built to find.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Reeta Aldous, the platform engineer who owns Fernglass's cost ceiling at Passerine, and had built flat usage caps once before without issue.
3 · THE OLD DEFAULT
What default, set at launch, quietly stopped making sense as Fernglass grew?
Tap to flip
ANSWER
Shipping a flat six hundred continuation cap with no usage meter attached. Fine at three hundred short story writers. Dangerous once a slice of writers were burning through six hundred continuations that each cost nine times more.
4 · THE GAP
What did the flat generation count hide?
Tap to flip
ANSWER
That a continuation from a manuscript over eighty thousand words cost close to nine times what a short story's did, so the same six hundred number meant a totally different bill depending on who hit it.
5 · THE OLD DECISION
What fix did Reeta write up first and kill herself, and why?
Tap to flip
ANSWER
Raising the flat cap from six hundred continuations to nine hundred. She killed it because it never touches the real mechanism, a count blind to cost that goes silent with no warning. It just delays the night it happens.
6 · THE NUMBER
Fill in the blank: of writers who hit the hard cap that month, ___ percent had manuscripts over eighty thousand words, against ___ percent of the subscriber base overall.
Tap to flip
ANSWER
Seventy one percent, against nine percent. Close to eight times their overall share, meaning the cap landed almost entirely on Fernglass's longest running writers.
7 · THE REPLAY
Same month, new design, what changes?
Tap to flip
ANSWER
The writer is warned at seventy and ninety percent of her real budget, then moved to a lighter mode at the ceiling instead of going silent. She finishes chapter forty the same night, and Passerine's cost on her account settles around twenty one dollars instead of running past thirty one uncapped.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what plays the role of manuscript length there?
Tap to flip
ANSWER
Hoofnote, a livestock health app from Grazeworth. Photos per diagnosis play the role manuscript length plays for Fernglass. A hard to photograph hoof case can cost far more than one clean eye photo.
Check yourself Score: 0 / 0
Fill in the blank
1. Fernglass's flat plan capped every writer at ___ continuations a month, while a continuation from a manuscript over eighty thousand words cost close to ___ times what a short story's continuation cost.
Show hint
Look at the two bars in the first chart in Section 1.
Show answer
600, and about 9 times. Six tenths of a cent for a short story versus about five cents for a long manuscript, the same monthly count of six hundred hid a completely different bill.
True or false
2. True or false: since only nine percent of Fernglass's subscribers had manuscripts over eighty thousand words, the flat cap wasn't worth redesigning.
True
False
Show hint
Look at the second chart in Section 3, and compare that nine percent to the share of cap hits it caused.
Show answer
False. That nine percent of subscribers accounted for seventy one percent of everyone who hit the hard cap, close to eight times their share. A small slice of writers absorbed almost all of the real harm, and a small slice is still worth fixing when it's the whole reason support tickets kept climbing.
Multiple choice
3. Why did pricing the ceiling in real dollars, with a graduated degrade, beat simply raising the flat cap from six hundred continuations to nine hundred?
A. It's technically simpler to build than a bigger flat number.
B. It fixes the actual mechanism, a call count blind to real cost that goes silent with no warning. Raising the number just delays the night it happens.
C. It makes every continuation higher quality than the flat cap did.
D. A flat cap is against the law to enforce on a subscription product.
Show hint
Look at the R step in the framework recap, and the rejected alternative named right after it.
Show answer
B. A bigger flat number is still a count blind to cost, and still a wall with no warning. Pricing the ceiling in real dollars, with warnings and a graceful degrade, is what actually changes what happens at the edge, not just when it happens.
True or false
4. True or false: the ninety one percent of writers with manuscripts under eighty thousand words also needed a usage meter and staged warnings built for them.
True
False
Show hint
Look at the "what I would leave alone" paragraph in Section 1.
Show answer
False. Their continuations already cost a fraction of a cent, and none of them were ever close to the ceiling. Building monitoring and warnings for them would spend engineering time defending a budget that was never actually at risk.
Short answer, apply it yourself
5. Pick a product you use, or one you've heard of, that has some kind of monthly usage limit. Name one thing that probably drives its real cost that the limit doesn't actually measure, and how you'd check.
Show hint
Think about what makes one use of the product far more expensive than another, even though the limit counts them the same.
Show answer
Model answer: A photo editing app that caps "edits per month" probably costs far more per edit on a large, high resolution image than on a small thumbnail. I'd check by pulling cost per account split by the size of image processed, not by a flat count of edits.
Multiple choice
6. If the share of Fernglass subscribers with manuscripts over eighty thousand words grew from nine percent to eighteen percent next year, with the flat six hundred cap still in place, what would you expect?
A. Blended cost would improve, since more writers are engaged enough to keep going that long.
B. More writers would hit the cap earlier in the month, and more of them would be cut off mid session, since the same expensive slice just got bigger.
C. Nothing would change, since the cap number itself didn't move.
D. It's impossible to say anything without knowing the exact subscription price.
Show hint
Think about what happens when the group that's nine times more expensive per call simply gets bigger, with the same flat number of calls allowed.
Show answer
B. The same mechanism behind the original ticket spike: a bigger long manuscript slice means more accounts running expensive continuations against an unchanged flat count, so hard stops, and the mid session cutoffs that come with them, would only get more common.
Before you close the answer
Why this works
Tests whether you'll design a ceiling around what a user actually costs, or a proxy that's easy to count but wrong, and whether your enforcement mechanism accounts for who it silently cuts off. Most candidates jump straight to a bigger number and stop there.
Follow-up traps
"Why not just raise the price and keep the flat cap simple?" Response: pricing doesn't fix a cap that's blind to real cost and still cuts someone off with no warning, and it also taxes the ninety one percent of writers who were never the problem.
"Isn't a dollar based ceiling just a slower way of doing the same cutoff?" Response: no, because the degrade step is gated by a continuity eval set and never fires mid generation, and the hard stop is priced with enough headroom that it essentially never triggers. It changes what happens at the edge, not just when.
If pressed
The seventy and ninety percent warning thresholds aren't the same fixed number for every writer. They're set relative to each writer's own trailing three month average cost, so a writer who's naturally heavy every month doesn't get an alarming warning during a normal month. Only a real move past a writer's own baseline trips the early warning.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.