Explain the relationship between latency and cost in model selection.
Adframe kept charging the same six tenths of a cent for every Sketchpass call, the whole quarter. What Hollowbrick Labs never priced separately was how many times a designer had to call it before one image was actually good enough to send to a client.
- Track cost per accepted draft per tier, not sticker price per call, and route by it.Why: Sketchpass's price per call never moved, but its real cost per accepted image nearly quadrupled once every retry got counted.
- Set two real thresholds on that number and act at each one, not just chart it.Why: a number nobody acts on past eight cents, then twelve, is a chart, not a decision.
- Default new requests to the tier that matches the measured acceptance rate for that request type, not to whichever tier is cheapest per call.Why: the old default made sense when almost every request was a rough concept, and stopped making sense once client facing requests started skipping the separate polish step.
- Reject routing everything to the expensive tier as the fix.Why: it triples the wait on early concepting work, which is most of Adframe's volume and never needed a first pass guarantee.
- Recheck acceptance rate against a fixed eval set every month, not just once at launch.Why: Sketchpass's acceptance rate on client facing work had been sliding for weeks before any number showed it.
- Leave rough, internal only exploration renders on the fast tier alone, with no routing change.Why: nothing downstream needs those to be accepted, so the retry cost there is real but harmless.
How to answer this, stage by stage
Nobody is grading whether you can say "faster models are usually cheaper." They are grading whether you know the two real ways that stops being true.
Let's learn
What does a cheap model actually cost, once you count how many times someone has to call it?
Adframe is the tool Hollowbrick Labs built so a designer can type a campaign brief and get a marketing image draft back, instead of waiting on a photo shoot or an illustrator's calendar.
Before something like Adframe, a rough concept for a campaign image took a freelance illustrator two to three days to turn around, and a design team of about a hundred and thirty people could only look at a handful of concepts a week for any single campaign.
With Adframe, a designer types a brief into Sketchpass, the fast tier, and gets a first draft back in about two seconds, for well under a cent. A team can now look at ten or fifteen variations in the time it used to take to get one.
The turn here isn't that Sketchpass gets things wrong sometimes. That was priced in from day one, a fast sketch tier is supposed to miss occasionally. The turn is what a designer does about a miss: instead of switching to the slower, pricier tier built for near final work, they just hit generate again on Sketchpass, and again, chasing a client ready look out of a tier that was built to sketch fast, not finish clean.
Here's what was actually driving that line. Sketchpass is a fast draft tier: a small number of generation steps, built to hand a designer six or eight quick variations to react to. Finalcut is the slow tier: many more steps, built to nail a client ready look close to the first try. The two tiers cost what they cost for the same reason they take the time they take, more steps means more compute, means more seconds and more cents, on both sides at once.
Once client facing requests started skipping the separate polish step and going straight into Adframe, Sketchpass kept getting asked to do Finalcut's job. It couldn't, not reliably, so designers regenerated it more and more to compensate. Nobody had told the routing rule that the traffic itself had changed.
At its worst, a cheap looking tier reused past its job costs more than just using the expensive tier once would have, and nobody sees it, because the sticker price per call never changes, only how many times the button gets pressed.
What I'd leave alone: rough, internal only exploration renders that never leave the team's private mood board barely enter this conversation. Nothing downstream needs those to be a finished, accepted asset, so however many times they get regenerated, the real cost stays trivial. Building the fix around that traffic would spend engineering time on a group that was never the problem.
The lesson: a cheap price per call is not the same thing as a cheap price per finished image. The gap between those two is exactly where a probabilistic model hides its real cost, and it only shows up if you're counting tries, not calls.
Now here is the same thing as a story
Read the long version below when you want to feel why a flat price tag hid so much, not just be told that it did.
Bryn Castrejon has run model routing for Adframe for a little over two years. She built the original routing rule herself, the week Adframe first opened up past the ten person pilot team.
In Adframe's first months, the rule was simple and it was right: every new request opened on Sketchpass, because nearly every request really was a rough first pass, someone's Tuesday afternoon attempt to see six headline treatments before lunch. Bryn checked the spend dashboard most weeks, one line, average cost per image, and it always read the same. About four cents, flat, since launch.
For a long stretch, that glance was enough. Somewhere around month five, marketing started sending Adframe requests that skipped the old separate polish step entirely, going straight from Adframe's draft into a client deck. Bryn noticed designers hitting generate more per request, but didn't think much of it. More variations felt like a good sign, not a warning. By month seven, she'd stopped opening the per tier breakdown at all, just the one blended line.
It came back in Hollowbrick's ordinary monthly FinOps review, the kind with forty line items and nobody's full attention. Priya, the finance partner who ran it, messaged Bryn afterward: "Sketchpass's line is more than four times what it was in January. Nothing else moved that much. What changed?"
Bryn hadn't budgeted an afternoon for this, but she pulled eight weeks of Sketchpass logs anyway. Cost per accepted draft, the actual finished image a designer kept and sent on, had climbed from three cents to eleven, while the sticker price per call hadn't moved at all. Designers weren't calling Sketchpass more because they liked it more. They were calling it more because it kept missing on client facing work it was never built to nail in one pass, and nobody had ever told the routing rule that the traffic itself had changed.
Hollowbrick's monthly Sketchpass line went from twenty three hundred dollars to ninety six hundred, over that same quarter, and none of it showed up anywhere Bryn was already looking, because the blended average cost per image barely moved, four cents holding steady the whole time, exactly the number everyone trusted.
It was never really about whether four cents was a healthy price. There was no single price per call number that could describe a cost that had quietly split into two very different jobs, one Sketchpass was built for and one it wasn't.
The decision that opened the door went back to Adframe's very first routing meeting, more than two years earlier. The room agreed Sketchpass would open every request by default, because at the time concepting was genuinely all Adframe did, and the final version still went through a separate design review before it ever reached a client. Nobody chose carelessly. It was the right default for the traffic that existed that week.
Run those eight weeks again with one change: cost per accepted draft tracked weekly, per tier, from day one, next to the blended average, not folded into it. By week five, the number crosses the eight cent watch line Bryn would have set, and her team goes looking, three weeks sooner than the FinOps review actually found it. The fix ships before the quarter closes: client facing requests default to Finalcut, and Sketchpass keeps doing what it was always good at. The Sketchpass line never passes four thousand dollars that quarter, and the routing rule finally matches the requests it's actually routing.
One design let a flat price per call number speak for a cost that had already split into two jobs. The other watches cost per accepted output, and it would have rung three weeks sooner, before a single quarter closed with a finance line nobody could explain.
What I'd tell myself, back at that first routing meeting: the default was never wrong, exactly. It was right for the requests that existed then, and nobody gave the week it stopped being right a name.
LEAD, the four letters behind the eleven cents
This isn't a story wearing a metric's clothes. It's a metric question, and LEAD is what separates a price tag on one call from what a request actually costs to finish.
Three things worth stating directly, since this is where the real judgment sits. The alternative Bryn's team considered first, and dropped, was routing every request to Finalcut by default, trading away Sketchpass's speed everywhere to stop the cost problem at the source. It lost because it triples the wait on concepting work, which is most of Adframe's volume and never needed a guarantee that a first pass would be client ready. The AI specific failure worth naming by name is hidden retry inflation: because a model's output is probabilistic, a cheap call doesn't guarantee an acceptable result the first time, so a tier's real cost is a distribution over tries per accept, not the sticker price of one call. The guardrail is a fixed eval set of about two hundred previously approved campaign assets that any routing default has to clear before it ships, re run monthly, so acceptance rate has a real number behind it instead of a hunch. That guardrail isn't free either: Finalcut runs about thirty five times Sketchpass's sticker price per call and takes sixteen times as long, a real cost and latency premium accepted on purpose for client facing work, not wished away. And the bar Adframe holds itself to was never zero regeneration across every request type, no product serving both five second concepting and client ready final art can promise that. It's a threshold specific bar: cost per accepted draft held under eight cents for the typical week on a given request type, checked against a real eval set, not one blended average standing in for a cost that had already split into two different jobs.
And if you want to be sure it really works, try it somewhere else
Same four letters, an AI invoice audit tool instead of an image generator, and this time it's paid for speed hiding the gap, not hidden retries.
Ledgergate is the AI tool Cargoline Produce, a regional produce distributor, uses to read vendor invoices, match line items against contracted pricing, and flag mismatches before anything gets paid. Dumisani Chikonzo runs finance ops on it.
The usual case, still holding: most of Ledgergate's volume runs on a fast, cheap line item match, an OCR read against the contracted rate sheet, approved if it matches clean. A slower, pricier model kicks in only when something doesn't match, a bundled discount, mixed units, a shipment split across two trucks, and it has to reason through the mismatch instead of just comparing two numbers. So far, latency and cost track the same shape they did at Hollowbrick: the slow model costs more per invoice, and the fast one costs less.
The reserved capacity added a flat thirty four hundred dollars a month, covering roughly forty invoices a day that genuinely needed sub two second turnaround during the three day close window. The other ninety six percent of invoices, arriving on ordinary days, ran fine on the standard queue for a fraction of that cost, but paid the reserved capacity premium anyway, because the contract charged it flat, not by the day. Nobody caught it in the per invoice cost line, since that number reports on the model call, not on the infrastructure tier sitting in front of it, until a routine annual vendor contract review set the flat fee against actual daily usage.
Same rank as before, different shape of break: track the real cost against what's actually driving it, not the number that's easy to read off one dashboard. The fix looks different because the split happened differently: reserve the low latency capacity only for the three day window it actually protects, and let the other twenty seven days run on the standard queue.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: latency and cost usually move together because both come from the same compute knob, but a cheap fast tier that needs many retries, or a speed guarantee bought on top of any model, can break that link. Track cost per accepted output, not the price on one call.
Cost: there's no budget this quarter for both the routing fix and a full FinOps dashboard rebuild. The routing fix wins, since it changes the real cost instead of just changing how the team watches it.
The model got better, for real: say Sketchpass's underlying checkpoint doubles in quality overnight. That's not proof the cost problem is solved. If its acceptance rate on client facing requests still sits below Finalcut's, the retries keep happening until someone actually re measures acceptance against the eval set.
Where people run it wrong.
They watch price per call because it's the number every billing dashboard already shows, and never divide by how many calls it actually took to finish the job.
They "fix" a rising cost complaint by telling a team to use the cheap tier less, instead of asking why the cheap tier stopped being good enough for what it's now being asked to do.
They wait for a finance line to go unexplained before acting, when the tier's own acceptance rate would have told the same story weeks earlier.
How to use it live. Say the real distinction out loud before naming a number: "the price on one call and the cost of one finished result are two different numbers, and only one of them tells you whether the cheap option is actually cheap." That buys a beat to think instead of repeating whatever a billing page already shows.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't you just cache good Sketchpass outputs and reuse them?" Response: caching helps repeat requests, not new campaign briefs, which is nearly all of Adframe's volume. It doesn't fix a tier being asked to do a job it was never built for.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Latency budgets and UX tradeoffs
- #1 What is a latency budget and how would you allocate one across a RAG pipeline?
- #2 Why do you set latency targets at the 95th percentile rather than the mean?
- #3 Describe how streaming changes perceived latency without changing actual latency.
- #4 At what point does latency stop mattering and quality take over?
- #5 How would you decide between a fast weak model and a slow strong one for autocomplete?
- #6 Explain the UX options available when a response will take 30 seconds.