What's your criteria in selecting a model?
Transcript
Read the full transcript (1,159 words)
[INTERVIEWER] What's your criteria in selecting a model? "How do you pick a model?" And the answer that ends the conversation is naming your favourite. "I'd use GPT-4o." Why? Because it's good. That's not selection, that's a preference. Model selection is a decision matrix run against your actual requirements. You start from the task, you set minimum bars, and then you optimise cost inside whatever clears those bars.
And more often than people expect, the right answer isn't one model at all. It's routing. Let me give you the framework. This is testing whether you'd make a defensible technical call under real constraints, or just chase the leaderboard. A PM who picks the biggest model for everything ships something that's brilliant and unaffordable. A PM who picks on price alone ships something cheap and broken.
The skill is holding the whole tradeoff at once, and that's what the matrix forces you to do. So first, score your candidates on the axes that actually matter for your task. Quality, but measured on YOUR eval set, not a public leaderboard, because leaderboards get contaminated, the test data leaks into training and the score inflates. Cost per token, both input and output, because output often costs more.
Latency, and specifically time-to-first-token and tokens per second, because those two drive whether the interactive experience feels alive or sluggish. Context window, does it actually fit your documents. Modalities and tool calling, do you need vision, audio, or reliable structured function calls. Privacy and deployment, API versus open-weights you host yourself, because regulated data might legally forbid sending it to a third party at all, and that single fact can force open weights on-prem.
Controllability, can you fine-tune it, does it follow instructions, does it emit reliable structured output. And the operational stuff people forget: rate limits, reliability, vendor lock-in, and the licence terms for commercial use. Now here's the method that keeps this from being a shrug. Set hard minimums first, then optimise. So you write down bars like: must fit thirty-two thousand tokens of context, p95 time-to-first-token under one second, at least ninety-five percent on our eval set, and no training on our data.
Then you filter. Anything that fails a bar is out, no debate. And among the survivors, you take the cheapest and fastest. That two-step, bars then optimise, is what turns a vague "it depends" into a decision you can defend line by line. And then the move that scores real points: consider routing instead of one model. Most traffic isn't hard.
So you send the easy, high-volume queries to a small cheap model, and you escalate only the genuinely hard ones to a frontier model. A cascade like that routinely beats "one model for everything," because you're not paying frontier prices to answer "what are your opening hours." The insight is that your query mix has a shape, and you match the model to the query, not to the average.
Say the tradeoff out loud. Quality, cost, and latency form a triangle, and you rarely max all three at once. The frontier model is the best quality, but it's many times the cost and it's slower. The tiny model is fast and cheap and worse. There's no free corner. So selection is always about which corner your specific task can afford to give up, and routing is how you refuse to give up any of them on average by splitting the traffic.
Let me make it concrete with a support assistant. Your bars: a hundred-and-twenty-eight-thousand-token context, p95 under two seconds, faithfulness at least ninety-six percent on the golden set, and a hard no-training-on-our-data guarantee. Now, rather than pick one model to clear all of that, you route. A small cheap model classifies the incoming intent and answers the simple FAQs directly. Complex, multi-document reasoning escalates to a frontier model.
On rough public pricing, a frontier model runs around two dollars fifty per million input tokens, versus about fifteen cents for a small one. That's roughly a fifteen to twenty times gap. So if you route eighty percent of the easy traffic to the small model, your blended cost per ticket lands around three cents, versus about eleven cents if every ticket hit the frontier model.
And you lost no quality on the hard twenty percent, because those still went to the big model. Same quality where it matters, roughly a quarter of the cost. That's the argument that wins the room. A good interviewer will push on the durability of the choice, because models move fast. The follow-up is "you picked one today, but a better, cheaper model ships next quarter.
Then what?" And the strong answer is that you don't hard-wire a single model into your product. You put an abstraction layer between your app and the model, so swapping providers is a config change, not a rewrite. You keep your eval set as the gate, so any new candidate has to clear the same bars before it's allowed in.
And you re-run that eval on a cadence, because pricing and quality both shift under you. So model selection isn't a one-time decision, it's a standing process with a test harness attached. Saying that tells them you've thought past the demo and into the eighteen months after launch, which is where the real cost lives. So what makes them lean in?
First, you measured quality on your own eval set and you distrusted the leaderboard, which is the mark of someone who's been burned by contamination. Second, you reasoned about routing and cascades instead of one model for all traffic, which shows you think about cost at scale, not just correctness on a demo. And third, you treated privacy and deployment as a hard constraint up front, not an afterthought, because that's the one that can quietly rule out your whole shortlist.
The traps are the ways people give this away. Trap one, naming a favourite model with no requirements behind it, which tells them you'd pick the same thing regardless of the problem. Trap two, ignoring cost and latency, so the choice is technically great and completely unshippable at scale. And trap three, forgetting data residency, because for regulated data, healthcare, finance, government, that single constraint can rule out every API model on your list, and if you didn't mention it, they'll assume you didn't know.
So, recap. Score the candidates on the axes that matter to your task, quality on your own set, cost, latency, context, modalities, privacy, controllability, and the operational terms. Set hard bars, filter to what clears them, then optimise for cost among the survivors. And where the traffic has a shape, route the easy stuff to a small model and reserve the frontier model for the hard tail.
The one line to carry in: set hard bars from the task, filter candidates against them, then optimise cost, and route easy queries to a small model so the frontier model earns its price.