ConceptFoundationalQuality, Cost & Token Economics / Eval design for product teams / #1

What makes an eval product-relevant rather than research-relevant?

Ninety four percent sounds like an answer. It is only an answer once you know ninety four percent of what.

The direct answer
An eval is product relevant when its example set is pulled from your own live traffic, in roughly the same mix of case types your users actually bring you, and that mix gets checked against live traffic on a schedule, not once at launch. Track a distribution match score between the eval set and real production queries every week. The week that score drops, stop trusting the offline number, no matter how flattering it still looks.
Do this, in order
  1. Build the golden set from real production queries and real misses, not a public benchmark or a curated demo set.Why: a benchmark that only resembles your traffic on the surface proves the model can search patents in general, not that it can search yours.
  2. Track a weekly distribution match score between the eval set and live query traffic.Why: this is the number that moves first. It fell for eight months while the golden score sat still.
  3. Wire that score to a real action at each threshold, not just a chart someone glances at.Why: a threshold with no owner is paint on a dashboard. Nobody stopped a release when this one slid, because nothing was wired to make them.
  4. Stratify the golden set by technology class and claim complexity, weighted to real query volume, not to what is easiest to grade.Why: a set that looks spread out can still be starving the one class where a miss costs the most.
  5. Refresh the set every quarter with new production examples, and retire ones that no longer occur.Why: this is what keeps the eval anchored to production as the docket changes, not to the year it was built.
  6. Leave the free quick scan tier's rough check alone.Why: nobody files a patent off a free directional result. That is not where the audit budget belongs.

How to answer this, stage by stage

Nobody is grading whether you can name a metric. They are grading whether you can tell a metric that describes your users from one that used to.

1
Scope it to one product before saying anything abstract
Say it like this
"Let's ground this in one product. Priorline is an AI patent and prior art search tool. An attorney describes a new invention, and it returns the existing patents and papers most likely to block or narrow the claim. Dorottya Szabo runs quality there."
Why this works
A metric question answered in the abstract turns into a list of buzzwords. One product makes it a real decision with real stakes.
2
Say your structure out loud
Say it like this
"I'll run this through LEAD. First the outcome the eval is actually protecting, then the signal that would move before that outcome does, then how that signal gets gamed, then what I'd actually do at each level of it."
Why this works
Tells the interviewer you have a plan instead of a definition, and stops you drifting into a vague answer about "quality."
3
Reframe the question before answering it
Say it like this
"This isn't really asking me to name a metric. It's asking whether I can tell an eval that looks like my product from an eval that actually predicts what happens when my product meets a real filing."
Why this works
Stops you giving the generic "we'd track accuracy and user satisfaction" answer most candidates default to.
4
Name the outcome you are actually protecting
Say it like this
"Not a search score. Whether an attorney can rely on Priorline's result list before they file, without a real blocking reference turning up two years later that we should have caught."
Why this works
The L step. Everything after this only makes sense once the interviewer knows what you're actually trying to move.
5
Give the leading signal, the actual answer to the question
Say it like this
"The eval is product relevant when its own mix of technology classes and claim types matches what real attorneys are searching this month. I track that as a distribution match score. Our golden score can sit at 94 percent for over a year while that match score quietly falls, and the match score is the one that would have warned us first."
Why this works
This is the E step, and it's the whole answer in one breath. A number that stays healthy while the ground shifts under it is not measuring the ground.
6
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happened when nobody was watching that number. A partner reviewing a biotech filing pulled one extra manual search out of habit and found a conflicting paper Priorline never surfaced. When we went back, our match score for that technology class had been sliding for eight months. Nobody had it wired to stop anything, so it just kept sliding."
Why this works
Names the AI specific failure mode the whole answer turns on: a silent miss, a real reference ranked low enough that nobody sees it, with no signal saying so.
7
Say what you'd do at each threshold, and what you'd leave alone
Say it like this
"Above 80 percent match, I trust the score and ship on it. Between 60 and 80, I flag it and check that class by hand before the next release gates on it. Below 60, I stop and rebuild the set before it's allowed to gate anything. And I leave the free quick scan tier alone. Nobody files a patent off a free directional result."
Why this works
The D step. A metric with no attached action is a chart nobody acts on, which is the failure this whole answer exists to prevent.
8
Close on the one line, not the story
Say it like this
"So: an eval earns the word product relevant when its own mix matches your live traffic and gets checked against it on a schedule, not because it once looked realistic in a launch deck."
Why this works
Closing on the rule, not the anecdote, is what makes this sound like a system you'd actually run, not a story you told once.

Let's learn

Say we build a search tool. An attorney types in a new invention, and it returns the existing patents and papers most likely to already cover it, so they know what to write around before they file.

Before a tool like this, a paralegal at a mid size firm spent two to three full days combing databases by hand for a single filing, reading something like 200 documents to find the handful that actually mattered.

Knowledge spark: what is prior art Prior art is anything that already existed before a new invention was filed, a patent, a paper, even an old product listing, that shows the idea isn't as new as the filer thinks. Finding it early saves a client from filing a claim that gets torn apart later.

Priorline, the tool in this story, does the first pass in under four minutes and returns the fifteen or so references most likely to matter, ranked by how closely they cut into the claim. Checked every week against 300 attorney reviewed queries held back just for grading it, Priorline's golden eval score has sat near 94 percent for well over a year.

Ninety four percent sounds like the whole story. It is not. The real trouble was something that number was never built to see.

A score that stays flat is not proof that nothing changed underneath it.
Golden eval score vs. distribution match score, month 1 to month 14
95% 55% month 14: the missed reference M1 M7 M14
Golden eval score, blendedDistribution match score, biotech class
The golden score holds between 93 and 95 percent the entire time. The match score for the biotech class falls from 92 percent at launch to 61 percent by month 14, the week a real reference got missed.

The golden score was reading a population that had stopped being the real one. Software and mechanical patents still made up most of the golden set, because that was Priorline's traffic when the set was built. Biotech and chemistry filings had grown to more than a third of real queries by month fourteen. The golden set barely had any.

Golden score vs. hand checked recall, by technology class
95% 93% 92% 71% Software patents Biotech and chemistry
Golden eval scoreReal recall, hand checked
On software patents, the golden score and the real hand checked recall nearly match, 95 versus 93. On biotech and chemistry, the golden score still reads 92, but real recall checked by hand comes in at 71. The eval set was never large enough there to know the difference.
The choice that mattered When the distribution match score was first added to the dashboard, it was built as a number to glance at, not a number wired to stop anything. That was fine while someone checked it every week by hand. It stopped being fine the day checking it became optional, and nothing else was watching in its place.

At its worst, a search tool attorneys quietly stop trusting is worse than no tool at all. The old way, slow as it was, forced a person to actually read every document. Priorline, once a class of query has been under matched for months, returns a clean looking list with a real gap sitting inside it, and nothing about the screen says so.

What I would leave alone Priorline's free quick scan, the version that gives a rough directional read before someone commits to a full search, genuinely does not need this. It is labeled as non final on every screen, and no attorney files a claim off it alone. Spending audit hours matching its eval to production would take time from the tier where a miss actually costs a client something.

The lesson: a score can be completely honest and still be measuring a docket that does not exist anymore. Ninety four percent told the team the model was fine. It never told them whether the eval was still asking about the same kinds of filings the firm was actually sending in.

Now here is the same thing as a story

The short version is above. Read this one when you want to feel why a passing score can still be lying to you.

Dorottya Szabo spent five years running a two person prior art search desk at a small Boston patent firm before she joined Priorline to build its quality program. She could smell a blocking reference before she finished reading the abstract.

When she joined, Priorline's golden set was 400 examples, built at launch from a public patent search benchmark plus the best hits from early sales demos, the ones that made the product look sharp in a pitch. The blended score sat at 94 percent. For the first year, Monday reviews were quick. She would pull ten hand reviewed queries, compare them against the golden set's pattern, and note that nothing had moved.

After a few months of nothing moving, she cut the manual cross check to two queries a month. By the start of year two, she had stopped opening it at all. The dashboard number was steady. Steady felt like proof.

It came back on a Thursday afternoon, from a client's outside counsel, not from the dashboard. A senior partner reviewing a filing on a modified antibody said the result list looked unusually clean for a technology this contested, and ran one extra manual search out of old habit. She found a conflicting publication out of a Japanese lab sitting on the first page of a plain database search. Priorline's ranked list had never surfaced it.

The score never moved. The population underneath it did.

Half the team wanted to retrain the model that week. Dorottya asked for the week instead, to pull every biotech and chemistry query from the past two quarters, roughly 340 filings, and check them by hand against the real published record instead of against the golden set.

The week was not spent proving the model wrong everywhere. It was spent finding that the model's judgment on software and mechanical filings, still most of the traffic, was fine. The miss was sitting exactly where the golden set had almost no coverage to see it.

It was never really about whether the blended score moved. It had not. What moved was whether a 400 example set, built once and left alone, could still answer for a docket that had grown a whole new technology area since the day it was built. It could not. It had stopped being able to months earlier.

The decision that opened the door went back to a planning meeting two years before, when someone asked whether the golden set should pull in more chemistry and biotech examples, since the firm's client mix was clearly headed that way. The room decided 400 well built software and mechanical examples were enough to launch with confidence, and chemistry could get added once there was volume there to justify it. Volume arrived. Nobody had put a date on when to go back and check.

Run the same Thursday again with one change. The distribution match score for the biotech class crossed below 80 percent around month seven, and this time that crossing is wired to something, not just a line on a chart. A trained searcher gets a flag, checks twenty real queries by hand that week, and finds the same pattern the partner found seven months early, before a client ever needed to catch it themselves.

One design trusted a person to keep checking a chart forever. The other design assumes she will not, and builds the check into the system instead of into her memory.

What I would tell myself, back in that planning meeting: "we'll add it once there's volume" is a promise about the future written into a document nobody is required to reopen. Put a date on it, or it never gets reopened at all.

Four checks, read back in the order that would have saved the quarter

This is LEAD, run start to finish on Priorline's own numbers.

LLink. What outcome is everyone actually racing to protect?
Not the golden eval score on its own. Whether an attorney keeps relying on Priorline's result list before they file, without a real blocking reference turning up later that the tool should have caught.
Name this first, or every later step is just an audit of a number nobody agreed matters.
EEarly signal. What moves before the outcome does?
The distribution match score, how closely the golden set's mix of technology classes and claim types tracks the real mix of queries coming in. This fell from 92 to 61 percent over fourteen months while the blended golden score barely moved.
This is the hard step, and the actual answer to the question. A score that stays healthy while the traffic underneath it changes is not measuring the traffic.
Hand sketched comparison titled how a flattering eval set gets you. Left panel a document icon labeled on the score sheet, caption golden set passes, looks like real traffic. Right panel a person icon labeled in the real filing, caption a chemistry claim, reference missed.
A set can look like your traffic and still not be sampled from it.
AAbuse. How does this get gamed?
A team under deadline can build a set that looks production like on the surface, real patent numbers, real claim language, a spread across several codes, while quietly hand picking the flattering half, the easy claims with clean citations, and avoiding the messy chemistry cases where two searchers might disagree. It passes the "does this look like our traffic" check while still not predicting a real filing.
This is the exact trap the question is testing. Looking representative and being representative are not the same claim.
DDecision. What changes at each threshold?
Above 80 percent match, trust the score, ship on it. Between 60 and 80, flag it, spot check the drifting class by hand before the next release gates on it. Below 60, stop. The score is no longer measuring this product, rebuild the set with fresh production examples before it gates anything again.
A metric with no wired action at each level is a chart, not a decision.

Three things worth stating directly, since this is where the real judgment sits. The alternative the team considered right after the near miss was buying a much bigger third party prior art benchmark, thousands of examples across every technology code, instead of building Priorline's own. It lost, because a bigger benchmark built for nobody's real traffic is still nobody's real traffic. It would have raised the headline number without changing what happens on Priorline's own biotech queries, since the new set's mix still would not have matched the docket any better than the old one did. The AI specific failure mode worth naming by name is a silent miss, the model ranking a real blocking reference low enough that nobody sees it, with no signal telling anyone it happened, which is worse than a wrong answer that at least announces itself. The guardrail is two part, the weekly match score, and a mandatory manual search on any claim in a class where match has slipped past the flag threshold, whatever the model's own ranking says. That guardrail is not free. Routing a chemistry or biotech claim to a mandatory manual search alongside Priorline's own results costs a trained searcher roughly ninety minutes a filing, on a class of filings that used to close in under ten. Priorline pays that cost on purpose, only on the class where the eval is currently thin, not everywhere. And the bar was never zero misses. A system searching millions of documents in seconds cannot promise that. It is an audited match score with three named zones, trusted, flagged, and stopped, checked every week, not a claim that the model is always right.

And if you want to be sure it really works, try it somewhere else

Same four letters, a city permits office instead of a patent firm, nothing about patents anywhere in sight.

CivicScope reviews building permit applications for a mid size city and flags which ones are ready to approve, which need a human plan reviewer's eyes, and which are missing something before they can even join the queue. Rikke Solaru runs quality for it.

The outcome being protected: not CivicScope's accuracy against a fixed test set, but whether a plan reviewer can trust a "ready to approve" flag without a code violation turning up after the permit has already been issued.

The early signal: CivicScope's golden set was built from two years of historic permits, mostly single family remodels. After a zoning change, mid size apartment retrofits grew to over a third of the real queue. The match score between the golden set and the live queue fell from 88 to 54 percent over five months, while the golden accuracy held near 91 percent the entire time.

How it got gamed: a plan review lead, under pressure to hit a quarterly review speed target, kept the golden set stocked with the applications that were fastest and cleanest to grade, mostly the single family jobs, since a strong score there was quick to prove and hard to argue with, even as the apartment retrofit backlog grew mostly unexamined.

Hand sketched decision tree titled does the eval set still match real queries. Root box check the match score. Three branches. Match holds above 80 leads to trust the score, ship. Match slips to 60 to 80 leads to flag it, spot check by hand. Match falls below 60 leads to stop, rebuild the golden set.
The same three zones, run on a permit queue instead of a patent docket.

Same decision, same three zones: CivicScope was sitting at 54, well past stop. Rikke's fix rebuilt the golden set with a stratified sample matching the real permit mix, apartment retrofits weighted to their actual third of the queue instead of their old sliver, and wired the below 60 zone to block a release until a fresh sample passed.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the decision, trust the score above 80 match, flag between 60 and 80, stop and rebuild below 60.
Cost: there is no budget this quarter for both a bigger golden set and a faster release schedule. The match score wins every time. A faster release on a score that has stopped measuring you is just speed toward the wrong answer.
The model got better, for real: say Priorline's headline score climbs to 96 percent. That is not proof the biotech class improved with it. A model can get better on average while the class costing you the most stays exactly as blind as before.

Where people run it wrong.
They read one flat, steady score as proof nothing changed, and never check what the eval set is actually sampling from.
They fix a drifting match score by adding whatever examples are easiest to grade, which flatters the number without fixing the mix.
They treat a rebuild as a one time fix instead of a schedule, so the same drift just takes longer to repeat.

How to use it live. Say the real tension out loud before answering it: "is this asking me to name a metric, or to say how I'd know my metric had quietly stopped meaning anything." That buys a beat to think instead of guessing out loud in front of the interviewer.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
LEAD: find the signal that moves before the outcome does. Built for metric questions where the obvious number lags behind the real story.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dorottya Szabo, quality lead at Priorline, an AI patent and prior art search tool for IP attorneys. Ran a two person search desk at a Boston patent firm for five years before that.
3 · THE OUTCOME
What outcome is the eval actually trying to protect?
Tap to flip
ANSWER
Whether an attorney can rely on Priorline's result list before filing, without a real blocking reference turning up later that should have been caught.
4 · THE EARLY SIGNAL
What's the signal that moves before the outcome does?
Tap to flip
ANSWER
The distribution match score, how closely the golden eval set's mix of technology classes and claim types tracks the real mix of live queries.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating the distribution match score as a number to glance at instead of one wired to stop a release, and never putting a date on revisiting the golden set as biotech volume grew.
6 · THE NUMBER
Fill in the blank: the distribution match score for the biotech class fell from 92 percent at launch to ___ percent by month 14, while the blended golden score held near 94 the whole time.
Tap to flip
ANSWER
61 percent. The golden score never showed the drift because it was blended across classes that had not drifted at all.
7 · THE REPLAY
Same near miss, new design, what changes?
Tap to flip
ANSWER
The match score crossing 80 percent around month seven gets wired to a mandatory hand check, catching the same pattern the partner found, seven months before a client would have needed to catch it themselves.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the same tension?
Tap to flip
ANSWER
CivicScope, a city permit review tool. Same tension: a golden set built for last year's traffic mix keeps passing while the real queue shifts to a technology class the set barely covers.

Check yourself Score: 0 / 0

True or false
1. True or false: Priorline's golden eval score dropping would have been the first sign that its eval had stopped matching real traffic.
  • True
  • False
Show hint
Look at what the two lines on the month 1 to 14 chart actually did.
Show answer
False. The golden score stayed near 94 percent the whole time. The distribution match score was the one that fell first, from 92 to 61 percent, months before the near miss.
Multiple choice
2. Why did the near miss happen even though Priorline's golden eval score sat near 94 percent for over a year?
  • A. The golden set's mix of technology classes had drifted away from the real query mix, so a high blended score no longer described the traffic Priorline actually served.
  • B. The model's confidence score was miscalibrated on every query, not just the biotech ones.
  • C. Priorline's search index had not been updated with recent patent filings.
  • D. The partner who caught the reference used a different search tool by mistake.
Show hint
Check what the by class chart shows about software patents versus biotech and chemistry patents.
Show answer
A. The golden set stayed heavy in software and mechanical examples while real traffic shifted toward biotech and chemistry, so the blended score kept looking healthy while the class that mattered most was barely covered.
Fill in the blank
3. On biotech and chemistry patents, the golden eval score read 92 percent, but real recall checked by hand came in at ___ percent.
Show hint
It's the smaller bar on the by class chart in Section 1.
Show answer
71 percent. A 21 point gap the golden set was never large enough in that class to catch.
Short answer, name the rejected alternative
4. What alternative did the team consider after the near miss, and why did it lose?
Show hint
Look at what "half the team wanted" right after the missed reference, and what the framework recap says lost instead.
Show answer
Model answer: Buying a much bigger third party prior art benchmark instead of building Priorline's own. It lost because a bigger benchmark built for nobody's real traffic is still nobody's real traffic, and its mix would not have matched Priorline's own docket any better than the old set did.
Short answer, apply it yourself
5. Pick an AI product you use yourself. Name one place its eval might be flattering rather than representative, and how you would check.
Show hint
Think of a product that reports one blended score, and ask what population that score is actually being measured against.
Show answer
Model answer: A grocery app's recipe suggestions report one accuracy score across all diets, but the team's test set is mostly ordinary meals. I would check by pulling a sample of suggestions made to accounts with a listed allergy or restriction, grading those by hand, and comparing that recall to the headline number.
Multiple choice
6. Priorline's decision rule sets three zones on the match score. If this month's biotech match score comes in at 68 percent, what should happen?
  • A. Trust the golden score and ship on the normal schedule.
  • B. Flag it and spot check the biotech results by hand before the next release gates on the score.
  • C. Stop everything and rebuild the entire golden set from scratch immediately.
  • D. Ignore it, since the blended golden score itself has not moved.
Show hint
68 sits inside which of the three named zones, trusted, flagged, or stopped?
Show answer
B. 68 percent sits in the 60 to 80 flag zone, which calls for a hand check on that class before the next release gates on the number, not full trust and not a full rebuild.
Before you close the answer
Why this works
Tests whether you can tell a metric that describes your users from one that used to. Most candidates name a metric like accuracy or user satisfaction and stop, without ever asking what population that metric is actually being measured against.
Follow-up traps
"Isn't a bigger eval set always safer than a smaller one?" Response: no. Size does not fix mix. A bigger set built off the same skewed sampling just makes the flattering score more confident, not more true.

"How do you know live traffic itself is the right target, and not also biased?" Response: because it is the actual population using the product, which is the honest definition of product relevant. If the live mix itself has a real gap, a technology class the product should serve but barely sees queries from yet, that is a separate coverage problem the eval should surface, not something the match score should quietly launder away.
If pressed
The real production bar in the flag zone was never measured on the blended average across all technology codes. It was measured on the worst represented class specifically, since a blended 75 percent match can still be hiding one class sitting at 30.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more