CaseAdvancedResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #6
How would you build an internal eval platform that other teams actually use?
ORDER the platform: Gullford Grocers, and ProofBench, its shared eval platform for internal AI features
Interviewer's question: "How would you build an internal eval platform that other teams actually use?" Gullford Grocers runs a regional chain of supermarkets. Ines Delacroix built ProofBench, the company's shared eval platform, starting with the loss-prevention team.
The direct answer
Build it in this order: capture one team's real production failure cases first, earn that one team's genuine trust, then generalize the schema for a second team, and only then build self-serve onboarding and dashboards for everyone. A wrong eval verdict early on, before any team trusts the platform, is nearly impossible to walk back. A missing dashboard is not. Build against the thing you can't undo, not the thing that's easiest to demo.
Do this, in order
Capture one real team's actual failure cases before building anything generic.Why: a second team will never trust a platform that's never proven itself on anyone's real mistakes.
Protect against a trust-destroying wrong verdict above everything else.Why: it's the one mistake in this whole build that doesn't have an undo button.
Generalize the schema only after the first team's trust is real.Why: a second team adopts based on evidence from the first, not a promise about the platform's design.
Keep any pass-rate number off a leadership-visible leaderboard until gaming risk is handled.Why: a visible score creates an incentive to curate the test set, not to fix the product.
Build self-serve onboarding and dashboards last, not first.Why: they're the easiest pieces to add later and the easiest to get wrong before anyone's used the platform for real.
How to answer this, stage by stage
Nobody is grading how many platform features you can list. They're grading whether you can say which piece breaks everything else if you build it out of order.
Stage 1
Scope it to one company, one platform
Say it like this
"I'll answer this for Gullford Grocers and ProofBench, its shared eval platform, starting with the loss-prevention team's shrink-detection model."
Why this works
Keeps "build an eval platform" from turning into an abstract architecture diagram.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what everything competes to move. Reversibility, the hardest gap to undo. Dependency, what unblocks what. Evidence, what's cheap to learn early. Rank, the actual order."
Why this works
Signals a method before listing a single platform feature.
Stage 3
Name the real outcome
Say it like this
"Every feature is competing to answer one question: will a second team actually gate their own launch on this platform's verdict, or quietly ship around it?"
Why this works
Without this, ranking platform features is just a matter of taste.
Stage 4
Find the least reversible gap
Say it like this
"A wrong eval verdict during someone's first real use of the platform is the hardest thing to undo. Once a team catches it being wrong, they stop trusting it, permanently."
Why this works
ORDER's hardest step: build first against whatever can't be recovered later.
Stage 5
Say what depends on what
Say it like this
"A second team's onboarding only means something once the first team has real, trusted failure cases already in the system. Building a generic schema before that is guessing."
Why this works
Shows part of the build order is forced by what each piece actually needs to be real.
Stage 6
Give the rank, and close
Say it like this
"So: one team's real failure cases first, earn their trust, generalize for a second team, then self-serve onboarding and dashboards last, once real use has proven the core."
Why this works
Restates the direct answer as an actual sequence, ready to defend if someone pushes on the order.
Let's learn
ProofBench is a shared platform that runs a team's AI feature against a set of real cases with known correct answers, and reports whether it passes.
Before it, each team at Gullford tested their own model informally, spot-checking a handful of outputs by eye before shipping. Ines built ProofBench's first version around the loss-prevention team's shrink-detection model, using actual flagged incidents from the past year as the test set.
Customer support team's reported ProofBench pass rate, vs. actual production error rate
The reported score and the real world stopped agreeing almost immediately. Nobody was checking whether they still matched.
The turn: the climbing pass rate was never proof the support bot was improving. It was proof the support team had learned which test cases made the number go up, which is a completely different skill from making the product better.
The first step is the only one that doesn't depend on anything else existing yet.
The decision that mattered
ProofBench showed every team's pass rate on a shared, leadership-visible dashboard from day one. That felt like healthy transparency at launch. It quietly turned the eval set into something worth curating, not just something worth passing.
At its worst: a team spends five months pruning its own test set instead of fixing its bot, reports a pass rate near ninety-six percent, and the actual production error rate never moves, because the two numbers stopped describing the same thing months earlier.
All three were true on Gullford's support team, for months, before anyone noticed.
What I would leave alone: the loss-prevention team's original golden set, pulled from real flagged incidents, never needed this level of concern. Nobody had an incentive to prune it, since it was built from real losses the company had already suffered, not a number anyone was being measured against.
The lesson: a shared eval platform's biggest risk isn't a wrong verdict on day one. It's a slow, plausible-looking drift where the number improves and the product doesn't, and nothing about the dashboard tells you which one is happening.
Now here is the same thing as a story
The short version above is what you'd say pitching ProofBench's build order to Gullford's engineering leadership. Read this one for how the drift actually got caught.
Ines Delacroix built ProofBench's first version around a real, painful problem: the loss-prevention team's shrink-detection model kept missing a specific kind of theft pattern, and nobody had a reliable way to measure whether a fix actually worked.
That first golden set, built from real flagged incidents, earned trust fast. Within three months, the loss-prevention team was gating every model change on ProofBench's verdict. Word got around, and the customer support team, building a bot to draft first-pass replies to routine complaints, asked to onboard next.
Knowledge spark: what makes a test case "golden" instead of just a test case?
A golden case is a real example with a known, correct answer that isn't up for debate. A test case someone invented to look plausible isn't golden, it's a guess dressed up as evidence.
The support team's onboarding went smoothly, on paper. Their pass rate climbed steadily for months. Then a new hire, three weeks into the data team, pulled up both the ProofBench dashboard and the actual customer complaint reopen rate, side by side, out of simple curiosity.
One of these you can still walk back easily. The other, once it happens, doesn't undo.
"Why does the pass rate keep going up," she asked in a team meeting, "when the reopen rate on customer complaints hasn't moved at all?" Nobody had an answer. It turned out the support team, watching their own number sit on a company-wide leaderboard next to loss-prevention's, had quietly been removing the ugliest, hardest-to-handle complaint examples from their test set every time the bot failed one, rather than fixing the bot.
The platform didn't lie. It faithfully reported the pass rate of a test set that had been slowly emptied of everything hard, one case at a time, by people who never meant to cheat, just to look better on a number everyone could see.
Here's the decision I'd take back: putting every team's pass rate on one shared, leadership-visible dashboard from the start. That felt like the transparent, healthy choice when only loss-prevention used the platform and nobody had reason to game a number built from real losses. It stopped making sense the moment a second team's incentive to look good outweighed their incentive to actually be good.
Replayed with pass rates kept private to each team's own leads, visible company-wide only as "onboarded, gated, and passing" or not, with no comparable number to game: the support team's test set stays intact, since there's no leaderboard rewarding a curated one. The reopen rate and the pass rate move together instead of apart, and the new hire's question never needs asking.
I built the shared leaderboard because I wanted every team to feel proud of using the platform, the same way loss-prevention did. It took one honest question from someone three weeks into the job to see that pride and gaming can look identical on the same chart.
ORDER, mapped onto one platform's rolloutNot a feature roadmap. ORDER is what tells you which capability breaks everything else if you build it first.
O
Outcome. What everything competes to move.
Whether a team genuinely gates their own launches on ProofBench's verdict, or quietly ships around it.
Without this named first, ranking platform features is just taste.
R
Reversibility. The hardest gap to undo.
A trust-destroying wrong verdict during a team's first real use. Once caught, that team stops trusting the platform for good.
The hardest step and the direct answer's foundation: build first against the thing you can't take back.
D
Dependency. What unblocks what.
A second team's onboarding only means something once the first team's real, trusted failure cases already exist in the system.
Part of the build order is forced by what each piece actually needs to be real.
E
Evidence. What's cheap to learn early.
Piloting with one motivated team's real failure cases costs far less than building generic multi-tenant infrastructure nobody's tested against reality yet.
Shows you'd learn something before spending a quarter building the wrong thing well.
R
Rank. The actual sequence.
Real failure cases first, earn trust, generalize the schema, then self-serve onboarding and dashboards last.
States the order and defends the top pick in one line.
The fourth callout is the one Gullford's support team quietly broke, one deletion at a time.
The leaderboard looked cheap and safe to build early. It was the one that quietly broke the platform's honesty.
The recap, one line per letter: outcome is real gating versus quiet workaround, reversibility is a trust-destroying wrong verdict, dependency is a second team needing the first team's real cases to exist first, evidence is a cheap single-team pilot, and rank is the four-step build order itself.
Months until a team's first genuinely trusted verdict, by build order
Building the impressive-looking infrastructure first didn't just delay trust, it delayed it by nearly a year.
And if you want to be sure it really works, try it somewhere elseSame five letters, a regional cinema chain instead of a grocery chain. Nothing about the two businesses is alike.
Palisade Cinema Group built a shared eval platform starting with its concession-demand forecasting model, before opening it to the staffing-prediction team. Dante Marchetti led that rollout.
Mapped onto ORDER: outcome is whether a second team genuinely gates decisions on the platform's verdict, same as Gullford. Reversibility: a wrong forecast verdict trusted blindly during a holiday weekend, before anyone had learned to double-check it, would have been the hardest thing to walk back, since a bad opening weekend of concession stockouts doesn't get a do-over. Dependency: the staffing team's onboarding only made sense once the concession team's golden set, built from three years of real holiday-weekend sales data, already existed and had proven itself. Evidence: piloting with just concessions, one team, one real problem, cost far less than building a platform for every team's hypothetical future use case. Rank restates the identical order, in a business with nothing else in common with a supermarket chain.
A different shape of picture than Section 2 used: a branching decision instead of two doors side by side. The same order still wins.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "one real team's failure cases first, earn trust, then generalize, dashboards last, because a wrong early verdict is the one mistake you can't undo," and stop.
Cost: if a full golden set is too expensive to build for every case type at once, start with just the highest-volume failure pattern and expand from there.
The model gets better, for real: even as underlying models improve, a shared leaderboard still creates the same incentive to curate a test set rather than fix a product, so the ranking risk doesn't go away just because accuracy does.
Where people run it wrong.
They build polished dashboards and self-serve onboarding first, because it's the most demoable part, before any team has a reason to trust the underlying verdict.
They put every team's score on one visible leaderboard, assuming transparency is automatically healthy.
They treat a climbing pass rate as proof of improvement, without ever checking it against a real-world number it's supposed to predict.
How to use it live. When someone asks how to build an eval platform other teams will use, ask yourself first: which mistake here can't be undone? Build against that one first, and let everything demoable wait.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "how would you build an internal eval platform that other teams actually use"?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Reversibility is the hardest step: a trust-destroying wrong verdict can't be undone.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ines Delacroix, who built ProofBench at Gullford Grocers starting with the loss-prevention team's shrink-detection model.
3 · THE HABIT
What did the support team start doing once their score sat on a visible leaderboard?
Tap to flip
ANSWER
They quietly started pruning the ugliest, hardest complaint examples from their test set instead of fixing the bot, whenever it failed one.
4 · THE FLIP
What's the two-setting switch here?
Tap to flip
ANSWER
Submitting real, including ugly, failure cases to be genuinely evaluated, versus quietly curating the test set to make the visible score look better.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Putting every team's pass rate on one shared, leadership-visible dashboard from day one, which turned the test set into something worth curating.
6 · THE NUMBER
Fill in the blank: the support team's reported pass rate climbed to 96 percent while actual production error rate stayed flat at about ___ percent.
Tap to flip
ANSWER
12 percent. The two numbers had stopped describing the same thing months before anyone noticed.
7 · THE REPLAY
Same onboarding, private scores instead of a shared leaderboard. What changes?
Tap to flip
ANSWER
The support team's test set stays intact, since there's no visible number to game. The reopen rate and the pass rate move together, and the new hire's question never needs asking.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what stayed the same?
Tap to flip
ANSWER
Palisade Cinema Group's concession-forecasting platform. The same four-step build order held: real cases first, trust earned, then generalize, dashboards last.
Check yourself Score: 0 / 0
Multiple choice
1. Why build one team's real failure cases before a generic multi-team schema?
A. It's cheaper to write generic code first.
B. Leadership prefers to see one team succeed before funding more.
C. A second team's trust depends on evidence from a first team's real, proven use, not a promise about the design.
D. Real failure cases are required by most compliance frameworks.
Show hint
Look at the dependency step.
Show answer
C. Dependency in ORDER is about what actually needs to exist before the next piece means anything, not cost or compliance.
True or false
2. True or false: a climbing ProofBench pass rate always means a team's AI feature is genuinely improving.
True
False
Show hint
Look at the two-line chart comparing pass rate and production error rate.
Show answer
False. The support team's pass rate climbed while their real production error rate stayed flat, because the test set had been quietly pruned rather than the product improved.
Fill in the blank
3. Fill in the blank: building dashboards and infrastructure first took ___ months to reach a team's first genuinely trusted verdict, versus 3 months building real cases first.
Show hint
Look at the bar chart.
Show answer
14 months. Building the impressive-looking infrastructure first delayed genuine trust by nearly a year.
Short answer, where it wouldn't matter
4. Name a part of ProofBench's build where the gaming risk described in this answer barely applies.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The loss-prevention team's original golden set, built from real flagged incidents with no visible leaderboard pressure attached to it.
Short answer, apply it yourself
5. Pick a metric you've seen tracked at work. What would make it easy to game without anyone intending to cheat?
Show hint
Think about a number that's visible to leadership and easier to move than the real thing it's supposed to represent.
Show answer
Model answer: A visible support-ticket resolution time metric is often easier to improve by closing tickets quickly than by actually solving the customer's problem.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Putting every team's score on one shared, visible dashboard. It felt transparent when only one team, with no gaming incentive, used the platform, and broke once a second team had a reason to look good instead of be good.
Before you close the answer
Why this works
Tests whether you'll rank platform features by what actually breaks trust irreversibly, or default to listing impressive infrastructure. Most candidates describe dashboards and pipelines and never mention gaming risk.
Follow-up traps
"Isn't a visible leaderboard good for driving adoption?" Response: it can be, once gaming risk is handled; visible without protection just rewards whoever curates their test set best, not whoever builds the best product.
"How would you have caught the drift sooner, without waiting for a new hire's question?" Response: pair every reported pass rate with its real-world outcome metric on the same dashboard, so a gap between the two is visible by design, not by luck.
If pressed
ProofBench's actual fix didn't remove the leaderboard entirely. It replaced the raw pass-rate number with a "cases added this quarter" count sitting next to it, so a team curating its set downward becomes visible on its own, instead of only a rising score being visible.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.