CaseAdvancedModel Fluency & the AI PM Role / Working with ML engineers and researchers / #12
Describe how you would build shared ownership of the eval set across PM, engineering and design.
SPARK · the monthly session that let a grant writer's page-break catch beat Proviso's own accuracy score to the truth
Tarnholm Systems builds Proviso, an AI tool that checks a nonprofit's grant application against a funder's rules before anyone submits it. Nyoka Ostrell is the AI PM who owns Proviso's roadmap. Iskren Kirrick leads the engineering team that built and, until recently, alone maintained Proviso's eval set. Almendra Lundgrave leads design, and for most of Proviso's first year had never once opened it.
The direct answer
Build a monthly session where the PM, the engineering lead, and the design lead review a sample of real outputs together and each nominate one new eval case: the PM from user feedback, design from a moment in the rendered report that looked wrong, engineering from something odd in a production log. Vote every nominated case into the shared eval set in that same session, on the spot. Nobody outside engineering decides how a case gets fixed. Everybody decides what belongs in the set.
Do this, in order
Put a fixed monthly session on the calendar, all three roles required.Why: this is the actual mechanism that stops the eval set drifting back into one team's private artifact.
Require each role to bring one real case, sourced from their own vantage point.Why: a PM sees what officers complain about, design sees what a report actually looks like on the page, engineering sees what the model actually gets wrong. No single role sees all three.
Vote every nominated case into the shared eval set in the room, not after.Why: a case that only gets talked about, not formally added, gets forgotten by the next session, the way the page-break case nearly was.
Keep the technical fix, the rule, the retrain, the layout change, with engineering alone.Why: shared ownership of what gets tested is not shared ownership of how a model gets built.
Track category coverage, not only accuracy, and watch for a category sitting at zero.Why: a rising accuracy number can sit on top of a whole class of failure the metric was never built to see.
Leave engineering's own regression tests, the deterministic parsing checks, entirely alone.Why: a form field either extracts correctly or it doesn't. That's not a judgment call for three people to debate.
1How to answer this, stage by stage
Nobody is grading whether you can say "we should all care about quality." They are grading whether you can name the actual room, the actual question asked in it, and the actual line between contributing to an eval set and deciding how a model gets fixed.
1
Scope it to one product, three named roles
Say it like this
"I'll ground this in one team. Tarnholm Systems builds Proviso, a tool that checks grant applications for compliance. I'm the PM. Iskren Kirrick leads engineering. Almendra Lundgrave leads design."
Why this works
Stops the answer from turning into a lecture on cross functional trust, which is a different, weaker question.
2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Situation, who actually owns the eval set today. Payoff, the habit I want it to build. Anchor, the one concrete session. Risk, what a stale eval set costs if I get this wrong. Keep out, what PM and design never touch."
Why this works
Two seconds of structure tells the interviewer you have a plan, not just a feeling about "everyone caring more."
3
Reframe what "shared ownership" is actually testing
Say it like this
"This isn't really 'how do three roles get along.' It's whether the eval set can catch a failure that a text accuracy score was never built to see, before a real nonprofit finds it by hand, at night, an hour before a deadline."
Why this works
Separates a real design answer from a vague endorsement of teamwork.
4
Give the anchor, the actual mechanism
Say it like this
"Once a month, forty five minutes, the three of us look at five real Proviso outputs together. I bring a case from a support ticket. Almendra brings a case from something that looked wrong on the printed page. Iskren brings a case from something odd in the logs. We vote every case into the shared eval set before anyone leaves the room."
Why this works
This is the direct answer, said as a rule someone could actually run, not a description of a value like "collaboration."
5
Name the two ways this breaks if you skip it
Say it like this
"Skip this and it goes one of two ways. Either the eval set stays Iskren's alone, and it never learns to check anything Almendra or I would have caught. Or we try to review everything together with no structure, and nothing ships, because three people end up arguing about implementation instead of nominating cases."
Why this works
Naming both failure directions tells the interviewer you're not about to over correct into the opposite mistake.
6
Prove it against the near miss
Say it like this
"Here's why this mattered. Ravensmoor Youth Trust used Proviso to check their application to the Underleigh Foundation. Every field passed. But the certification statement got split across a page break, and Proviso's eval set only ever tested the extracted text, never the printed page. A grant writer named Xenna caught it by printing a hard copy at 9:40pm, about eighty minutes before the portal closed."
Why this works
One real near miss carries more weight than any abstract claim about a missing category of coverage.
7
Name the keep out, then close on one line
Say it like this
"I'd never let Almendra or me decide how a flagged case actually gets fixed, that stays with Iskren's team. And I'd never spend this session on deterministic parsing bugs that already pass or fail cleanly, those don't need three people in the room. So: one monthly session, one nominated case each, voted in on the spot, and the fix always stays engineering's call."
Why this works
Shows judgment about where the idea stops, then restates the whole decision in one breath.
2Let's learn
This is the gap the monthly session exists to close. Not a lack of care in the room. A lack of any road for what design or the PM already notice.
Proviso reads a nonprofit's grant application, checks it against a funder's rulebook, and marks every rule PASS or FAIL with the exact line it's flagging. Before Tarnholm built anything, a grant writer at a place like Ravensmoor Youth Trust checked her own application by hand against a rulebook of about forty items: page limits, required certifications, budget math, font size, section order. That took her close to three hours, usually the night before the deadline, the way it always got done.
Knowledge spark: what is an eval set, for a tool like this?
A pile of real applications with the right answer already worked out by a person: which rules should pass, which should fail, and why. Engineers run Proviso against that pile after every change, to check it still gets the right answer, before anything reaches a real nonprofit.
Proviso now reads the same application in under two minutes. Extraction accuracy, whether it reads the right words off the right line, climbed from 89 percent to 98 percent over Proviso's first nine months. Iskren's team built that eval set, ran it in the same repo as the code, and kept it passing every release. By every number they tracked, Proviso was getting better every month.
A model can get every word right and still fail the one thing nobody thought to check.
Here's what that accuracy number never touched. Proviso's eval set tested the text it pulled out of a document. It never once rendered a page and looked at it, because nobody had asked it to. For eight straight months, zero of Proviso's cases tested how a compliance report actually looks once it's printed or exported, the thing a grant officer or a funder's reviewer actually opens.
Same failure, same product. What changes is only whether anyone was ever asked to look at the printed page.
Ravensmoor Youth Trust ran their application to the Underleigh Foundation's youth literacy grant through Proviso. Every field passed. But a page limit trim, a routine step that removes extra white space to fit a report inside its allowed pages, had pushed the signature line of the required non discrimination certification onto its own page, orphaned, with nothing above it to say what it belonged to. The extracted text was still whole and correct, so Proviso's eval set, which only ever checked extracted text, had nothing to flag.
Xenna Menzies, the grant writer, still prints a hard copy before every submission, an old habit from years of doing this by hand. At 9:40pm, with the portal closing at midnight, she opened page two and found a lone signature line with no context above it. To a tired reviewer at Underleigh, that reads like a missing certification, not a formatting accident, and a missing certification is grounds for a flat rejection. She fixed it by hand and resubmitted with about eighty minutes to spare.
Nothing in this gap was Proviso lying. Every word it read was correct. It was simply never asked whether the page still made sense.
Proviso's accuracy, against how much of the eval set ever tested the printed page
Text extraction accuracyRender and formatting coverage
Accuracy looked healthy every single month. Coverage of anything render or formatting related sat at zero for eight straight months, the whole time Ravensmoor's near miss was possible, and only started moving once the monthly session began pulling in cases from outside engineering.
What it costs at its worst: if a certification like this had reached Underleigh uncaught, Ravensmoor risks a flat rejection on a formatting accident, not a real compliance failure, on a grant they were otherwise fully eligible for. Worse, if it happened more than once, grant writers using Proviso would stop trusting its PASS mark and go back to checking every application by hand, the exact three hours a night Proviso was built to remove.
The choice I would take back
When Proviso launched, Tarnholm decided the eval set would live entirely inside engineering's repo, reviewed only by engineers, so it stayed versioned and safe from accidental changes. That was the right call when Proviso only checked extracted text against a rulebook. It stopped being the right call once Proviso started checking formatting and layout too, because now real signal about a whole category of failure was being produced outside engineering, with nowhere designed to land. Nobody decided to exclude design and the PM on purpose. Nobody ever redesigned the arrangement once the product outgrew it, either.
What I would leave alone: Iskren's team runs its own regression tests for PDF field extraction, whether a form field is read correctly at all, entirely separate from this session. That's a deterministic pass or fail with no judgment call in it. Handing it to a three person vote would slow down real engineering work for a decision that doesn't need three people in the room.
The lesson: a model's own scorecard only ever answers the questions somebody thought to ask it. If only one team ever writes the questions, whole categories of failure stay invisible, no matter how high the score climbs.
3Now here is the same thing as a story
The short version above is what you'd actually say in the room. Read this one for why the fix had to be a standing session, not a promise to "loop design in more."
Nyoka Ostrell has run product for compliance tools for six years, and she can usually tell within a week whether a rule engine is actually solid or just quiet because nobody's tested it hard yet. When she joined Tarnholm to lead Proviso, Iskren Kirrick had already built something genuinely good: a rules engine that read a grant application and caught the kind of mistakes that used to sink applications on technicalities. She trusted it, because the numbers backed her up. Every release, extraction accuracy climbed a little further.
For most of Proviso's first year, that trust felt earned. Nyoka watched the eval set's pass rate every Friday from a dashboard Iskren's team maintained. Ninety one percent. Ninety four. Ninety seven. She never opened the eval set itself, the actual pile of test cases underneath the number, because there was no reason to. It lived in a repo she didn't have much cause to visit, reviewed by engineers who plainly knew what they were doing.
Almendra Lundgrave, who led design, had it even further from her desk. She designed the report Proviso generates for a nonprofit to review before submitting, the layout, the page breaks, the print styling. She never once looked inside the eval set that was supposed to prove the tool worked, because as far as anyone had told her, that wasn't a design concern. It was a text accuracy number, and text accuracy wasn't her department.
This was the line Nyoka would eventually draw on purpose. Early on, it wasn't a line at all. It was just the whole eval set, sitting on the other side of a door nobody thought to open.
Then came the Tuesday morning Nyoka got a message from a member of Ravensmoor Youth Trust's staff, not angry, mostly confused. Their application to the Underleigh Foundation had passed every Proviso check the night before, and Xenna Menzies had still nearly missed her deadline fixing a certification page that looked, on paper, like it was missing entirely.
Nyoka pulled up the case herself. The extracted text was fine, word for word, exactly what Proviso's eval set would have graded correct. The printed page was not fine. A trim step meant to keep the report inside its page limit had pushed the certification's signature line onto page two, alone, with nothing above it. Proviso hadn't lied about a single word. It had simply never been asked whether the page still made sense once a human opened it.
We didn't miss one rule. We missed an entire way of looking at the page.
She brought it to Iskren first, expecting frustration, and instead got something closer to relief. His team had been maintaining an eval set built for one kind of question, is the text right, for over a year. Nobody outside engineering had ever asked it a different kind of question, does the page still read the way it's supposed to, because nobody outside engineering had ever been in the room where that question would occur to anyone.
Nyoka's first instinct was the tempting one: ask Iskren's team to just add rendering checks to their own eval set, quietly, without changing who owned it. Almendra talked her out of it within a day. A rendering check written by an engineer who's never sat with a grant writer's real, panicked 9:40pm moment tests what an engineer imagines could go wrong on a page, not what actually goes wrong. The people who'd catch the next one weren't in the room where the eval set got written, and adding cases from the outside without adding the outside people wouldn't fix that.
They also floated an open ticket queue instead, anyone at Tarnholm could file a suspected eval case whenever they noticed one, no meeting required. Nyoka killed that idea inside a week of trying it: three tickets came in, none got triaged, and by the second week nobody was filing them, because writing up a case nobody would look at felt like shouting into a drawer.
The whole answer in one picture. A small fixed session, one case from each vantage point, a vote before anyone leaves the room.
What they built instead: once a month, forty five minutes, all three of them sit down with five real Proviso outputs from the last few weeks. Each person is expected to bring one case, sourced from wherever they actually sit. Nyoka's usually comes from a support ticket or a comment on a call. Almendra's comes from something that looked wrong the moment she opened the rendered report, a run of text too close to a page break, a checkmark sitting next to something that visually reads as broken. Iskren's comes from a production log, a pattern in what Proviso actually got flagged for by users who complained. Every case that survives the discussion gets voted into the shared eval set before the meeting ends, no exceptions, no "we'll circle back."
The first session ran two months after the Ravensmoor near miss. Almendra brought the exact shape of the page break problem, generalized: any certification statement sitting within four lines of a page limit trim. Nine other render cases followed over the next two sessions, plus nine cases Nyoka pulled from support tickets. By month nine, formatting and render coverage in Proviso's eval set had gone from zero to sixty one percent of the known failure patterns the team could name, without touching how engineering actually builds or ships a fix.
What I'd tell myself, sitting across from Iskren that Tuesday: the mistake was never that engineering owned the eval set too tightly on purpose. The mistake was that nobody ever asked whether the questions it answered were still the right questions, once the product outgrew the ones it started with.
4SPARK, five checkpoints for who's in the room
Not a script for sounding collaborative. SPARK forces you to name the one concrete session, and prove it survives both ways a shared eval set actually breaks.
SSituation. Who owns the eval set today, and why does that break?
Nyoka Ostrell owns Proviso's roadmap. Iskren Kirrick's engineering team built and, until recently, alone maintained the eval set. Almendra Lundgrave leads design and had never opened it. Whether a real failure ever reaches the eval set depends entirely on whether an engineer happens to notice it, because nobody outside engineering has ever been asked to look.
One PM, one engineering lead, one design lead, one real product. Never a segment called "cross functional alignment."
PPayoff. What habit do I want this to build?
I want all three of us to keep bringing real cases from wherever we actually sit, and I want everyone to treat a failing eval case as a shared problem, not something that's only engineering's to notice. The habit is the product. Catching the next page break case before it reaches a real deadline is downstream of that habit, not the goal itself.
Name the thing each role stops assuming is someone else's job. That's the payoff, not the meeting's length.
Letting Ravensmoor's own staff submit suspected cases directly was tempting early. Nyoka deferred it: without a PM triaging first, most submissions would describe a symptom, not the actual page level cause.
AAnchor. The one decision everything else hangs on.
A forty five minute session, once a month, all three roles. Each person nominates one case from a real Proviso output, sourced from their own vantage point: PM from user feedback, design from a jarring render, engineering from a production log. Every nominated case gets voted into the shared eval set in that same session. Nyoka also considered letting grant officers submit suspected cases directly to the eval set. She rejected it: a submission with no PM triage in front of it usually describes what a user noticed going wrong, not what actually caused it, and would flood the session with symptoms instead of testable cases.
Concrete enough to argue with. This is the answer to the question.
RRisk. What breaks the first time it's wrong?
Run it like a status meeting where cases get discussed but never formally voted in, and it quietly becomes theater, the way the open ticket queue did in its first two weeks. Let it turn into a design review where anyone outside engineering starts dictating how a fix should work, and engineering time gets spent defending implementation choices instead of shipping. Nyoka accepts a real cost here: forty five minutes of three senior people's time, every month, forever, whether or not that particular session turns up anything dramatic, because the cost of a missed formatting category is a nonprofit's grant, not a line on a dashboard.
Not "trust dropped." What each role actually starts doing next, in either direction.
Same finding, same team, same failed render. What changes the outcome is only whether the road to the eval set exists.
KKeep out. What I deliberately will not build into this.
No decision about which detection method catches a voted-in case. No decision about whether the fix is a new rule, a prompt change, or a retrain. No decision about how the eval harness itself gets built or run. Those stay with Iskren's team, always. And this session never touches the deterministic regression tests, whether a form field extracts correctly at all has no judgment call in it worth three people's time.
Shows judgment instead of a wish for endless input. Ties straight back to Risk: the wrong kind of silence and the wrong kind of overreach are both real ways to lose this.
The recap, one line per letter: situation is an eval set only one role ever looked at, payoff is all three roles treating a failing case as shared, anchor is the monthly session and the vote that makes a case real, risk is either extreme costing months of blind spots or months of stalled shipping, and keep out draws the line at implementation, never at nomination.
5And if you want to be sure it really works, try it somewhere else
Same five letters, a textile mill instead of a grant checker, and this time the blind spot isn't a page break. It's a defect a camera model was never trained to name.
Elmswood Textiles builds Loomwatch, an AI tool that watches fabric on a conveyor belt through a mounted camera and flags defects before a bolt gets shipped to a garment maker. Praxton Yellin owns Loomwatch's roadmap. Loomwatch's eval set, like Proviso's, lived entirely inside the computer vision team's own test harness for its first year, built from thousands of labeled defect photos nobody outside engineering had ever reviewed.
Same ritual, a loom instead of a grant form. The eval set still only learns to see what someone was in the room to point at.
Mapped onto SPARK: the situation is a defect eval set that only ever learned from the defects a computer vision engineer already knew to label. The payoff is a quality floor supervisor and a line operator both contributing real cases the moment they see something Loomwatch missed, not months later in a formal incident report. The anchor is the same monthly session, one case each: the PM from a client complaint, the floor supervisor from a defect she caught by eye that never triggered a flag, the vision engineer from a pattern in near-miss confidence scores. The risk runs the same both ways: skip it and a whole class of defect, say a subtle sheen difference that only shows under certain mill lighting, stays permanently invisible to the model's own scorecard; run it as a design review and engineers spend the session defending model architecture instead of nominating cases. Keep out draws the same line: the floor supervisor helps decide what belongs in the eval set, never how the vision model itself gets retrained.
Eval cases contributed by role at Tarnholm, before the session existed vs after three of them
Before the session existedEngineering, afterPM, afterDesign, after
Engineering's own regression set didn't shrink, it stayed the backbone of the eval set. What changed is that 23 cases, nine from Nyoka and fourteen from Almendra, entered a set that had held zero cases from outside engineering for Proviso's entire first year.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, the monthly session and the vote, and what it protects against.
Cost: no budget for a standing monthly ritual across every product line. Run it quarterly instead, coverage still climbs off zero, just slower, and the PM's time cost drops to a quarter of the monthly version.
The model got better, for real: say Proviso's extraction accuracy hits 99.5 percent. The session still matters, because accuracy was never measuring the thing that broke. A better text score doesn't teach the eval set to check a printed page it was never asked about.
Where people run it wrong.
They let the session become a status update, cases get mentioned, nobody votes, and by the third month everyone quietly stops bringing anything, the way Tarnholm's open ticket queue died in two weeks.
They let a PM or a designer start dictating the technical fix inside the session, which turns a forty five minute nomination meeting into an hour and a half architecture debate nobody scheduled.
They add render or formatting cases without ever checking them against a broader held out set, and a fix built to patch one dramatic near miss quietly breaks a different, ordinary case nobody was watching.
How to use it live. Before answering a "how would you build shared ownership" question cold, ask yourself one thing: what's the smallest standing room where a real case, from any of the three roles, has somewhere to land and get voted on. Naming that room, not a value word like alignment or collaboration, is usually exactly what the question is listening for.
6Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question about designing shared ownership, not reacting to one failure?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It runs forward from how the eval set is actually owned today, instead of working backward from a single incident.
2 · THE THREE PEOPLE
Who owns Proviso's eval set in this answer, and what does each person do?
Tap to flip
ANSWER
Nyoka Ostrell, the AI PM, at Tarnholm Systems. Iskren Kirrick leads engineering and built the original eval set. Almendra Lundgrave leads design and designs the rendered report Proviso produces.
3 · THE PAYOFF
What habit does the monthly session exist to build?
Tap to flip
ANSWER
All three roles keep bringing real cases from their own vantage point, and everyone treats a failing eval case as a shared problem, not something only engineering is responsible for noticing.
4 · THE ANCHOR
What's the one concrete decision in this answer?
Tap to flip
ANSWER
A 45 minute session, once a month, all three roles. Each nominates one case from a real output. Every nominated case gets voted into the shared eval set in that same session.
5 · THE DECISION WALKED BACK
What old decision would Nyoka take back?
Tap to flip
ANSWER
Keeping the eval set entirely inside engineering's own repo and review process. That made sense when Proviso only checked extracted text. It stopped making sense once Proviso started checking layout and formatting too, and nobody redesigned the arrangement to match.
6 · THE NUMBER
Fill in the blank: for eight straight months, ___ of Proviso's eval cases tested how a report actually renders on the printed page.
Tap to flip
ANSWER
Zero. Extraction accuracy climbed to 97 percent across that same stretch, and the eval set never once measured the gap that nearly cost Ravensmoor their grant.
7 · THE RISK, SURVIVED
What breaks if the session goes wrong in either direction, and how does the anchor survive it?
Tap to flip
ANSWER
Run it like a status update and cases get mentioned but never voted in, so nothing sticks. Let it turn into an implementation debate and engineering time gets spent defending choices instead of shipping. It survives because the vote is mandatory in the room, and the fix decision always stays with engineering, never up for a vote itself.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs SPARK again on a different product. Which one, and what's the equivalent anchor?
Tap to flip
ANSWER
Loomwatch, Elmswood Textiles' camera tool for spotting fabric defects. The equivalent anchor is the same monthly session, with a PM, a quality floor supervisor, and a vision engineer each nominating a case.
7Check yourself Score: 0 / 0
True or false
1. True or false: Proviso's rising extraction accuracy score, 89 percent up to 98 percent, would eventually have caught the certification page-break problem on its own, given enough time.
True
False
Show hint
Think about what extraction accuracy is actually measuring, versus what broke on the printed page.
Show answer
False. Accuracy measured whether Proviso read the right words. The certification text was read correctly the whole time. The failure was in how the page rendered once trimmed, a category accuracy was never built to test, no matter how high the score climbed.
Fill in the blank
2. Fill in the blank: after three monthly sessions, Proviso's eval set held 210 engineering cases, ___ cases from Nyoka, and ___ cases from Almendra.
Show hint
Look at the grouped bar chart in Section 4, "Eval cases contributed by role."
Show answer
9 and 14. Nine cases came from Nyoka's side, sourced from user feedback. Fourteen came from Almendra, sourced from what actually looked wrong on a rendered page.
Multiple choice
3. Why didn't Proviso's own accuracy metric ever catch the page-break problem, even as it climbed to 98 percent?
A. Because the model was hallucinating certification text that didn't actually exist.
B. Because Ravensmoor's staff never reported the issue to Tarnholm.
C. Because the metric only tested extracted text, never how the report actually rendered as a printed page.
D. Because the eval set had gone stale and hadn't been updated in over a year.
Show hint
Look at what Iskren's team's eval set actually checked, versus what Xenna caught by printing a hard copy.
Show answer
C. The extracted text was correct the entire time. The eval set simply never asked a rendering question, because nobody outside engineering had ever been in the room where that question would come up.
Short answer, name the reversal
4. What old decision would Nyoka take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Keeping the eval set entirely inside engineering's own repo and review process, with no room for PM or design to contribute cases. It made sense at launch, when Proviso only checked extracted text and keeping the eval set inside CI, reviewed only by engineers, kept it safe from accidental changes. It stopped making sense once Proviso grew to check formatting and layout too, a category only design would ever notice going wrong.
Short answer, apply it yourself
5. Think of an AI product you use or have worked on. What's one real failure a designer or a non-technical teammate would probably notice before an engineer's own accuracy metric would flag it?
Show hint
Look for something about how an output looks, reads, or feels once a real person opens it, not whether the underlying data was technically correct.
Show answer
Model answer: A resume-screening tool's accuracy metric might track whether it correctly extracts a candidate's skills. It would never flag that the tool's rejection email reads as cold and dismissive, something a designer reading real output would catch in the first pass, and an accuracy score never would.
Short answer, work the number
6. Across the first three monthly sessions, PM and design together contributed 23 cases to a 233-case eval set. What share of the total set is that, and why does that small share matter more than its size?
Show hint
Divide 23 by 233. Then think about what category those 23 cases represent, not just their count.
Show answer
Model answer: About 9.9 percent, roughly 1 in 10 cases. The size looks small next to engineering's 210. What matters is that those 23 cases are the entire category of render and formatting coverage, which sat at zero for eight months before the session existed. A small share can still be the only share that catches a whole kind of failure.
Before you close the answer
Why this works
Tests whether you can turn "shared ownership" into an actual mechanism, a session, a nomination, a vote, instead of a value statement about caring more. And whether you know exactly where the line sits between contributing to what gets tested and deciding how a model gets fixed.
Follow-up traps
"Isn't a monthly cadence too slow if something urgent breaks?" Response: nothing stops an out-of-cycle case from being filed and escalated straight to engineering if it's a live incident. The monthly session guarantees a floor of coverage, it isn't the only door.
"What stops the PM from just voting their favorite complaints into the eval set and skewing it toward PM concerns?" Response: every case still gets voted on by all three people in the room, and design or engineering can push back on a case that doesn't hold up under real scrutiny. No single role's nominations get in unchallenged.
If pressed
Render and formatting cases get checked against an actual rendered PDF snapshot diff, not just the extracted text, specifically because that's the layer the accuracy metric can't see. That harness is versioned separately from the text-extraction eval, so a layout change can be tested without touching, or accidentally breaking, the text checks Iskren's team already trusts.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.