Artifact critiqueAdvancedAI Opportunity & Model Strategy / Model selection from a PM lens / #19

Build the scorecard you would use to present a model recommendation to leadership.

SPARKnobody had agreed on a single row this scorecard should have

Ashgrove Metrics sells an observability dashboard to engineering teams, and it recently added an AI feature that reads an incident's logs and drafts a root-cause summary. Freya Danthorpe is the senior PM who has to decide, every time a new model is worth considering, what leadership actually needs to see before they approve it.

The direct answer
Build one fixed card, five rows, reused every single time: task fit against a real example, cost per call at your actual expected volume, a latency number checked against the feature's own requirement, the failure mode and its guardrail, and a named owner with a re-review date. Fill in the same five rows every time a model gets evaluated, never a fresh deck, so this quarter's recommendation can be checked against next quarter's, and so someone other than you can build it and still be trusted.
Do this, in order
  1. Fix the five rows once, and never let the format change quarter to quarter.Why: a scorecard that changes shape every time can't be compared to the one before it.
  2. Put a real latency and cost number in every row, at your actual volume, not a vendor's demo numbers.Why: a model can pass every quality question and still blow a constraint nobody wrote down.
  3. Name the failure mode and its guardrail for every candidate, not just the winner.Why: leadership is approving a risk, not just a score, whether they say so or not.
  4. Give every card a named owner and a re-review date.Why: a recommendation with no owner is nobody's job to revisit when the model changes.
  5. Let a junior analyst fill in the card, and review it in twenty minutes instead of rebuilding it.Why: a fixed template is what makes delegation trustworthy instead of a gamble.
  6. Don't collapse the five rows into one weighted score.Why: leadership needs to see which specific row they're trusting, not a single number hiding the judgment call.

How to answer this, stage by stage

Nobody is scoring whether your slide looks polished. They're scoring whether the same five things get checked no matter who builds it or how rushed the quarter is.

Stage 1
Scope it to one real recommendation
Say it like this
"Let's ground this in Ashgrove's incident-summary feature. Every time a new model gets considered for it, I need one artifact that leadership can approve in one meeting, not a fresh argument each time."
Why this works
Stops the answer from turning into a list of generic scorecard categories with nothing real behind them.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as SPARK. Situation, how this gets decided today without a scorecard. Payoff, the habit I want it to build. Anchor, the actual five rows. Risk, what breaks the first time a recommendation is wrong. Keep out, what I won't try to cram onto one card."
Why this works
Signals a repeatable design method for the artifact, not just good taste in slide layout.
Stage 3
Reframe: it isn't "which model is better," it's "what does leadership need to see the same way, every time"
Say it like this
"This isn't really a question of picking a winner. It's a question of what five things have to be true before anyone approves a model, stated so plainly that two different quarters' cards can sit side by side and mean the same thing."
Why this works
This is where a strong answer separates from someone who just describes a nice-looking comparison table.
Stage 4
Give the one decision: the anchor
Say it like this
"Here's the anchor: five fixed rows, always in this order. Task fit against a real incident. Cost per call at our actual volume. Latency against the feature's own requirement. Failure mode and its guardrail. Owner and re-review date. Every candidate gets the same five rows, filled with real numbers, not a paragraph of vibes."
Why this works
This is the direct answer, drawn as an actual artifact you could hand someone, not a description of "good criteria."
Stage 5
Prove it with the compressed evidence, and name the trade-off
Say it like this
"We used to skip the latency row when the demo sounded sharp enough. One quarter that cost us a model that answered in 3.4 seconds against an 800-millisecond target, caught only after launch. The fixed card costs Priya about four hours to fill in and me twenty minutes to review, more upfront discipline than a freeform slide, in exchange for never missing a constraint like that again."
Why this works
Compresses the whole case into the one missing row that actually caused the failure, and names the real cost of fixing it.
Stage 6
Say what you'd keep out, then close
Say it like this
"I wouldn't collapse the five rows into one weighted score, since that hides exactly which judgment call leadership is actually trusting. For Ashgrove, the card holds: same five rows, every quarter, whoever builds it."
Why this works
Closes with real judgment about what the artifact deliberately refuses to do, and restates the direct answer in one breath.

Let's learn

Picture the meeting before anyone had agreed on a single row this scorecard should have.

Before any fixed format, Freya built the whole model-comparison case herself every quarter: pulled sample incidents, timed responses, priced out calls, and built a new slide deck from scratch, about two full days of work each time.

Once she started delegating it to Priya Anand, a newer analyst on her team, a quarter's comparison came back in about four hours of Priya's time, and Freya's own review dropped fast: three hours the first quarter, forty five minutes the second, and by the third quarter she forwarded the deck to leadership without opening it at all.

Hand sketched icon list titled SPARK the five letters. Five rows: Situation how the work gets done today without you. Payoff the habit you want this to build, shown in a different color. Anchor the one decision everything hangs on. Risk what breaks the first time you're wrong. Keep out what you won't build on day one.
The five letters, held up as one page. Anchor is the row that has to survive contact with a real quarter.

Here's the turn: the real problem was never that Priya's work got worse. It was that "compare the models" and "build the leadership deck" had been merged into one loose task with no fixed shape, so each quarter's deck quietly measured different things, and nobody was checking whether they still matched.

Freya's review time per quarterly comparison
200 min 100 0 0 min, unreviewed 20 min, fixed card Quarter 1 Quarter 2 Quarter 3 After the card
Review time didn't shrink gradually toward a healthy number. It fell straight to zero, then had to be rebuilt from a different design.

At its worst, a model gets recommended and approved with no one checking whether it can even meet the feature's own speed requirement, and the gap surfaces only after the feature is already live.

The scorecard didn't get worse. It stopped being a scorecard and became a form letter nobody read.
The choice I would take back Ashgrove's original process merged "compare the candidate models" and "build the leadership presentation" into one open-ended task with no fixed template. That made sense when Freya did both steps herself and held the whole comparison in her head. It stopped making sense the moment a second person took over the work and had nothing written down to match against.

What I would leave alone: I wouldn't force every internal, low-stakes experiment through the same five-row process, since a quick internal test with no leadership approval attached doesn't need the same weight as a production recommendation.

The lesson: a scorecard is not a nicer slide deck. It is the one artifact that lets someone else do this job as well as you did, without you checking their work line by line.

Now here is the same thing as a story

The short version above is what you'd say walking someone through the artifact. Read this one for what it felt like the week a new hire's question exposed how thin the process had gotten.

Freya could look at a rough model comparison and tell within a minute which row was doing real work and which was padding.

When she first handed the quarterly comparison to Priya, it felt like a clean win: Priya was sharp, thorough, and the first deck came back better organized than anything Freya had built solo. Freya read every slide carefully, checked the numbers herself, and signed off with confidence.

Hand sketched flow diagram titled Today, without a scorecard, fifth step emphasized. Five steps left to right: Compare models in a Slack thread. Pick a gut favorite. Build a one off slide. Present, hope it lands. Redo it from scratch next quarter.
Five steps, and the fifth one is where the whole process quietly resets itself every quarter.

By the second quarter, Freya was busier, and she skimmed Priya's deck in forty five minutes instead of building her own confidence in it from scratch. The numbers looked reasonable. She signed off.

Knowledge spark: why would a model pass every quality check and still fail leadership's real bar? A model's quality on a sample question doesn't tell you its cost at your real volume, or its speed against your feature's own requirement. Those are separate facts that have to be checked every time, on purpose, because a model can be excellent at one and quietly wrong on another.

By the third quarter, a deadline was tight. Freya forwarded Priya's deck straight to leadership, unread, trusting that the same good instincts from quarter one were still guiding it.

Hand sketched comparison titled The anchor, close up. Left panel, a document icon labeled Before, caption a fresh slide deck invented each quarter. Right panel, a gauge icon labeled After, caption the same five row Model Fit Card every time.
The actual design decision, close enough to argue with: five fixed rows, never a fresh invention.

A newly hired analyst, reading through the last four quarters' decks to get up to speed, asked Freya a plain question: "why does this quarter's deck skip latency completely? Last quarter had a whole slide on it." Freya didn't have a good answer, because she hadn't read the deck closely enough to know it was missing.

Hand sketched decision tree titled The day the recommended model is wrong. Root, chosen model underperforms after launch. Three branches: card has a latency row leads to gap was flagged before ship. Card has no latency row leads to nobody knew until launch. No card ever existed leads to redo the whole case from memory.
Whether the card has a latency row at all decides how badly a wrong pick surfaces, and when.

The model that quarter's deck had recommended, chosen without a latency row, answered incident summaries in 3.4 seconds against an 800 millisecond target for the live chat feature it was meant to power. Nobody found out until the feature had already shipped and engineers started complaining that the assistant felt sluggish.

Freya reclaimed the whole process herself after that, back to two full days a quarter, plus time spent explaining to Priya what had gone wrong, which came out to more total work than either extreme on its own.

The real question was never whether Priya was careful enough. It was whether the process gave her, or anyone, a fixed shape to fill in, so a missing row would be obvious the moment it was empty instead of invisible inside a slide that looked complete.

Hand sketched labeled parts diagram titled What's in the Model Fit Card. A document icon at the center labeled Model Fit Card, with five labeled callouts around it: Task fit, Cost per call, Latency budget, Failure mode, Owner and re-review date.
Five named parts. A card missing any one of them is visibly incomplete, not just thin.

When the delegation was first set up, someone said, "Priya's sharp, she'll figure out what to include," and it sounded reasonable, since Priya really was sharp. Nobody had written down what "include" actually meant.

Model latency against the live chat feature's requirement
4000ms 2000 0 requirement: 800ms 3400ms Chosen without latency row 650ms Chosen with the Model Fit Card
One model missed the requirement by more than four times over. The other one wasn't even close to the line.

Rerun the same quarter with the Model Fit Card in place: Priya fills in all five rows in about four hours, the latency row flags the 3.4 second model before it ever reaches a slide, and Freya reviews the finished card in twenty minutes, checking that the rows are real, not rebuilding the case herself.

What I'd tell myself, hearing that new hire's plain question land: a scorecard with no fixed rows isn't a lighter process. It's the same process, minus the one thing that made it trustworthy enough to hand to someone else.

SPARK, the artifact that makes delegation survive a busy quarterNot a script for never trusting a junior analyst again. SPARK is what tells you exactly which five things they need in front of them.

S
Situation. How does this get decided today, without you?
A Slack thread comparison, a gut favorite, and a one-off slide deck built fresh by whoever has the time that quarter.
Grounding in the real, messy default keeps the anchor from becoming an abstract wish list.
P
Payoff. What habit do you want this to build?
Trust that any completed card, from any analyst, means the same five things were actually checked, so leadership can approve it without re-deriving the case themselves.
The habit, not the pretty slide, is the actual thing worth designing for.
A
Anchor. The one design decision everything hangs on.
Five fixed rows, always in the same order: task fit, cost per call at real volume, latency against the feature's requirement, failure mode and guardrail, owner and re-review date.
This is the hardest step, and the answer to the question, stated as a real artifact, not a philosophy about rigor.
R
Risk. What breaks the first time you're wrong?
If a recommended model underperforms after launch, the card's owner and re-review date make it someone's clear job to catch it, instead of a mystery nobody claims.
The anchor has to visibly survive its own risk, or the card is just a nicer-looking guess.
K
Keep out. What you deliberately won't build yet.
No single weighted score collapsing the five rows into one number, and no mandatory card for quick internal experiments with no leadership approval attached.
Naming what stays out shows judgment instead of process for its own sake.

The recap, one line per letter: situation is the ad hoc Slack thread and fresh slide deck built each quarter, payoff is trusting any analyst's completed card the same way, anchor is the five fixed rows in a fixed order, risk is the owner and re-review date that catch a wrong call, and keep out is the single weighted score and the low-stakes exemption Ashgrove deliberately isn't building.

And if you want to be sure it really works, try it somewhere elseSame five letters, a veterinary imaging tool instead of an observability dashboard. The rows barely change.

Odalys Ferraro runs product at Ashmere Veterinary Group, deciding which model reads x-ray images to flag likely fractures for a vet to confirm. Mapped onto SPARK: situation is a vet manually reviewing every x-ray with no assistance, about six minutes each. Payoff is trusting a flagged image enough to review it first, without re-checking every calm image the model already cleared. Anchor: the same five rows, task fit against a set of real, vet-confirmed x-rays, cost per scan at the clinic's real monthly volume, latency against the time a vet is willing to wait mid-exam, the failure mode of a missed fracture and its guardrail, a named clinical owner and a re-review date tied to each new model version. Risk: a missed fracture is dangerous, so the anchor needs a conservative default, an uncertain image still gets flagged for review rather than cleared silently. Keep out: no single combined risk score hiding which specific row is driving a recommendation.

Hand sketched timeline titled When the card gets reused, third milestone emphasized. Four milestones: Launch, card built once. Quarter 2, same card, new numbers. New hire question, gap gets caught early, shown in a different color. Quarter 4, trend visible across cards.
Different industry, same rhythm: the card gets reused, not reinvented, every time a new version comes up.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "five fixed rows, task fit, cost, latency, failure mode, owner, every time," and stop.
Cost: no budget for a dedicated analyst to build the card every quarter. Say so honestly, and cut the card down to three rows before cutting the discipline of filling it in consistently.
The model got better, for real: if a new model claims a lower cost and the same quality, that's still worth a fresh card, since a cheaper model with a worse failure mode is not automatically the better recommendation.

Where people run it wrong.
They build a beautiful one-off deck for the first recommendation and never standardize it, so no two quarters can be compared.
They collapse everything into a single weighted score, hiding the actual judgment call from the people approving it.
They let review depth quietly shrink each quarter without ever noticing, since nothing marks the moment it happened.

How to use it live. The moment an interviewer asks you to build a scorecard, ask yourself: what five things would I need to see, in writing, before I'd stake my own name on this recommendation? Name those five, in a fixed order, and the rest of the artifact follows on its own.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Delegation flip: Freya handed the quarterly comparison down to Priya, and when quality dropped, she took it back, doing both the original work and a conversation about why.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Freya Danthorpe, senior PM at Ashgrove Metrics, who built a fixed Model Fit Card after delegating the recommendation process broke down.
3 · THE HABIT
What did Freya stop doing because delegation worked at first?
Tap to flip
ANSWER
She stopped reading Priya's decks closely, going from a careful three hour review to a forty five minute skim to forwarding a deck unread.
4 · THE ANCHOR
What's the actual anchor decision in this answer?
Tap to flip
ANSWER
Five fixed rows, always in the same order: task fit, cost per call, latency, failure mode and guardrail, owner and re-review date.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging "compare the models" and "build the leadership deck" into one loose task with no fixed template, instead of splitting them into a repeatable card.
6 · THE NUMBER
Fill in the blank: the model chosen without a latency row measured ___ milliseconds against an ___ millisecond requirement.
Tap to flip
ANSWER
3,400 milliseconds against an 800 millisecond requirement.
7 · THE REPLAY
Same quarter, Model Fit Card in place. What changes?
Tap to flip
ANSWER
The latency row flags the slow model before it reaches a slide, Priya fills the card in four hours, and Freya reviews it in twenty minutes instead of rebuilding the whole case.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product. Which one, and what stays the same?
Tap to flip
ANSWER
Ashmere Veterinary Group's x-ray fracture flagging tool. The rows barely change: task fit, cost, latency, failure mode, owner, still in a fixed order.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Merging "compare the models" and "build the leadership deck" into one loose task with no fixed template. It made sense while Freya did both steps herself and held the whole comparison in her head, and stopped making sense once a second person took it over with nothing written down to check against.
Multiple choice
2. What was the flip in Freya's story?
  • A. She decided Priya wasn't good enough for the job.
  • B. She went from carefully reviewing every delegated deck to forwarding one unread, then swung back to reclaiming the whole process herself.
  • C. She asked engineering to slow down the model's response time.
  • D. She stopped using Slack for model comparisons.
Show hint
Look at "the habit thinning" across three quarters.
Show answer
B. The delegation flip: trust thinned in three beats until she reclaimed the work entirely, doing more total work than either extreme alone.
Fill in the blank
3. Fill in the blank: Freya's review time went from ___ minutes in quarter one, to ___ minutes in quarter two, to ___ minutes in quarter three.
Show hint
Look at the line chart of review time per quarter.
Show answer
180 minutes, 45 minutes, 0 minutes. A steady decline to nothing, with no single moment that felt like a decision to stop checking.
True or false
4. True or false: this answer recommends collapsing the five rows into one weighted score so leadership can compare candidates faster.
  • True
  • False
Show hint
Look at the K step, keep out.
Show answer
False. A single collapsed score hides which specific row leadership is actually trusting, which is exactly the judgment the card exists to make visible.
Multiple choice, where it wouldn't matter
5. Where does this answer say the full five-row process is NOT worth requiring?
  • A. Any model that costs less than a competitor's model.
  • B. A quick internal experiment with no leadership approval attached.
  • C. Any model used outside the incident-summary feature.
  • D. Models built by a vendor instead of in house.
Show hint
Look at "what I would leave alone."
Show answer
B. A low-stakes internal test with no approval on the line doesn't need the same weight as a production recommendation.
Short answer, apply it yourself
6. Think of a recurring decision at work or school that gets redone from scratch every time, like a status report or a budget request. What five fixed things could turn it into a reusable template?
Show hint
Think about what would need to stay identical for two different versions to be fairly compared.
Show answer
Model answer: A monthly budget request: last month's actual spend, this month's ask, the reason for any change, what gets cut if it's denied, and who approved the last one. Fixed rows turn a fresh argument into a five-minute check.
Before you close the answer
Why this works
Tests whether you'll design an artifact that survives being handed to someone else, instead of one that only works as long as you personally build it every time.
Follow-up traps
"Isn't a fixed template too rigid for a genuinely novel model?" Response: the five rows are fixed, the content inside them isn't. A genuinely novel model still needs a task-fit number and a latency number, it just might score differently on them.

"Couldn't you just trust a senior analyst's judgment instead of a template?" Response: judgment is exactly what a fixed template protects, since even a strong analyst's review depth quietly thinned over three quarters with nothing written down to catch it.
If pressed
The re-review date on each card is set automatically to 90 days out or the vendor's next announced model version, whichever comes first, so a card never quietly goes stale without anyone noticing.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more