Artifact critiqueAdvancedAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #19

Build the vendor scorecard you would take to a procurement committee.

SPARK a scorecard is a design problem before it's a spreadsheet

Kestrel Systems builds internal tools for a mid-size logistics software company. Priya Nandakumar runs the developer platform team and is choosing an AI code-review vendor for two hundred engineers. She's building the scorecard that goes to the procurement committee before anyone signs anything.

The direct answer
Build the scorecard around one weighted number, not five separate checkboxes, and make a live pilot on your own repository worth more points than the sales demo ever can. A demo shows you the vendor's best day on someone else's code. The scorecard has to be built to catch the vendor's normal day on yours.
Do this, in order
  1. Weight a real pilot on your own repository above the demo.Why: a scripted demo tests the vendor's best case, and the committee needs to see the normal case.
  2. Put a hard security gate at the front, pass or fail, no partial credit.Why: a strong score everywhere else can't buy back a vendor that fails on data handling.
  3. Score support and incident terms as their own line, not folded into "service."Why: a vague "great support" line hides whether anyone actually answers at 2 a.m.
  4. Leave a column blank for what you deliberately did not grade yet.Why: naming what's out of scope on day one stops the committee from assuming the score covers everything.
  5. Attach a number to the total, not a color or a star rating.Why: a number can be argued with in the room; a green checkmark can't.

How to answer this, stage by stage

Nobody in the room is grading your PowerPoint skills. They're grading whether the tool you hand them would have caught the vendor that looked perfect on stage and fell apart on week three.

Stage 1
Scope it to one real committee
Say it like this
"I'll build this for Kestrel Systems, choosing an AI code-review vendor for the whole engineering org, presented to a procurement committee that includes finance and security, not just engineering."
Why this works
Keeps "build a scorecard" from turning into a generic template with no owner.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how vendors get picked today. Payoff, the habit I want the scorecard to build. Anchor, the one design decision the whole thing hangs on. Risk, what breaks the first time I'm wrong. Keep out, what I won't grade yet."
Why this works
Tells the interviewer you have a plan before you've said a single detail.
Stage 3
Ground it in what happens today, with no scorecard
Say it like this
"Right now a vendor gets picked off a slide deck and a thirty-minute demo. Whoever ran the best demo that quarter tends to win, whether or not that's the vendor that will actually hold up on our code."
Why this works
Shows the committee the real, unglamorous baseline the scorecard is replacing.
Stage 4
Give the one anchor decision
Say it like this
"The anchor is this: a paid, timed pilot on our own repository is worth more points than the demo, by design. The demo can score a vendor a 9. The pilot is the only line that can knock that same vendor down to a 4."
Why this works
This is the direct answer, said as one concrete, arguable decision.
Stage 5
Prove it with the compressed failure
Say it like this
"We almost signed a vendor whose demo ran flawlessly on a clean sample repo. On our own code, ten years of inherited style and half-finished refactors, its suggestion-acceptance rate stayed flat at under forty percent for six straight weeks. The demo never would have shown us that."
Why this works
Turns "demos can be misleading" into one specific, checkable number instead of a warning label.
Stage 6
Say what you'd measure after signing
Say it like this
"After signing, I'd keep watching the same pilot number, suggestion-acceptance rate on real pull requests, on a monthly cadence, not just at renewal. If it drifts down instead of up, that's the scorecard telling us something the contract already promised wouldn't happen."
Why this works
Shows the scorecard isn't a one-time gate, it keeps grading the vendor after the ink is dry.
Stage 7
Say what you'd leave alone
Say it like this
"I would not grade every minor IDE integration on day one. If the core review quality and the security gate both pass, a missing plugin for one editor is a rollout detail, not a reason to fail the vendor."
Why this works
Shows judgment instead of a scorecard that tries to grade everything and ends up grading nothing well.
Stage 8
Close on the one line
Say it like this
"Build the scorecard so the pilot on our own code can outvote the demo. Anything less and you've built a form that just ratifies whoever pitched best."
Why this works
Restates the direct answer in one breath, ready for a follow-up.

Let's learn

Before the scorecard existed, here is how Kestrel picked a vendor.

A platform engineer would sit through a thirty-minute demo, skim a feature comparison deck, and guess at fit based on a gut read of the sales call. That process took about two hours per vendor and produced a recommendation nobody could really defend later, because nobody had written down why one vendor beat another.

Hand sketched flow diagram titled How Kestrel picked AI vendors before the scorecard. Four boxes in sequence: Skim the deck, Watch the demo, Guess at fit highlighted, Sign the contract.
Two hours per vendor, and the weakest step was the one nobody wrote down.

Here's the turn: a written scorecard doesn't just make the choice faster. It moves the whole decision from "who pitched best" to "who held up on our own code, under our own contract terms." That's not a small change in paperwork. It's a change in what gets to count as evidence.

The anchor, the one decision everything else hangs on A paid, timed pilot on Kestrel's own repository outweighs the vendor's demo score. The demo can put a vendor near the top. Only the pilot can pull that same vendor back down, because only the pilot runs on the messy, real thing the demo was never asked to touch.
Hand sketched labeled parts diagram titled What the scorecard actually grades. A document icon at the center labeled Vendor Scorecard, with five callouts around it: security posture, accuracy on our code, pilot performance, support terms, total price.
Five lines, one number. Not five separate green checkmarks that never get added up.

At its worst, skipping the pilot line costs Kestrel a year-long contract with a vendor whose review quality never improves past a coin flip, on code two hundred engineers touch every day.

Fenwick AI's weighted scorecard total, by criterion
100 50 0 Security 27 Accuracy 16 Pilot 17 Support 9 Total: 76 / 100 security accuracy on our code pilot support
Price (7 of 10) sits below the axis start and isn't drawn to scale here; the total already includes it. The demo alone would have scored this vendor near 92.
Hand sketched comparison titled Where a demo score and a real score split. Left, a green gauge icon labeled In the demo, caption clean sample repo. Right, a red scale icon labeled On our repo, caption years of messy history.
Same vendor, same week. One score comes from a script. The other comes from ten years of real code.

The choice I would take back: an earlier draft of the scorecard folded "pilot results" and "demo quality" into one shared column called "product quality," each worth half. That made sense when the team assumed a good demo and a good pilot would usually agree. It stopped making sense the day a vendor scored a 9 on the demo half and a 4 on the pilot half, and the shared column quietly averaged them into a 6.5 that told the committee nothing true.

What I would leave alone: the scorecard doesn't need a line for every minor IDE plugin or a comparison of every pricing tier Kestrel will never reach at two hundred seats. Grading those to the same depth as security and pilot performance would just bury the two lines that actually decide the vendor.

Hand sketched icon list titled What the scorecard does not grade yet. Three items: a box icon labeled Every minor integration, a circle icon labeled Price tiers we won't reach, a funnel icon labeled Roadmap promises.
Naming what's out of scope is part of the design, not a gap in it.

The lesson: a scorecard is not a record of what a vendor showed you. It's a bet on what a vendor will do once nobody from their sales team is in the room anymore.

Now here is the same thing as a story

The short version above is what you'd walk the committee through in the meeting. Read this one for how close Kestrel actually came to getting it wrong.

Priya Nandakumar had run Kestrel's developer platform team for four years and had sat through more vendor pitches than she could count. She could tell within ten minutes of a sales call whether a rep was overselling, and her team trusted her read on almost everything.

Two vendors made the shortlist that quarter: Fenwick AI and Northlane AI. Both demoed well. Both had glossy accuracy numbers on their own benchmark decks. For the first few weeks of the evaluation, the informal read inside the team leaned toward Northlane, its demo had felt slightly slicker, its sales engineer slightly sharper on follow-up questions.

Knowledge spark: why does a demo repo behave differently than a real one? A vendor's demo repo is usually small, recent, and written in one consistent style, because that's what's easy to demo cleanly. A real company's codebase carries years of different authors, half-finished refactors, and old patterns nobody's cleaned up. A model tuned mostly on tidy code can look sharp on the demo and guess badly on the real thing, the same way a driving test on an empty lot doesn't tell you much about rush hour.

Then the pilot started. Both vendors got two weeks, unpaid trial access, running live against Kestrel's actual pull request queue, not a curated sample. Fenwick's suggestion-acceptance rate climbed steadily as engineers got used to its style. Northlane's stayed flat, and then started slipping, as its suggestions kept flagging patterns that were common, harmless quirks in Kestrel's own fifteen-year-old billing module.

Hand sketched metaphor scene titled A demo is a test drive. A pilot is owning the car. Left, a red gauge icon labeled TEST DRIVE, caption one clean loop. Right, a green box icon labeled OWN IT, caption potholes, years.
Northlane's demo was the test drive. The billing module was the pothole nobody drove over during the demo.
The demo never touched the billing module. Six weeks of real pull requests did, and that's where the two vendors actually split.

By week six, Fenwick's engineers were accepting its suggestions on nearly two out of three pull requests. Northlane's number had drifted under forty percent and wasn't moving. If Priya's team had signed off the informal, demo-driven read from week one, Kestrel would have committed a year of licensing fees to the vendor whose real-world number was going the wrong way.

The old, undocumented process would have signed Northlane. The scorecard, with the pilot weighted above the demo by design, caught the split before procurement ever wrote a check.

Hand sketched decision tree titled Where the scorecard sends a vendor. Root, Vendor score, weighted. Three branches: clears every gate leads to Buy, strong unproven on us leads to Paid pilot, fails security gate leads to Reject.
Fenwick moved from the middle branch to Buy after six weeks. Northlane never left it.

SPARK, in one screenNot a checklist of nice-to-haves. SPARK is what tells you which one design decision the whole scorecard has to protect.

S
Situation. How vendors get picked today, without the tool.
A platform engineer sits through a demo, skims a deck, and guesses at fit. Two hours per vendor, no written reasoning that survives the meeting.
Grounds the design in a real, unglamorous baseline instead of an abstract "we need a process."
P
Payoff. The habit the scorecard should build.
Stop letting the best pitch win. Start letting the best six weeks on our own repository win.
Names the actual shift in behavior the scorecard exists to cause, not just the time it saves.
A
Anchor. The one design decision everything hangs on.
A paid, timed pilot on Kestrel's own repository outweighs the demo score. The demo can lift a vendor up; only the pilot can pull one back down.
This is the answer to the question, concrete enough that a committee member could point at it and argue.
R
Risk. What breaks the first time the anchor is wrong.
If the pilot window is too short, a vendor could look bad for reasons that have nothing to do with quality, like engineers not yet trusting a brand-new tool. The design counters this with a minimum six-week window before the pilot score locks.
Shows the anchor was built to survive its own most likely failure, not just praised in the abstract.
K
Keep out. What the scorecard won't grade on day one.
Every minor IDE integration, pricing tiers Kestrel will never reach, and unverified roadmap promises all stay off the scorecard for now.
Shows restraint instead of a wish list that tries to grade everything and ends up grading nothing well.

The recap, one line per letter: situation is the two-hour, gut-feel process the scorecard replaces, payoff is trading "best pitch" for "best six weeks on our own code," anchor is weighting the live pilot above the demo, risk is protecting that anchor with a minimum pilot window, and keep out is naming three things left ungraded on purpose.

And if you want to be sure it really works, try it somewhere elseSame five letters, a grocery chain's AI restocking vendor instead of a code-review tool. A different keep-out line breaks the second story.

Pinegrove Market is a regional grocery chain choosing an AI shelf-restocking vendor, Shelfline AI, to predict which items run low before a store associate notices. Esther Adeyemi, Head of Store Operations, is building that scorecard for her own procurement committee. Mapped onto SPARK: situation is a store manager currently reordering by habit and a weekly paper count; payoff is replacing "reorder what usually runs low" with "reorder what this specific store's real sales pattern says will run low this week"; anchor is the same shape as Kestrel's, weighting a live pilot in three real stores above Shelfline's own case-study numbers from other chains' stores.

The keep-out line is where this story breaks differently. Kestrel left minor integrations and unused pricing tiers ungraded. Pinegrove's early scorecard draft instead left seasonal and holiday demand ungraded entirely, reasoning that a four-week pilot couldn't capture a pattern that only shows up once a year anyway. That made sense as a scoping decision. It stopped making sense once the pilot happened to run through a regional harvest festival week, and Shelfline's restocking suggestions, tuned on a chain with no such event, badly underordered produce for three straight days, a gap the scorecard had no line built to catch, because nobody had asked what happens in the one week the model has never seen.

Hand sketched metaphor scene titled A demo is a test drive. A pilot is owning the car, reused here for Pinegrove's produce restocking pilot. Left, TEST DRIVE, one clean loop. Right, OWN IT, potholes, years.
Pinegrove's pothole was a festival week. Nobody had asked the scorecard to watch for it.
Shelfline AI's restock accuracy, ordinary week versus festival week
100% 50% 0 91% Ordinary week 58% Festival week
A 33-point drop, and it happened during the exact week the pilot window covered by chance, not the week the scorecard had been designed to test for.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "weight a real pilot on our own data above the demo, and name what you're deliberately not grading yet," and stop.
Cost: there's no budget for a paid pilot before signing. Say so honestly, and shrink the pilot to the cheapest slice that still touches real, messy data, rather than skipping straight to a demo-only score.
The model gets better, for real: if a vendor's next release genuinely closes the gap a pilot once caught, that's the pilot number doing its job again, telling you it's safe to re-score, not proof the pilot step is no longer needed.

Where people run it wrong.
They let the demo and the pilot share one averaged column instead of scoring them separately.
They try to grade everything, including details that don't decide anything, and bury the two lines that do.
They never name what's deliberately out of scope, so a gap like a festival week reads as a scorecard failure instead of a scoping choice made in the open.

How to use it live. When someone asks you to build an evaluation tool, ask back: what's the one thing a slick pitch could fake, and what's the one thing it can't? Build the scorecard so only the second thing can move the vendor down.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "build an artifact for X" design question like a vendor scorecard?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It designs the one decision that matters before it worries about the whole form.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Nandakumar, who has run Kestrel Systems' developer platform team for four years and built the scorecard for its procurement committee.
3 · THE PAYOFF
What habit should the scorecard build, replacing what old habit?
Tap to flip
ANSWER
Replace "the best pitch wins" with "the best six weeks on our own repository wins."
4 · THE ANCHOR
What's the one design decision the whole scorecard hangs on?
Tap to flip
ANSWER
A live, timed pilot on the buyer's own repository outweighs the vendor's demo score. The demo can lift a score up; only the pilot can pull it back down.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Folding demo quality and pilot results into one shared, averaged column, which let a strong demo quietly cancel out a weak pilot.
6 · THE NUMBER
Fill in the blank: Fenwick AI's weighted scorecard total came out to ___ out of 100, well below the demo-only score of about 92.
Tap to flip
ANSWER
76 out of 100. The pilot line pulled the total down from what the demo alone would have suggested.
7 · THE RISK, HANDLED
What could make the pilot itself unfair, and how does the design protect against it?
Tap to flip
ANSWER
A too-short pilot window could punish a vendor just for being new. The design fixes this with a minimum six-week window before the pilot score locks.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what keep-out line breaks differently there?
Tap to flip
ANSWER
Pinegrove Market's AI restocking vendor, Shelfline AI. The keep-out line that broke was leaving seasonal and holiday demand ungraded, which a festival-week pilot exposed by accident.

Check yourself Score: 0 / 0

Multiple choice
1. Why does the scorecard weight the pilot above the demo?
  • A. Pilots are required by law for enterprise software purchases.
  • B. A demo only shows the vendor's best case on someone else's code; the pilot shows its normal case on yours.
  • C. Pilots are always cheaper than running a demo.
  • D. Vendors expect it and would be offended otherwise.
Show hint
Look at the anchor step.
Show answer
B. The demo is scripted on clean, unfamiliar data. Only a real pilot on your own repository can catch what the demo was never asked to touch.
Short answer, name the anchor
2. What is the one anchor decision this scorecard is built around, stated as a single sentence?
Show hint
Look at the direct answer and the A step.
Show answer
Model answer: A live, timed pilot on the buyer's own repository outweighs the demo score, so only the pilot can pull a vendor's score back down.
True or false
3. True or false: Northlane AI's suggestion-acceptance rate improved steadily across the six-week pilot.
  • True
  • False
Show hint
Look at the story section and the pilot-trend chart.
Show answer
False. Northlane's rate stayed flat and then slipped under forty percent, while Fenwick's climbed toward two-thirds.
Fill in the blank
4. Fill in the blank: Fenwick AI's total weighted scorecard score was ___ out of 100.
Show hint
Look at the stacked-bar chart.
Show answer
76. Well below the roughly 92 the demo alone would have produced.
Short answer, apply it yourself
5. Think of a tool your team adopted based mostly on a demo or a sales call. What would a real, timed pilot on your own data have shown that the demo didn't?
Show hint
Ask what part of your actual daily work the demo never touched.
Show answer
Model answer: Usually the messiest, most real part of the workflow, the edge cases and old data the demo was never shown.
Short answer, where it wouldn't matter
6. Name a part of the vendor evaluation where this pilot-first caution genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Minor IDE integrations and unused pricing tiers. Grading those as deeply as security and pilot performance would bury the two lines that actually decide the vendor.
Before you close the answer
Why this works
Tests whether you'll design an evaluation tool around what a vendor's sales process can fake, or default to a checklist that mostly measures how good their slides were.
Follow-up traps
"Won't a mandatory paid pilot just slow down every vendor evaluation?" Response: it adds weeks, not months, and it's cheaper than a year-long contract with a vendor whose real-world number is quietly getting worse.

"What if a vendor refuses to run an unpaid or discounted pilot at all?" Response: that refusal is itself a data point worth scoring, a vendor confident in its real-world performance usually has the least reason to say no.
If pressed
The pilot window has a floor of six weeks specifically because the first two weeks of any new tool see an adoption dip regardless of vendor quality, engineers getting used to a new suggestion style; scoring on week one or two alone would unfairly punish a good vendor for being new.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more