ConceptIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #17

Describe the difference between evaluating a vendor and evaluating a model.

PICK the accuracy number that never moved, and the update nobody announced

KinetixAI sells an injury-risk model to sports science departments. Thandeka Moyo is Product Lead for Sports Science Technology at Meridian FC. Petra Nyholm joined the department three weeks before the incident below.

The direct answer
Evaluating a model asks whether its outputs are good enough on your own hardest cases. Evaluating a vendor asks whether everything around those outputs, its change process, its communication, its accountability, will hold up over years. Score both, on two separate sheets. A great model score tells you nothing about whether the vendor will silently change that model out from under you.
Do this, in order
  1. Split evaluation into two tracks from day one: model quality and vendor operations.Why: one combined score hides which kind of problem you actually have the day something breaks.
  2. Negotiate a change-notification clause for any model version update.Why: an unannounced update is a vendor-process failure the model's own accuracy number will never show you.
  3. Test the model on your own hardest cases, not just the vendor's validation study.Why: their number and your reality can diverge exactly where it matters most.
  4. Set a real bar below which no vendor relationship can compensate for a weak model.Why: at some point the model itself is the actual risk, not the vendor's process around it.
  5. Track daily flag-rate drift, not just accuracy at signing.Why: it's the number that would have shown the silent update the same morning it happened.

How to answer this, stage by stage

Nobody is scoring you on whether you can define both terms. They're scoring whether you can show a real case where mixing them up actually cost something.

Stage 1
Scope it to one vendor, one distinction
Say it like this
"I'll answer this for Meridian FC's contract with KinetixAI, an injury-risk model for first-team players."
Why this works
Keeps an abstract definitional question anchored to one real decision.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my actual pick. Impact, who feels each kind of gap. Cost asymmetry, which one is worse. Kill criteria, what would flip my priority."
Why this works
Shows you're treating this as a real tradeoff to commit to, not just two terms to define.
Stage 3
Give the position, before any reasoning
Say it like this
"If time only allows doing one deeply before signing, do the vendor evaluation. You can retrain or replace a mediocre model without switching partners. You cannot easily fix a bad vendor relationship no matter how good the model scores."
Why this works
This is the direct answer's core claim, stated as a real, defensible commitment.
Stage 4
Name who feels each kind of gap
Say it like this
"A model gap lands on the sports science staff sorting through extra false alarms, annoying but fixable by Friday. A vendor gap lands on the whole department's trust in the tool, and it doesn't get fixed by Friday."
Why this works
Names both sides in felt terms instead of an abstract compliance checklist.
Stage 5
Prove it with the silent update
Say it like this
"KinetixAI's own accuracy number never moved. But they pushed a new model version mid-season with no notice, and our daily flag rate jumped from 8 percent of players to 34 percent overnight. We spent a week assuming our players had gotten fragile before a new hire asked if the model itself had changed."
Why this works
Turns "vendor eval matters too" from a compliance point into a specific, costly week.
Stage 6
Close on the one line
Say it like this
"A model tells you if the outputs are good today. A vendor tells you if you'll still trust those outputs next year. Score them separately, and never let one hide a problem in the other."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.

Let's learn

Say a sports science team buys an AI model that reads training load and biometric data and flags which players carry elevated injury risk each morning.

Before it, Meridian FC's staff reviewed GPS load data and wellness questionnaires by hand every morning, about 45 minutes of judgment calls across roughly 30 first-team players, correctly flagging real risk in about 58 percent of the soft-tissue injuries that later occurred. KinetixAI's own validation study reported 89 percent sensitivity. Review time dropped to about 10 minutes a day, spent only on the players it flagged.

Hand sketched labeled parts diagram titled Two evaluation tracks, not one. A document icon at the center labeled KinetixAI Eval, with four callouts: model accuracy, vendor process, change control, support response.
Meridian's contract review only ever filled in one of these four boxes.

Here's the turn: the model's accuracy number was never actually the problem, it stayed close to 89 percent the whole time. The real problem showed up when KinetixAI quietly shipped a new model version mid-season, with no notice to the club, and the daily share of players flagged jumped from about 8 percent to about 34 percent overnight.

Daily share of players flagged, three weeks around the silent update
40% 20 0 silent v2 update Week 2: 8% Week 3: 34%
The model's own accuracy score never flagged this. Only the daily flag-rate number, tracked separately, would have caught it the same morning.

At its worst: during that week, two players were pulled from a match warmup on a flagged-but-likely-fine risk score nobody trusted anymore, a visible, costly call made on a number the staff no longer believed, and the department nearly abandoned the tool altogether out of frustration.

The choice I would take back Meridian's original vendor review treated KinetixAI's published validation accuracy, 89 percent sensitivity, as the entire evaluation, technical and operational both. That made sense at the time: the number was strong, and a strong number felt like it covered the whole question. It stopped covering the whole question the moment the vendor changed something the number could never show.
Knowledge spark: what is model versioning? It's a vendor swapping in a newer version of its underlying model, sometimes retrained on new data, sometimes recalibrated differently. Done quietly, it looks to you like nothing changed. It looks to the model like everything did.

What I would leave alone: the model's actual technical accuracy, on the cases it was built for, doesn't need constant second-guessing. Retraining or swapping the model for a better one is a normal, healthy part of the relationship. The problem was never that KinetixAI improved the model. It was that nobody told Meridian it happened.

The lesson: a model evaluation and a vendor evaluation ask two different questions, and a single strong number from one can make you forget to ever ask the other.

Now here is the same thing as a story

The short version above is what you'd say defending this distinction to Meridian's sporting director. Read this one for how the confusion actually played out on the ground.

Her name is Thandeka Moyo. She has run sports science technology procurement for Meridian FC for four years, long enough to have signed the original KinetixAI contract herself, back when the pitch deck's 89 percent sensitivity number felt like the only question worth asking.

For a season and a half, the daily flag list ran quiet and steady. About two or three players a day, out of thirty, came up elevated. The staff trusted the rhythm of it the way you trust a well-worn habit.

Hand sketched flow diagram titled How the silent update actually happened. Five boxes in sequence: KinetixAI ships v2, No notice sent highlighted, Flag rate jumps, Staff assume players worse, Petra asks the question.
The whole incident lives in the second box. Everything after it is just people reacting to a change nobody told them about.

Then, on a Wednesday in October, the morning list came back with eleven players flagged instead of the usual two or three. The staff's first read was the obvious one: a hard training block the week before, a run of matches, maybe the squad really was more fragile than usual. They tightened recovery protocols. Two days later, the number was worse, not better.

Petra Nyholm had joined the department three weeks earlier. She hadn't built the habit of trusting the rhythm yet, because she'd never seen the rhythm. In a Thursday meeting, she asked the question nobody else had thought to ask out loud: "Wait, did anything change on KinetixAI's side? Are we sure it's the players, and not the model?"

Nobody had an answer, and that was the actual finding. Not that the model had changed. That nobody at Meridian would have known either way.

A call to KinetixAI's support line confirmed it within the hour: a new model version had gone out to all clients two weeks earlier, recalibrated on a larger dataset, technically an improvement, sensitivity had actually ticked up to 91 percent. Nobody at KinetixAI had flagged the change to any client, because their contract had no clause requiring it.

Hand sketched comparison titled The asymmetry, drawn. Left, a small gauge icon labeled Model gap, caption false positives, fixable fast. Right, a larger box icon labeled Vendor gap, caption silent change, costly trust hit.
The model got slightly better. The vendor relationship got a lot worse, for a week nobody could explain.

The two extra players pulled from that Saturday's warmup had scored elevated on the new, more sensitive version. Neither had anything actually wrong. The decision wasn't unreasonable given what the staff knew. It just cost a warmup, a lineup change, and a week of the whole department half-trusting a tool that, technically, had gotten better.

Hand sketched quadrant titled Where problems actually come from. Axes, kind of problem from technical to operational, and how easy to spot from invisible to visible. False positive rate sits technical and visible. Silent model swap sits operational and invisible. Support ticket delay sits operational and visible. Stale training data sits technical and invisible.
The incident that actually hurt Meridian sits in the one quadrant a pure accuracy number will never reach.

Here's what I'd take back. Thandeka's original review scored KinetixAI on one number, the vendor's own validation accuracy, and treated that single figure as the whole due diligence, model and vendor both. That was a reasonable shortcut when the number looked this strong. It stopped being reasonable the moment "the model is good" and "the vendor tells you when it changes" turned out to be two completely different questions with two completely different answers.

I would go back and build a second evaluation track from day one, not the technical one, the operational one: does this vendor notify us before a change, can we pin a version, what's the actual support response time. None of that shows up in a sensitivity score.

And the part I'd tell myself: we didn't get a worse model. We got a better one, silently, and silence was the entire cost.

PICK, in one screenNot "define both terms." PICK is what tells you which one to prioritize when you can't do both deeply, and why the same mistake keeps happening.

P
Position. The pick, before any reasoning.
If time only allows one deep review before signing, do the vendor evaluation. A mediocre model can be retrained or replaced without switching partners. A bad vendor relationship can't be fixed by a better model.
Interviewers are testing whether you can commit to a real answer instead of saying "both matter equally."
I
Impact. Who feels each kind of gap.
A model gap lands on staff sorting extra false alarms, annoying but fixable by Friday. A vendor gap lands on the whole department's trust in the tool, and that doesn't come back by Friday.
Naming both in felt terms is what keeps this from being an abstract definitions question.
C
Cost asymmetry. Which one is actually worse.
A model gap is visible and cheap: retrain, adjust a threshold, add a human check. A vendor-process gap is invisible until it isn't, and by then it's cost a lineup decision and a week of half-trust.
The hardest step, and the whole pick turns on it.
K
Kill criteria. What would flip the pick.
If the model's own accuracy fell below a real safety bar, missing genuinely serious injury risks, model quality becomes the urgent priority regardless of how strong the vendor relationship is.
Naming this in advance is what separates a real priority from a rule of thumb.
Hand sketched icon list titled What a vendor eval checks that a model eval doesn't. Four items: change notification for updates, version pinning rights, a real support response SLA, data handling policy over time.
None of these four appear on any model's accuracy report. All four are exactly where this incident actually lived.

The recap, one line per letter: position is prioritizing the vendor review when time is short, impact is a fixable staff annoyance against a costly trust hit, cost asymmetry is invisible-until-expensive beating visible-and-cheap, and kill criteria is a real accuracy floor below which the model itself becomes the urgent problem.

And if you want to be sure it really works, try it somewhere elseSame four letters, a demand-forecasting vendor for a restaurant supply company instead of a sports science tool. A different silent swap, the same missing clause.

Larchmont Foods uses a demand-forecasting vendor to predict weekly order volumes across its regional distribution centers. Emil Kowalczyk runs vendor relationships for supply planning there. Mapped onto PICK: position is prioritizing the vendor-operations review, since a forecasting model's error rate can be corrected with a manual override, but a vendor who changes methods without notice can't be un-trusted overnight. Impact names a planner absorbing a bad week of over-ordering, against the whole company's holiday-peak planning built on a forecast that quietly stopped matching its own historical pattern. Cost asymmetry is the identical shape: a forecast error is visible on a single week's order sheet; a silent methodology change is invisible until the numbers stop making sense for reasons nobody can point to.

The old decision here isn't a single validation number, it's a related reversal: Larchmont's original vendor selection scored candidates only on forecast accuracy, mean absolute percentage error, with no separate checklist item for change-management practices. That made sense when accuracy was the only number every vendor's sales team offered to compare. It stopped making sense the week before a holiday peak, when the vendor updated its underlying model without notice and order quantities swung far outside their normal range, with no warning attached.

Cost of each kind of gap, in staff hours spent recovering from it
240 hrs 120 0 40 hrs Model-quality gap 220 hrs Vendor-process gap
The model-quality gap is a maintenance task. The vendor-process gap is a five-times-larger fire drill, from one missing clause.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "score the model and the vendor separately, and prioritize the vendor review when time is short," and stop.
Cost: there's no budget for a full vendor audit before every renewal. Say so honestly, and start with the one clause that costs nothing to add, change notification.
The model gets better, for real: if a vendor's update genuinely improves accuracy, that's still a vendor-process question first, whether they told you, before it's a model-quality win.

Where people run it wrong.
They treat a strong accuracy number as proof the whole vendor relationship is sound.
They only re-evaluate at contract renewal, missing changes that happen quietly in between.
They blame "the model" for a problem that was actually a vendor's undisclosed process change.

How to use it live. The moment someone asks "how do we evaluate this vendor," ask back: are we scoring the outputs, or the operation producing them? Answer both, on two separate sheets, and never let one hide behind the other.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "A or B" tradeoff and priority questions, including a two-concept comparison?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It forces a real committed answer instead of "both matter equally."
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Thandeka Moyo, Product Lead for Sports Science Technology at Meridian FC, four years running the KinetixAI relationship.
3 · THE HABIT
What did Meridian's staff stop doing once the daily flag list ran quiet and steady?
Tap to flip
ANSWER
Questioning whether the model itself had changed. A season and a half of steady numbers built a trust nobody thought to re-check.
4 · THE COST ASYMMETRY
Which kind of gap is actually worse, and why?
Tap to flip
ANSWER
The vendor-process gap. It's invisible until it isn't, and it costs a week of trust and a real, visible lineup decision, not just extra review time.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating KinetixAI's published validation accuracy as the entire evaluation, technical and operational both, instead of building a separate vendor-process track.
6 · THE NUMBER
Fill in the blank: the daily flag rate jumped from about 8 percent to about ___ percent overnight, with no change to the model's own reported accuracy.
Tap to flip
ANSWER
34 percent. The model's accuracy actually rose slightly, to 91 percent, but nobody was told the version had changed.
7 · THE REPLAY
Same silent update, but a change-notification clause is already in the contract. What changes?
Tap to flip
ANSWER
KinetixAI has to tell Meridian before the update ships. Staff know to expect a shift in flag rate, and no player gets pulled from a warmup on a number nobody trusts.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, using the same framework. Which product, and what's the reversal?
Tap to flip
ANSWER
Larchmont Foods' demand-forecasting vendor. The reversal is scoring vendor candidates only on forecast accuracy, with no change-management checklist item.

Check yourself Score: 0 / 0

True or false
1. True or false: the flag-rate spike at Meridian FC happened because KinetixAI's model got worse.
  • True
  • False
Show hint
Look at what happened to the model's actual accuracy after the update.
Show answer
False. The model's sensitivity actually rose to 91 percent. The problem was that nobody was told the version had changed.
Fill in the blank
2. Fill in the blank: the daily share of players flagged jumped from about 8 percent to about ___ percent overnight.
Show hint
Look at the line chart in Section 1.
Show answer
34 percent. A more than four-times jump, with no change to the model's own reported accuracy score.
Multiple choice
3. According to the Position step, which review should get priority if time is short before signing?
  • A. The vendor's marketing materials.
  • B. The vendor's operational practices, like change notification.
  • C. The model's exact algorithm.
  • D. The price per player, per month.
Show hint
Look at the Position step in the PICK recap.
Show answer
B. A mediocre model can be retrained or replaced. A bad vendor relationship can't be fixed by a better model.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "The choice I would take back."
Show answer
Model answer: Treating KinetixAI's validation accuracy as the whole evaluation. It made sense because the number was strong, and a strong number felt like it covered everything.
Short answer, where it wouldn't matter
5. Name a case where a silent model version update genuinely wouldn't be a problem.
Show hint
Look at "What I would leave alone."
Show answer
Model answer: The model's own technical accuracy improving over time. Retraining or upgrading the model is healthy; the problem was only ever the silence around it.
Short answer, apply it yourself
6. Think of a vendor or tool you rely on. Have you ever separately checked its actual output quality versus its process for telling you about changes?
Show hint
Ask whether you'd know if the tool changed behind the scenes, separate from whether you like its current output.
Show answer
Model answer: Most people only ever check the output quality. The vendor-process question, would they tell you if something changed, is the one usually skipped.
Before you close the answer
Why this works
Tests whether you'll treat "the model is accurate" and "the vendor is trustworthy" as the same claim, and whether you can name a real, costly moment where confusing them actually mattered.
Follow-up traps
"Isn't a change-notification clause just extra paperwork if the model keeps improving?" Response: no, because the cost isn't the change itself, it's the week spent assuming the players got worse instead of knowing the model did.

"What if the vendor won't agree to notify you of updates?" Response: that refusal is itself a vendor-evaluation finding, and it should weigh heavily even if the model's own accuracy looks strong.
If pressed
Meridian's renegotiated contract added a 14-day advance-notice window for any model version change, plus the right to stay on the prior version for up to 30 days while staff validated the new one on their own data.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more