Describe the difference between evaluating a vendor and evaluating a model.
KinetixAI sells an injury-risk model to sports science departments. Thandeka Moyo is Product Lead for Sports Science Technology at Meridian FC. Petra Nyholm joined the department three weeks before the incident below.
- Split evaluation into two tracks from day one: model quality and vendor operations.Why: one combined score hides which kind of problem you actually have the day something breaks.
- Negotiate a change-notification clause for any model version update.Why: an unannounced update is a vendor-process failure the model's own accuracy number will never show you.
- Test the model on your own hardest cases, not just the vendor's validation study.Why: their number and your reality can diverge exactly where it matters most.
- Set a real bar below which no vendor relationship can compensate for a weak model.Why: at some point the model itself is the actual risk, not the vendor's process around it.
- Track daily flag-rate drift, not just accuracy at signing.Why: it's the number that would have shown the silent update the same morning it happened.
How to answer this, stage by stage
Nobody is scoring you on whether you can define both terms. They're scoring whether you can show a real case where mixing them up actually cost something.
Let's learn
Say a sports science team buys an AI model that reads training load and biometric data and flags which players carry elevated injury risk each morning.
Before it, Meridian FC's staff reviewed GPS load data and wellness questionnaires by hand every morning, about 45 minutes of judgment calls across roughly 30 first-team players, correctly flagging real risk in about 58 percent of the soft-tissue injuries that later occurred. KinetixAI's own validation study reported 89 percent sensitivity. Review time dropped to about 10 minutes a day, spent only on the players it flagged.
Here's the turn: the model's accuracy number was never actually the problem, it stayed close to 89 percent the whole time. The real problem showed up when KinetixAI quietly shipped a new model version mid-season, with no notice to the club, and the daily share of players flagged jumped from about 8 percent to about 34 percent overnight.
At its worst: during that week, two players were pulled from a match warmup on a flagged-but-likely-fine risk score nobody trusted anymore, a visible, costly call made on a number the staff no longer believed, and the department nearly abandoned the tool altogether out of frustration.
What I would leave alone: the model's actual technical accuracy, on the cases it was built for, doesn't need constant second-guessing. Retraining or swapping the model for a better one is a normal, healthy part of the relationship. The problem was never that KinetixAI improved the model. It was that nobody told Meridian it happened.
The lesson: a model evaluation and a vendor evaluation ask two different questions, and a single strong number from one can make you forget to ever ask the other.
Now here is the same thing as a story
The short version above is what you'd say defending this distinction to Meridian's sporting director. Read this one for how the confusion actually played out on the ground.
Her name is Thandeka Moyo. She has run sports science technology procurement for Meridian FC for four years, long enough to have signed the original KinetixAI contract herself, back when the pitch deck's 89 percent sensitivity number felt like the only question worth asking.
For a season and a half, the daily flag list ran quiet and steady. About two or three players a day, out of thirty, came up elevated. The staff trusted the rhythm of it the way you trust a well-worn habit.
Then, on a Wednesday in October, the morning list came back with eleven players flagged instead of the usual two or three. The staff's first read was the obvious one: a hard training block the week before, a run of matches, maybe the squad really was more fragile than usual. They tightened recovery protocols. Two days later, the number was worse, not better.
Petra Nyholm had joined the department three weeks earlier. She hadn't built the habit of trusting the rhythm yet, because she'd never seen the rhythm. In a Thursday meeting, she asked the question nobody else had thought to ask out loud: "Wait, did anything change on KinetixAI's side? Are we sure it's the players, and not the model?"
A call to KinetixAI's support line confirmed it within the hour: a new model version had gone out to all clients two weeks earlier, recalibrated on a larger dataset, technically an improvement, sensitivity had actually ticked up to 91 percent. Nobody at KinetixAI had flagged the change to any client, because their contract had no clause requiring it.
The two extra players pulled from that Saturday's warmup had scored elevated on the new, more sensitive version. Neither had anything actually wrong. The decision wasn't unreasonable given what the staff knew. It just cost a warmup, a lineup change, and a week of the whole department half-trusting a tool that, technically, had gotten better.
Here's what I'd take back. Thandeka's original review scored KinetixAI on one number, the vendor's own validation accuracy, and treated that single figure as the whole due diligence, model and vendor both. That was a reasonable shortcut when the number looked this strong. It stopped being reasonable the moment "the model is good" and "the vendor tells you when it changes" turned out to be two completely different questions with two completely different answers.
I would go back and build a second evaluation track from day one, not the technical one, the operational one: does this vendor notify us before a change, can we pin a version, what's the actual support response time. None of that shows up in a sensitivity score.
And the part I'd tell myself: we didn't get a worse model. We got a better one, silently, and silence was the entire cost.
PICK, in one screenNot "define both terms." PICK is what tells you which one to prioritize when you can't do both deeply, and why the same mistake keeps happening.
The recap, one line per letter: position is prioritizing the vendor review when time is short, impact is a fixable staff annoyance against a costly trust hit, cost asymmetry is invisible-until-expensive beating visible-and-cheap, and kill criteria is a real accuracy floor below which the model itself becomes the urgent problem.
And if you want to be sure it really works, try it somewhere elseSame four letters, a demand-forecasting vendor for a restaurant supply company instead of a sports science tool. A different silent swap, the same missing clause.
Larchmont Foods uses a demand-forecasting vendor to predict weekly order volumes across its regional distribution centers. Emil Kowalczyk runs vendor relationships for supply planning there. Mapped onto PICK: position is prioritizing the vendor-operations review, since a forecasting model's error rate can be corrected with a manual override, but a vendor who changes methods without notice can't be un-trusted overnight. Impact names a planner absorbing a bad week of over-ordering, against the whole company's holiday-peak planning built on a forecast that quietly stopped matching its own historical pattern. Cost asymmetry is the identical shape: a forecast error is visible on a single week's order sheet; a silent methodology change is invisible until the numbers stop making sense for reasons nobody can point to.
The old decision here isn't a single validation number, it's a related reversal: Larchmont's original vendor selection scored candidates only on forecast accuracy, mean absolute percentage error, with no separate checklist item for change-management practices. That made sense when accuracy was the only number every vendor's sales team offered to compare. It stopped making sense the week before a holiday peak, when the vendor updated its underlying model without notice and order quantities swung far outside their normal range, with no warning attached.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "score the model and the vendor separately, and prioritize the vendor review when time is short," and stop.
Cost: there's no budget for a full vendor audit before every renewal. Say so honestly, and start with the one clause that costs nothing to add, change notification.
The model gets better, for real: if a vendor's update genuinely improves accuracy, that's still a vendor-process question first, whether they told you, before it's a model-quality win.
Where people run it wrong.
They treat a strong accuracy number as proof the whole vendor relationship is sound.
They only re-evaluate at contract renewal, missing changes that happen quietly in between.
They blame "the model" for a problem that was actually a vendor's undisclosed process change.
How to use it live. The moment someone asks "how do we evaluate this vendor," ask back: are we scoring the outputs, or the operation producing them? Answer both, on two separate sheets, and never let one hide behind the other.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the vendor won't agree to notify you of updates?" Response: that refusal is itself a vendor-evaluation finding, and it should weigh heavily even if the model's own accuracy looks strong.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Evaluating AI vendors as a buyer
- #1 List the ten questions you would ask every AI vendor before a pilot.
- #2 How do you evaluate a vendor's quality claims without running your own eval?
- #3 Design the pilot you would run to evaluate two competing AI vendors.
- #4 What contractual terms matter specifically for AI vendors and not for other software?
- #5 How do you assess a vendor's model dependency and what happens if their provider changes terms?
- #6 Describe the data handling questions you would put to a vendor on behalf of your security team.