List the ten questions you would ask every AI vendor before a pilot.
Ferronova Bearing Works machines precision bearings for pumps and gearboxes. Marek Dubicki has kept the floor's motors running for eleven years. Kelmoor sells a vibration-sensor AI that predicts bearing failure before it happens.
- Ask what happens when their model version changes, and whether you'll be told.Why: a silent version change is the single most common way a working pilot turns into a broken one, with no warning and no owner.
- Ask for a real exit clause, tested against a worse new version, not just a bad month.Why: without it, the only leverage you have is a favor, not a right.
- Ask to see the actual eval set and the pass bar, not a marketing number.Why: a benchmark score with no visible test set is a claim, not evidence.
- Ask what happens to your data after the pilot, and whether they retrain on it.Why: this is the one item procurement forgets to renegotiate once the pilot goes well and everyone relaxes.
- Ask who is responsible when the model's output causes a real loss.Why: liability left vague at signing gets decided under pressure later, on the vendor's terms.
How to answer this, stage by stage
Nobody is scoring you on whether you can rattle off ten questions fast. They're scoring whether you can say which two matter most, and why an AI vendor needs different ones than a normal software vendor does.
Let's learn
Here is what happens when a plant screens an AI vendor with the same checklist it has always used for ordinary software.
Before Kelmoor, Marek and three other technicians walked the floor on a rotation, reading vibration gauges by hand on a clipboard, forty checks a week, about thirteen hours between them. It caught roughly seven of every ten failures early enough to matter. Kelmoor's sensors read every motor continuously and flagged risk scores all day, catching closer to nine of every ten, and freeing Marek's rotation down to a handful of spot checks a week.
Here's the turn: catching more failures was never the risk. The risk was that Ferronova's procurement team ran Kelmoor through the same twelve-point questionnaire it used for payroll software and inventory tools, uptime, price tiers, a security certificate, and nothing that asked what happens when the model itself changes. That questionnaire had worked fine for years. It had just never been asked to screen something that could quietly become a different product overnight.
Ferronova's own ten questions, the ones the plant wishes had been asked before signing:
- What exact eval set and pass bar did you test this on, and can we see it?
- What happens the day your model version changes? Will you tell us?
- Can we exit the contract if a new version performs worse than the one we piloted?
- Where does our data go after the pilot ends, and can you delete it on request?
- Who is responsible if the model's output causes a real loss?
- What is the single failure mode you see most often, in plain words?
- How is this priced, and does the price change if our usage grows?
- What is your uptime and response-time commitment, in writing?
- Do you retrain on our data, and can we opt out?
- What did the last version change break for another customer, and how did you find out?
At its worst, this doesn't just cost a few hours of confusion. It risks a healthy machine getting pulled off a production line on a false alarm nobody had a way to catch, because nobody had asked what a version change would even look like.
What I would leave alone: the sensor hardware itself doesn't need this level of scrutiny. A vibration sensor is a fixed piece of equipment with a spec sheet; it isn't going to quietly behave differently next quarter the way the model reading its output can.
The lesson: a checklist built for software that never changes will always miss the one risk that matters most in software that does. The fix isn't a longer list. It's two new lines near the top, about the day the thing you bought stops being the thing you tested.
Now here is the same thing as a story
The short version above is what you'd say to a hiring panel. Read this one for how an ordinary Tuesday turned into a near miss on the floor.
Marek Dubicki had spent eleven years learning the sound a bearing makes in the week before it fails, a faint high note under the normal hum, usually two or three days' warning if you knew to listen. When Kelmoor's sensors went live, they caught things Marek's ear sometimes missed, and within two months his rotation shrank from daily rounds to a Friday spot check.
Three months in, Kelmoor pushed a routine model update meant to cut false alarms industry-wide. On Ferronova's older gearbox line, the new version read a harmless vibration pattern, one that had always been there, as a fresh warning sign. Alerts on that line tripled inside a single shift.
A floor supervisor, new to the rotation and trained to trust the dashboard, queued the line's newest gearbox motor for an emergency pull. Marek caught it walking past: he recognized the pattern as the same one the line had always had, checked the raw vibration trace himself, and called it off ten minutes before a crew would have shut down a healthy machine.
Ferronova's ops VP, reviewing the near miss, pulled ten of the plant's recent vendor contracts for a routine audit the following week and found the same gap in every one that touched AI: no notice clause, no exit tied to version quality, no visible eval set. The checklist wasn't Kelmoor's problem. It was the plant's.
Ferronova rewrote its vendor questionnaire the following month, adding a version-notice clause and a right to exit tied to measured performance, not just to nonpayment. Kelmoor kept the contract. It just started having to say something the day it changed.
ORDER, in one screenNot a longer checklist. ORDER is what tells you which two questions on the list you genuinely cannot skip.
The recap, one line per letter: outcome is keeping trust in the pilot's results after the model inevitably changes, reversibility is the exit clause being the hardest gap to fix later, dependency is needing the eval set agreed before an exit clause means anything, evidence is asking for one real story about a past version change, and rank is putting notice and exit at the top of ten, not buried in the middle.
And if you want to be sure it really works, try it somewhere elseSame five letters, a school district's attendance-risk vendor instead of a bearings plant. A pricing decision breaks the second story, not a contract clause.
Halveston County Schools piloted Attendix, a vendor selling AI that flags students at rising risk of chronic absence, so counselors can step in early. Mapped onto ORDER: outcome is trusting which students the model flags as the school year goes on, reversibility is the same top risk, whether the district can exit if a later version starts flagging the wrong students, dependency is needing the district's own definition of "at risk" agreed before any exit clause can be tested against it, evidence is asking Attendix for one real case where a version update changed which students got flagged, and rank puts exit rights and version notice at the top again, the same two questions, a different building entirely.
The old decision here isn't a missing contract clause, it's a pricing one: the district paid Attendix per flagged student, which meant the vendor's own incentive quietly leaned toward flagging more students, not fewer, whenever a new version shipped. That made sense when the district wanted the tool to feel thorough. It stopped making sense once counselors were spending afternoons chasing flags on students who'd never been at risk, while the actual list they needed shrank in the noise.
Swap the trigger and it still runs.
Speed: an interviewer caps you at forty-five seconds. Say "version notice and exit rights first, everything else after," and stop there.
Cost: there's no time to negotiate all ten before a pilot has to start next week. Say so honestly, and get the two that matter most in writing even if the rest wait for the full contract.
The vendor gets better, for real: if a later version genuinely improves accuracy across the board, that's still a version change nobody may notice without the notice clause, so the question stays worth asking even when the news is good.
Where people run it wrong.
They treat all ten questions as equally important and ask them in whatever order the meeting happens to go.
They accept a vendor's benchmark number without asking what it was actually tested on.
They assume a good pilot result means the contract doesn't need an exit clause, right when they have the most leverage to ask for one.
How to use it live. If an interviewer asks for all ten and you're running short on time, name the two that rank first, say why, and offer the rest as a shorter follow-up list. Ranking under pressure is the actual skill being tested, not recall.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the vendor refuses to commit to a version-notice clause?" Response: that refusal is itself the answer, a vendor unwilling to say when it changes its own product is telling you exactly how much control you'll have later.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Evaluating AI vendors as a buyer
- #2 How do you evaluate a vendor's quality claims without running your own eval?
- #3 Design the pilot you would run to evaluate two competing AI vendors.
- #4 What contractual terms matter specifically for AI vendors and not for other software?
- #5 How do you assess a vendor's model dependency and what happens if their provider changes terms?
- #6 Describe the data handling questions you would put to a vendor on behalf of your security team.
- #7 What does a good vendor eval report look like and what should make you suspicious?