ConceptFoundationalAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #1

List the ten questions you would ask every AI vendor before a pilot.

ORDER the twelve-point checklist that never had a line for a model changing its mind

Ferronova Bearing Works machines precision bearings for pumps and gearboxes. Marek Dubicki has kept the floor's motors running for eleven years. Kelmoor sells a vibration-sensor AI that predicts bearing failure before it happens.

The direct answer
Before anything else, ask two things every AI vendor can answer but almost never volunteers: what happens the day your model changes, and can we walk away if the new version is worse than the one we piloted. Every other question on the list matters. Those two are the ones that turn an ordinary pilot into a trap you can't get out of, because a model is not a fixed product the way old software was, it keeps changing underneath the contract you signed.
Do this, in order
  1. Ask what happens when their model version changes, and whether you'll be told.Why: a silent version change is the single most common way a working pilot turns into a broken one, with no warning and no owner.
  2. Ask for a real exit clause, tested against a worse new version, not just a bad month.Why: without it, the only leverage you have is a favor, not a right.
  3. Ask to see the actual eval set and the pass bar, not a marketing number.Why: a benchmark score with no visible test set is a claim, not evidence.
  4. Ask what happens to your data after the pilot, and whether they retrain on it.Why: this is the one item procurement forgets to renegotiate once the pilot goes well and everyone relaxes.
  5. Ask who is responsible when the model's output causes a real loss.Why: liability left vague at signing gets decided under pressure later, on the vendor's terms.

How to answer this, stage by stage

Nobody is scoring you on whether you can rattle off ten questions fast. They're scoring whether you can say which two matter most, and why an AI vendor needs different ones than a normal software vendor does.

Stage 1
Name the one contract this is about
Say it like this
"I'll answer this for a real case: a manufacturing plant piloting a vendor's vibration-sensor AI for predicting bearing failure."
Why this works
Keeps a ten-item checklist question from turning into a generic procurement lecture.
Stage 2
Say the method out loud
Say it like this
"I'll use ORDER: outcome, reversibility, dependency, evidence, rank. It's built for exactly this, picking what matters most out of a long list."
Why this works
Shows the interviewer you have a repeatable way to rank, not just a list you memorized.
Stage 3
Name what the ten actually protect against
Say it like this
"Every one of these ten questions exists to catch one thing: a model that behaves differently next month than it did during the demo, without anyone telling you."
Why this works
Reframes a checklist as a single real risk instead of ten unrelated boxes to tick.
Stage 4
Rank them, not just list them
Say it like this
"Top of the list: what happens when your model changes, and can we leave if the new one is worse. Those two are hardest to undo once you've signed without them. The rest, pricing, uptime, retraining, matter, but you can renegotiate most of those later. You can't renegotiate a contract that never gave you an exit."
Why this works
This is the direct answer, stated plainly, which is exactly what a "list ten things" question is quietly testing for: can you rank them.
Stage 5
Prove it with the near miss
Say it like this
"Three months into Ferronova's pilot, Kelmoor quietly shipped a model update. False alarms tripled on one line overnight. Nobody had asked what happens when the model changes, so nobody was told it had. A floor supervisor nearly pulled a healthy machine out of production before Marek caught it by checking the readings himself."
Why this works
Turns "ask good vendor questions" from a lecture into a specific afternoon with a name attached to it.
Stage 6
Close on the one line
Say it like this
"Ask all ten if you have time. But ask about version changes and exit rights before you ask about anything else, because those are the two you can't fix after the ink is dry."
Why this works
Restates the direct answer in one breath, ready for a live follow-up question.

Let's learn

Here is what happens when a plant screens an AI vendor with the same checklist it has always used for ordinary software.

Before Kelmoor, Marek and three other technicians walked the floor on a rotation, reading vibration gauges by hand on a clipboard, forty checks a week, about thirteen hours between them. It caught roughly seven of every ten failures early enough to matter. Kelmoor's sensors read every motor continuously and flagged risk scores all day, catching closer to nine of every ten, and freeing Marek's rotation down to a handful of spot checks a week.

Hand sketched flow diagram titled Before the pilot starts. Five boxes in sequence: Data rights, Eval method, Drift plan, Exit terms highlighted, Sign pilot.
This is the order that actually protects a pilot. Most checklists skip straight from the first box to the last.

Here's the turn: catching more failures was never the risk. The risk was that Ferronova's procurement team ran Kelmoor through the same twelve-point questionnaire it used for payroll software and inventory tools, uptime, price tiers, a security certificate, and nothing that asked what happens when the model itself changes. That questionnaire had worked fine for years. It had just never been asked to screen something that could quietly become a different product overnight.

Ferronova's own ten questions, the ones the plant wishes had been asked before signing:

  1. What exact eval set and pass bar did you test this on, and can we see it?
  2. What happens the day your model version changes? Will you tell us?
  3. Can we exit the contract if a new version performs worse than the one we piloted?
  4. Where does our data go after the pilot ends, and can you delete it on request?
  5. Who is responsible if the model's output causes a real loss?
  6. What is the single failure mode you see most often, in plain words?
  7. How is this priced, and does the price change if our usage grows?
  8. What is your uptime and response-time commitment, in writing?
  9. Do you retrain on our data, and can we opt out?
  10. What did the last version change break for another customer, and how did you find out?
Hours lost per quarter to vendor-related pilot surprises, before and after the checklist changed
40 hrs 20 0 34 hrs Old checklist 7 hrs Ten-question checklist
The sensors didn't change. What changed was how many surprises the contract let through before anyone knew to ask.

At its worst, this doesn't just cost a few hours of confusion. It risks a healthy machine getting pulled off a production line on a false alarm nobody had a way to catch, because nobody had asked what a version change would even look like.

The choice I would take back Procurement used the standard twelve-point software questionnaire on Kelmoor, the same one used for every other vendor, because it had always worked before. That made sense when every past vendor sold a fixed product that behaved the same in month twelve as it did in the demo. It stopped making sense the moment Kelmoor shipped a model update nobody had a clause requiring them to announce.

What I would leave alone: the sensor hardware itself doesn't need this level of scrutiny. A vibration sensor is a fixed piece of equipment with a spec sheet; it isn't going to quietly behave differently next quarter the way the model reading its output can.

The lesson: a checklist built for software that never changes will always miss the one risk that matters most in software that does. The fix isn't a longer list. It's two new lines near the top, about the day the thing you bought stops being the thing you tested.

Now here is the same thing as a story

The short version above is what you'd say to a hiring panel. Read this one for how an ordinary Tuesday turned into a near miss on the floor.

Marek Dubicki had spent eleven years learning the sound a bearing makes in the week before it fails, a faint high note under the normal hum, usually two or three days' warning if you knew to listen. When Kelmoor's sensors went live, they caught things Marek's ear sometimes missed, and within two months his rotation shrank from daily rounds to a Friday spot check.

Hand sketched comparison titled Reversible or not. Left, a green box labeled Swap vendor, caption exit clause holds. Right, a red scale icon labeled Locked in, caption no clean exit.
Ferronova had this backwards. Nothing in the contract said which side of this line they were actually on.
Knowledge spark: why would a vendor change a model without telling anyone? Vendors update models to fix bugs or improve accuracy on their whole customer base, not just one plant. A change that helps most customers can shift the balance for a customer whose machines look different from the average, and unless the contract requires a notice, the vendor may not even know your specific case got worse.

Three months in, Kelmoor pushed a routine model update meant to cut false alarms industry-wide. On Ferronova's older gearbox line, the new version read a harmless vibration pattern, one that had always been there, as a fresh warning sign. Alerts on that line tripled inside a single shift.

A floor supervisor, new to the rotation and trained to trust the dashboard, queued the line's newest gearbox motor for an emergency pull. Marek caught it walking past: he recognized the pattern as the same one the line had always had, checked the raw vibration trace himself, and called it off ten minutes before a crew would have shut down a healthy machine.

Nobody had asked Kelmoor what a version change would look like. So when one happened, nobody on the floor had a way to tell a real warning from an old habit the new model had never seen before.

Ferronova's ops VP, reviewing the near miss, pulled ten of the plant's recent vendor contracts for a routine audit the following week and found the same gap in every one that touched AI: no notice clause, no exit tied to version quality, no visible eval set. The checklist wasn't Kelmoor's problem. It was the plant's.

Hand sketched quadrant titled Which questions matter most, axes How easy to verify and Cost if skipped. Exit clause sits high cost easy to verify. Data rights sits high cost harder to verify. Uptime SLA sits low cost easy to verify.
The top right is where the real risk lives. Most standard questionnaires only ever check the bottom right.

Ferronova rewrote its vendor questionnaire the following month, adding a version-notice clause and a right to exit tied to measured performance, not just to nonpayment. Kelmoor kept the contract. It just started having to say something the day it changed.

Hand sketched labeled parts diagram titled What a pilot agreement should hold. A document icon at the center labeled Pilot Deal, with four callouts: data rights, exit clause, eval method, version notice.
This is what got added to the contract after the near miss. None of it existed before.

ORDER, in one screenNot a longer checklist. ORDER is what tells you which two questions on the list you genuinely cannot skip.

O
Outcome. What all ten questions are competing to protect.
Keeping the pilot's results trustworthy after the vendor's model has changed at least once, which it will.
Without naming this, ranking ten questions is just personal taste.
R
Reversibility. Which gap is hardest to undo.
Signing without an exit clause is nearly impossible to fix later; signing without a price-tier detail can usually be renegotiated at renewal.
This is the hardest step, and it's why exit and version-notice rank first.
D
Dependency. What unblocks what.
You can't meaningfully ask for an exit clause tied to performance until you've agreed on the eval set that defines performance in the first place.
Some of the ten have to be asked in a specific order, not just in priority order.
E
Evidence. What you can learn cheaply before signing.
Ask for one real example of a version change that affected another customer. A vendor with a good answer has thought about this before; one with no answer hasn't.
A single sharp question here saves the audit Ferronova had to run after the fact.
R
Rank. State the order and defend the top pick.
Version-notice and exit rights first, eval set and data handling next, liability and pricing after that, uptime and retraining last.
A ranked list is defensible. An alphabetical one just looks organized.
Hand sketched icon list titled Six of the ten, in the room. Show us the eval set and the pass bar. What happens when your model version changes. Who owns it when the output is wrong. Where does our data go after the pilot. What is your worst failure mode, plainly. Can we exit if a new version gets worse.
Six of the ten, the ones that come up first in the actual room with the vendor.

The recap, one line per letter: outcome is keeping trust in the pilot's results after the model inevitably changes, reversibility is the exit clause being the hardest gap to fix later, dependency is needing the eval set agreed before an exit clause means anything, evidence is asking for one real story about a past version change, and rank is putting notice and exit at the top of ten, not buried in the middle.

Hand sketched timeline titled When to ask each one. Four milestones: Before demo, ask model basis. Before contract, ask exit terms, highlighted. During pilot, ask drift signal. Before renewal, ask cost changes.
The two most important questions have exactly one moment where they still work: before signature.

And if you want to be sure it really works, try it somewhere elseSame five letters, a school district's attendance-risk vendor instead of a bearings plant. A pricing decision breaks the second story, not a contract clause.

Halveston County Schools piloted Attendix, a vendor selling AI that flags students at rising risk of chronic absence, so counselors can step in early. Mapped onto ORDER: outcome is trusting which students the model flags as the school year goes on, reversibility is the same top risk, whether the district can exit if a later version starts flagging the wrong students, dependency is needing the district's own definition of "at risk" agreed before any exit clause can be tested against it, evidence is asking Attendix for one real case where a version update changed which students got flagged, and rank puts exit rights and version notice at the top again, the same two questions, a different building entirely.

The old decision here isn't a missing contract clause, it's a pricing one: the district paid Attendix per flagged student, which meant the vendor's own incentive quietly leaned toward flagging more students, not fewer, whenever a new version shipped. That made sense when the district wanted the tool to feel thorough. It stopped making sense once counselors were spending afternoons chasing flags on students who'd never been at risk, while the actual list they needed shrank in the noise.

Hand sketched decision tree titled Which of the ten to ask first, root New AI vendor at the table. Three branches: pilot going to production leads to ask exit and version notice. Vendor is a thin wrapper on someone else's model leads to ask about their upstream provider too. Short trial only leads to ask about data deletion first.
Not every pilot needs all ten asked with equal force on day one. This is what decides where to start.
Share of Attendix's flagged students later found not to be at risk, by quarter after a version update
50% 25 0 per-student pricing starts Q1 Q4 12% 41%
The model got more sensitive right when the pricing gave the vendor a reason to want it more sensitive. Nobody had asked about that link before signing.

Swap the trigger and it still runs.
Speed: an interviewer caps you at forty-five seconds. Say "version notice and exit rights first, everything else after," and stop there.
Cost: there's no time to negotiate all ten before a pilot has to start next week. Say so honestly, and get the two that matter most in writing even if the rest wait for the full contract.
The vendor gets better, for real: if a later version genuinely improves accuracy across the board, that's still a version change nobody may notice without the notice clause, so the question stays worth asking even when the news is good.

Where people run it wrong.
They treat all ten questions as equally important and ask them in whatever order the meeting happens to go.
They accept a vendor's benchmark number without asking what it was actually tested on.
They assume a good pilot result means the contract doesn't need an exit clause, right when they have the most leverage to ask for one.

How to use it live. If an interviewer asks for all ten and you're running short on time, name the two that rank first, say why, and offer the rest as a shorter follow-up list. Ranking under pressure is the actual skill being tested, not recall.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "rank these" or "what matters most" prioritization question?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. It ranks candidates by what's hardest to undo if skipped.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marek Dubicki, a reliability engineer with eleven years on Ferronova's floor, who caught the near miss by checking the raw vibration trace himself.
3 · THE HABIT
What did Marek's team stop doing because Kelmoor worked?
Tap to flip
ANSWER
Daily manual vibration rounds on the clipboard rotation, cut down to a single Friday spot check within two months.
4 · THE TOP TWO
Which two of the ten questions rank above the rest, and why?
Tap to flip
ANSWER
What happens when the model version changes, and whether you can exit if the new version is worse. Both are nearly impossible to secure after signing.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Screening Kelmoor with the same twelve-point questionnaire used for ordinary software, with no line for a model version changing underneath the contract.
6 · THE NUMBER
Fill in the blank: hours lost per quarter to vendor-related surprises fell from 34 to ___ after the ten-question checklist replaced the old one.
Tap to flip
ANSWER
7 hours, roughly a fifth of the old rate, once version-notice and exit clauses were in writing.
7 · THE REPLAY
Same version update, same gearbox line, but the notice clause now exists. What changes?
Tap to flip
ANSWER
Kelmoor flags the update before it ships. Marek's team re-checks the baseline on that one line first, so the false-alarm spike never reaches the floor supervisor at all.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what old decision gets taken back?
Tap to flip
ANSWER
Halveston County Schools' Attendix pilot. The reversal is per-student pricing, which quietly rewarded the vendor for flagging more students as the model changed.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: hours lost per quarter to vendor surprises at Ferronova fell from 34 hours to ___ hours after the checklist changed.
Show hint
Look at the bar chart in Section 1.
Show answer
7 hours. The sensors didn't improve, the contract just stopped hiding version changes.
Multiple choice
2. According to the reversibility step, why do version-notice and exit rights rank above pricing and uptime questions?
  • A. They're required by law in every industry.
  • B. Vendors expect them and get offended if you skip them.
  • C. They're nearly impossible to secure after the contract is already signed.
  • D. They cost the vendor the most money to provide.
Show hint
Look at the R step in the ORDER recap.
Show answer
C. Pricing and uptime can usually be renegotiated later. A missing exit clause can't be added after the fact.
True or false
3. True or false: the near miss on Ferronova's floor happened because Kelmoor's sensors were faulty.
  • True
  • False
Show hint
Look at the knowledge spark about model updates.
Show answer
False. The sensors worked. A routine model update changed how alerts were scored, and nobody had asked to be told when that would happen.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Using the standard software questionnaire on an AI vendor. It made sense because every past vendor sold a fixed product that never changed after the demo.
Short answer, where it wouldn't matter
5. Name a part of Ferronova's setup where this extra scrutiny genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The vibration sensor hardware itself. It's a fixed piece of equipment with a spec sheet, not something that quietly changes behavior next quarter.
Short answer, apply it yourself
6. Think of a vendor tool you rely on at work or at home. What's the one question about it you've never actually asked, but probably should?
Show hint
Ask yourself what would happen if the tool changed behavior overnight, and whether you'd even find out.
Show answer
Model answer: Most people have never asked "will you tell me when this changes," because it never mattered with ordinary software. It matters with anything built on a model.
Before you close the answer
Why this works
Tests whether you can rank a list under pressure instead of just reciting it, and whether you understand why an AI vendor needs different questions than an ordinary software one.
Follow-up traps
"Isn't asking for all ten just going to slow down every pilot?" Response: no, most of the ten can be a quick email exchange; only the top two are worth holding up a signature over.

"What if the vendor refuses to commit to a version-notice clause?" Response: that refusal is itself the answer, a vendor unwilling to say when it changes its own product is telling you exactly how much control you'll have later.
If pressed
Ferronova's rewritten contract defined "worse" for the exit clause using the same eval metric from question one, a rise in false-alarm rate above a stated percentage on their own machines, not Kelmoor's aggregate customer base.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more