ConceptIntermediateAI Opportunity & Model Strategy / Evaluating AI vendors as a buyer / #21
How do you evaluate the support and incident response terms of an AI vendor?
LEAD the contract number that fails quietly, weeks before the outage that makes it matter
Elkhorn Veterinary Group runs six emergency animal hospitals. Naledi Okonjo is VP of Clinical Operations. Vetrix is the AI vendor whose model flags likely internal bleeding and fractures on ultrasound and x-ray images, before a radiologist reviews the case.
The direct answer
Evaluate the terms by their leading number, not their promised one: measured time to a live human acknowledgment on a real past incident, not the response-time figure printed in the SLA. A vendor's contract can promise fifteen minutes and still leave you waiting ninety, and the only way to catch that gap before an emergency exposes it is to ask for their own incident logs, not their sales sheet.
Do this, in order
Ask for real acknowledgment-time logs from past incidents, not the SLA's promised number.Why: a promised response time and a measured one can differ by an hour or more, and only the logs show which one is real.
Pin down what counts as "acknowledged," in writing.Why: an automated ticket-received email and a human starting triage are not the same event, and vague wording lets a vendor claim the cheaper one.
Require a named escalation path with a backup contact, not one email address.Why: a single point of contact fails exactly when you need it most, during a major outage when that one person is unreachable.
Set the service credit to scale with real harm, not a flat percentage of monthly fees.Why: a capped credit that's cheaper than the actual cost of an outage gives the vendor no real reason to fix the underlying problem.
Track acknowledgment time monthly after signing, not just at renewal.Why: a slow drift upward is the early warning; waiting for the next big outage to notice it is waiting too long.
How to answer this, stage by stage
Picture the interviewer isn't grading whether you know what "SLA" stands for. They're grading whether you'd trust a printed number or go find the real one.
Stage 1
Scope it to one real vendor relationship
Say it like this
"I'll answer this for Elkhorn Veterinary Group, evaluating Vetrix's incident response terms for its AI diagnostic imaging tool, across six emergency animal hospitals."
Why this works
Keeps "evaluate the terms" from turning into a generic legal checklist.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the real business outcome at stake. Early signal, the number that moves first. Abuse, how a vendor could game that number. Decision, what I'd actually do at each threshold."
Why this works
Shows you're not just going to read the contract line by line, you have a method for what to check.
Stage 3
Name the real outcome at stake
Say it like this
"The outcome isn't 'the software is back up.' It's how long a case sits without a working diagnostic read during an actual emergency, because that's the minutes that matter to the animal in front of us."
Why this works
Grounds the whole evaluation in the thing that actually costs something, not a vague notion of "uptime."
Stage 4
Give the one decision, before any reasoning
Say it like this
"Ask for the vendor's actual acknowledgment-time logs from past incidents, not the number printed in the contract. That measured number is the real early signal, weeks before it ever shows up as a bad night in the clinic."
Why this works
This is the direct answer, said plainly, before any story does the convincing.
Stage 5
Prove it with the compressed near miss
Say it like this
"During one real outage, Vetrix's contract promised acknowledgment within fifteen minutes. The actual first human response came at over thirteen hours. A tech caught the case by hand before anything went wrong, but the printed number and the real one were nowhere near each other."
Why this works
Turns "check the SLA" into one specific, checkable gap instead of a vague warning.
Stage 6
Say the threshold and the action
Say it like this
"If measured acknowledgment time crosses thirty minutes for two months running, that's not a fluke, that's a trend, and it triggers a mandatory renegotiation clause, not just a strongly worded email."
Why this works
Shows the metric isn't a dashboard decoration, it has a real, pre-agreed action attached.
Stage 7
Close on the one line
Say it like this
"Don't evaluate the number a vendor promises you. Evaluate the number their own past incidents actually produced, and watch it every month, not just at renewal."
Why this works
Restates the direct answer in one breath, ready for a live follow-up.
Let's learn
Here is what a real incident looks like when a diagnostic AI vendor's support terms only look good on paper.
Before Vetrix, a night-shift vet tech at Elkhorn would flag a suspicious ultrasound to an on-call radiologist by phone, getting a read back in about twenty-five minutes on average. Vetrix's model, once installed, flagged the same kind of case in under ninety seconds, and cut the average wait for a first read to about four minutes.
The whole contract lives or dies on the second box. Everything after it depends on how long that one takes.
Here's the turn: the SLA Elkhorn signed promised a fifteen-minute acknowledgment on any outage. Nobody had ever measured whether Vetrix's own support desk actually hit that number, because it had never mattered enough to check, until the night it did.
Average time to first human acknowledgment on Vetrix support tickets, by month
The line crossed the 30-minute concern mark in month four. The major outage didn't happen until month seven, three months of warning nobody was watching for.
Four separate promises. Elkhorn had only ever read the first one closely.
At its worst, an unmeasured support gap doesn't just cost a few hours of frustration. It costs the exact minutes a genuinely sick animal doesn't have, with nobody at the vendor answering the phone.
The decision I would take back
Elkhorn accepted a fifteen-minute acknowledgment promise in the contract without ever asking Vetrix for logs showing what their support desk actually delivered on past incidents. That made sense at signing, when the relationship was new and no incident history existed yet. It stopped making sense a year in, once real incident data existed and nobody had gone back to check it against the number on paper.
What I would leave alone: Vetrix's actual diagnostic model doesn't need this same scrutiny on uptime philosophy. A model that's briefly unavailable during a scheduled update is a routine maintenance question, not an incident-response one, and treating every short planned outage like a crisis would burn trust for no reason.
The lesson: a number in a contract is a claim, not a fact, until you've asked to see what actually happened the last time it was tested for real.
Now here is the same thing as a story
The short version above is what you'd say defending this evaluation method to Elkhorn's board. Read this one for how close the near miss actually came.
Naledi Okonjo had run clinical operations at Elkhorn for nine years and knew every one of her six hospitals' overnight rhythms by heart. The Tuesday the imaging system went dark started like any other overnight shift, quiet until it wasn't.
At 8:02 a.m., a Labrador named Milo came in showing signs of internal bleeding after being hit by a car. The vet tech pulled up an ultrasound and reached for Vetrix's flagging tool. Nothing loaded. The system had gone down eleven minutes earlier, mid-scan, for reasons nobody at the clinic could see.
Knowledge spark: why does an AI diagnostic tool "go down" at all?
Most of these tools run the actual model on the vendor's own servers, not inside the clinic's building. If that server, or the connection to it, has a problem, every clinic using the tool loses the flag at the same time, even though nothing changed on the clinic's own equipment. It looks like a local glitch. It's actually a shared outage.
At 8:14, the tech called Elkhorn's emergency contact for Vetrix, the number printed at the top of the support agreement. The contract promised acknowledgment within fifteen minutes. By 9:40, over ninety minutes later, there had been no reply at all, not even an automated one.
Thirteen hours between the call and the reply. Fifteen minutes was the number on paper.
Milo's case never went unread. A senior vet, not the AI, caught the bleeding by hand from the raw scan. The contract's fifteen-minute promise, the thing Elkhorn had actually paid for, never showed up at all.
The senior vet on duty that morning, with nineteen years of practice behind her, read the raw ultrasound herself and caught the bleed without the AI flag, the same way clinics worked before Vetrix existed. Milo recovered. But the first real human response from Vetrix's support line didn't arrive until 9:36 that night, thirteen hours and thirty-four minutes after the call went in.
Naledi's team pulled Vetrix's own incident log after the fact and found the gap wasn't a one-time fluke. Acknowledgment time had been quietly climbing for three months before that Tuesday, and nobody at Elkhorn had been watching the number closely enough to notice it crossing thirty minutes, then forty-five, then past an hour.
Same clinic, same kind of outage, a year apart. The renegotiated contract is the left panel now.
LEAD, in one screenNot a lecture on reading contracts closely. LEAD is what tells you which number to actually go verify.
L
Link. The real outcome at stake.
Not "the software is back online," but how long a real emergency case sits without a working diagnostic read.
Grounds the whole evaluation in what actually costs something.
E
Early signal. The number that moves first.
Measured, real acknowledgment time on past incidents, not the number printed in the contract. It had been climbing for three months before the outage that exposed it.
This is the hardest step and the answer to the question: the leading number is a measured one, not a promised one.
A
Abuse. How this metric gets gamed.
A vendor can satisfy "acknowledged within fifteen minutes" with an automated ticket-received email, while no human actually starts triage for hours.
Explains why the contract's wording, not just its number, needs to be pinned down in writing.
D
Decision. What you'd actually do at each threshold.
Cross thirty minutes for two months running, and it triggers a mandatory renegotiation clause, not a strongly worded email.
Turns the metric into an action instead of a number nobody ever acts on.
A vendor can be fast to acknowledge and still slow to actually fix the problem. Ask about both separately.
The recap, one line per letter: link is the real minutes a sick animal waits, not a generic uptime number, early signal is the measured, not promised, acknowledgment time, abuse is watching for a vendor that counts an automated email as a response, and decision is a pre-agreed renegotiation trigger at thirty minutes, sustained for two months.
Any one of these three is worth pushing back on before signing, not after the first bad night.
And if you want to be sure it really works, try it somewhere elseSame four letters, an airport's AI baggage-routing vendor instead of a diagnostic imaging tool. A different gamed metric breaks the second story.
Fenmoor Ground Services runs baggage handling at a regional airport, using Taxiline AI to route checked bags between connecting flights during tight layovers. Colm Fitzgerald, Head of Ground Operations, is evaluating Taxiline's incident terms after a routing outage nearly stranded a full connecting flight's bags. Mapped onto LEAD: link is minutes of delay per missed connection, not "system uptime"; early signal is the measured time between an outage starting and a human at Taxiline actually rerouting bags manually, tracked incident by incident, not the promised figure in the contract; abuse is the same shape as Vetrix's, a vendor could count an automated status-page update as "acknowledged" while no human touches the actual reroute for an hour.
The old decision here isn't an unmeasured promise, it's a different reversal: Fenmoor's contract defined "acknowledged" as any automated system notification, written that way early on because it seemed like a fair, objective, hard-to-dispute trigger. That made sense when outages were rare and short. It stopped making sense the day Taxiline's own status page updated within ninety seconds of an outage, satisfying the letter of the contract, while the first human who could actually reroute bags by hand didn't log in for fifty-five minutes.
Fenmoor's contract had all four boxes filled in. Only "response time" had been defined loosely enough to be gamed.
Minutes of connecting-flight delay per baggage incident, before and after the SLA's wording was fixed
Same outage frequency, same airport. Naming a real human contact instead of an automated one cut average delay by roughly five times.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask for the measured acknowledgment time on past incidents, pin down what counts as acknowledged, and set a real threshold that triggers renegotiation," and stop.
Cost: there's no budget or leverage to renegotiate an existing contract right now. Say so honestly, and start by just requesting the incident logs, since seeing the real number costs nothing and tells you whether renegotiation is actually urgent.
The model gets better, for real: if a vendor's measured acknowledgment time genuinely improves and stays under the threshold for six months, that's the early signal doing its job in reverse, telling you the relationship has actually earned more trust, not less scrutiny forever.
Where people run it wrong.
They read the SLA's printed number and never ask whether the vendor has ever actually hit it.
They leave "acknowledged" undefined, letting an automated email count the same as a human starting triage.
They wait for a bad night to notice the drift, instead of watching the number every month like any other leading indicator.
How to use it live. When someone asks how you'd evaluate support terms, ask back: what's the number this contract promises, and has anyone ever checked what actually happened last time? Let the gap between those two numbers, if there is one, decide how worried to be.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "how do you evaluate this" metric question like incident response terms?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It finds the number that moves first, weeks before the outcome the contract is supposed to protect.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Naledi Okonjo, VP of Clinical Operations at Elkhorn Veterinary Group, who has run clinical operations there for nine years.
3 · THE LINK
What's the real business outcome this evaluation is protecting, in plain terms?
Tap to flip
ANSWER
How long a real emergency case sits without a working diagnostic read, not a generic "system uptime" number.
4 · THE EARLY SIGNAL
What's the leading indicator this answer says to track, and why not just the SLA's printed number?
Tap to flip
ANSWER
Measured, real acknowledgment time from past incidents. The printed number is a promise; only the logs show what actually happened.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Accepting the fifteen-minute acknowledgment promise at signing without ever requesting Vetrix's own incident logs to check it against reality.
6 · THE NUMBER
Fill in the blank: the night Milo came in, Vetrix's contract promised acknowledgment in 15 minutes. The real first human response came after ___ hours.
Tap to flip
ANSWER
About 13.5 hours (13 hours 34 minutes). The gap between the promised number and the real one is the whole point of the answer.
7 · THE REPLAY
Same outage, same clinic, but the contract now requires a named human contact with a backup. What changes?
Tap to flip
ANSWER
Acknowledgment drops to around 11 minutes, matching the renegotiated terms shown in the comparison sketch, instead of the old vendor's 94-minute pattern.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what gamed definition breaks it there?
Tap to flip
ANSWER
Fenmoor Ground Services' Taxiline AI baggage-routing tool. The gamed definition is counting an automated status-page update as "acknowledged," while no human actually rerouted bags for nearly an hour.
Check yourself Score: 0 / 0
True or false
1. True or false: the fifteen-minute acknowledgment promise in Vetrix's contract had been independently verified against real incidents before the outage that exposed it.
True
False
Show hint
Look at "the decision I would take back."
Show answer
False. It had never been checked against real incident logs until after the near miss forced the question.
Multiple choice
2. Why is measured acknowledgment time the "early signal" here, rather than the incident itself?
A. It's the easiest number for a vendor to report accurately.
B. It had been drifting upward for three months before the major outage that finally made it visible.
C. It's required by veterinary industry regulation.
D. It's the only number that appears in most vendor contracts.
Show hint
Look at the line chart.
Show answer
B. The drift crossed the concern threshold three months before the outage that exposed it, which is exactly what makes it a leading indicator.
Fill in the blank
3. Fill in the blank: after the SLA's wording was fixed at Fenmoor, average connecting-flight delay per incident fell from 47 minutes to about ___ minutes.
Show hint
Look at the bar chart in Section 4.
Show answer
9 minutes. Roughly a five-times improvement, from naming a real human contact instead of an automated one.
Short answer, apply it yourself
4. Think of a support contract or SLA you've relied on. What's the number it promises, and have you ever actually checked whether it holds up in a real incident?
Show hint
Ask whether "acknowledged" or "resolved" has ever been pinned down in writing.
Show answer
Model answer: Most people can name the promised number but not the real one, exactly the gap this answer is built to close.
Short answer, why no middle setting
5. Why couldn't Elkhorn have just "read the SLA more carefully" instead of asking for real incident logs?
Show hint
Look at the Abuse step.
Show answer
Model answer: A carefully read promise is still just a promise. Only the vendor's own past incident data shows whether that promise has ever actually been kept.
Short answer, where it wouldn't matter
6. Name a part of the Vetrix relationship where this same incident-response scrutiny genuinely doesn't need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Brief, scheduled maintenance windows. Treating a planned, announced update like an incident would waste scrutiny on something that was never a surprise.
Before you close the answer
Why this works
Tests whether you'll trust a printed contract number or go verify it against the vendor's own history, and whether you can turn a vague word like "acknowledged" into something specific enough to enforce.
Follow-up traps
"What if the vendor refuses to share their incident logs?" Response: that refusal is itself useful information, ask for it as a signed contract term before renewal, not as an optional favor.
"Isn't a thirty-minute threshold arbitrary?" Response: it's set relative to the real outcome, how long a case can safely wait, not picked as a round number; it should move if the underlying clinical tolerance changes.
If pressed
The renegotiated Vetrix contract also added a named backup escalation contact, reachable by phone, specifically because the original single-contact design meant one person being asleep or unreachable was enough to blow past the whole promised window.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.