Artifact critiqueAdvancedResponsible AI & Advanced Practice / AI product case study teardowns / #12

Examine an enterprise AI product's trust and security page as a product artifact.

AUDIT the artifact is Halcyon Ridge's public Trust & Security page, for its AI document-review copilot

Halcyon Ridge sells an AI copilot that reviews loan documents for banks and credit unions, flagging clauses a human should double check. Vashti Orenthal does vendor risk review for Northscale Credit Union, and her job this quarter is deciding whether Northscale signs with them.

The direct answer
Halcyon Ridge's trust page names no eval set, no version pin, and no failure rate anywhere on it. Don't score it on tone or badge count. Send it back with four specific questions, and treat a non-answer to any one of them as your real answer.
Do this, in order
  1. Ask which model version the page's claims were tested against, and on what date.Why: a claim with no version pin can't be checked again next quarter, or ever.
  2. Ask to see the eval set behind any accuracy number.Why: a percentage with no named test set is a number, not evidence.
  3. Ask what the page leaves out: failure rate, error categories, incident history.Why: what a report omits usually tells you more than what it includes.
  4. Run ten of your own real loan documents through it before signing anything.Why: a vendor's own number and your own result are two different facts.
  5. Note the badges as a floor, not a ceiling.Why: a compliance certificate proves a process was followed once, not that the model is good.
  6. Keep trusting badges for things badges are actually built to prove.Why: not every claim on the page is empty, and treating all of them as worthless is its own mistake.

How to answer this, stage by stage

An interviewer handing you a trust page wants to see you read past the design. Six moves get you there.

Stage 1
Name the artifact and the job it's doing
Say it like this
"This is a marketing page written to close a sale. That's not a criticism, it's just what it's for, and it changes how I read every sentence on it."
Why this works
Stops you from grading the page as if it were a neutral document.
Stage 2
Say your structure out loud
Say it like this
"I'll run AUDIT on it: ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, then test it myself."
Why this works
Shows the interviewer a repeatable method, not a one-off gut reaction.
Stage 3
Ask who paid for it
Say it like this
"Every number on this page comes from Halcyon Ridge itself. There's no independent auditor named anywhere, so I'd treat every figure as a claim, not a finding."
Why this works
Separates a vendor's own marketing copy from an outside party's finding.
Stage 4
Demand the version pin
Say it like this
"Which model version was this tested against, and on what date? If they can't answer that in one sentence, the number they gave me is already out of date."
Why this works
A model updates quietly. A claim with no date attached can't be reproduced by anyone, including the vendor.
Stage 5
Isolate what's missing
Say it like this
"There's no failure rate anywhere on this page. Not a range, not a category, nothing. That silence is the loudest thing on the page."
Why this works
What a report leaves out is usually more informative than what it shows.
Stage 6
Test it yourself, and close
Say it like this
"I'd run our own forty real loan files through it before anyone signs a contract. If our number is far from their number, that gap is the real due diligence finding, not a footnote."
Why this works
Replicating the claim on your own data is the only step that turns a hope into a fact.

Let's learn

Halcyon Ridge's copilot reads a loan file and flags the three or four clauses inside it that a human reviewer should look at twice.

Before the copilot, Northscale's own review team read every file cover to cover: about eleven pages per file, forty files a week, catching almost everything that mattered because a trained eye reads everything, slowly.

With the copilot, that same team now reads Halcyon Ridge's flags first: three or four clause snippets per file, in under two minutes, and moves on to the next file.

Knowledge spark: what's a version pin? A note saying exactly which build of a model, on which date, produced a number. Without one, nobody, not even the vendor, can go back and reproduce the claim, because the model behind it may already be a different one.

The turn: the copilot's flags are not the problem. Vashti's actual job right now is not judging whether the tool works. It's judging whether the page in front of her gives her any way to check.

Claimed pass rate vs. Vashti's own forty-file replication
100% 50% 0 98% Vendor's own page 81% Vashti's own 40 files
A seventeen-point gap between the claim and the replication, on a page that never named an eval set to begin with.

At its worst: Northscale signs a two-year contract on the strength of a badge wall and a single unlabeled percentage, and the seventeen-point gap only surfaces the first time a flagged clause turns out to be wrong on a loan that already closed.

The decision I would take back Vendor review used to treat a trust page's badge count as most of the score, since a page with SOC 2 and ISO logos looked more serious than one without. That made sense back when most vendors on Northscale's list were plain software with no model inside them, where a badge really did cover most of the real risk. It stopped making sense the day the product itself started making judgment calls a badge was never built to certify.

What I would leave alone: the badges themselves are still worth something. SOC 2 genuinely tells you the vendor has a real security process. Don't throw that signal out just because it can't tell you whether the model is accurate too.

The page was never hiding a bad number. It was hiding the fact that no real number exists yet, dressed up to look like one does.

The lesson: a trust page with no version pin and no eval set isn't halfway to a real audit. It's a promise with none of the parts that would let anyone check it, and the badge wall is what makes that easy to miss.

Now here is the same thing as a story

The short version above is what you'd say in the room. Read this one for how Vashti actually caught the gap.

Vashti has reviewed vendor contracts for six years, and she can usually tell within ten minutes of a sales call which claims are going to survive a follow-up question.

Hand sketched comparison diagram titled What the page shows vs what it hides. Left panel, a document icon labeled Shown, caption badges logos a promise. Right panel, a box icon labeled Hidden, caption no eval set no date.
Two sides of the same page, and only one of them is written to be checked.

The first read went well. Halcyon Ridge's page had a clean layout, three compliance badges, a quote from a bank CISO, and a single bold percentage: 98% flag accuracy.

Then a new hire on Vashti's team, doing a practice run before her own first vendor call, asked a question nobody on the team had asked out loud in months: "wait, accuracy on what, exactly?"

Hand sketched labeled parts diagram titled What a real trust page names. Center document icon labeled Trust Page, with four callouts: named eval set, version pin date, failure rate range, incident contact.
Four things a real trust page names. Halcyon Ridge's page named none of them.

Vashti went back to the page with that question in hand and found nothing answering it: no eval set named, no date, no version number, no failure category, nothing.

Hand sketched flow diagram titled Where the evidence trail should be. Five boxes: claim made, badge shown, evidence question mark highlighted, buyer trusts, contract signed.
The evidence step, the one in the middle, is the one Halcyon Ridge's page skips straight past.

She emailed Halcyon Ridge asking for the eval set and the version pin behind the 98% figure. Three days passed before a reply came back: a restated version of the same marketing sentence, no set, no date, no version number.

Hand sketched timeline titled The version pin, across six reviews. Five milestones, all reading pin 14 months old, the fifth emphasized as same pin again.
Vashti later found the same 14-month-old version pin quoted in five other credit unions' vendor files. Nobody had ever asked for a newer one.

She ran her own forty real loan files through the copilot before writing her recommendation. Its flags matched a human reviewer's judgment on 81 of them, not 98.

Hand sketched quadrant titled Sorting the page's own claims. Axes how specific from vague to exact, and how checkable from take our word to provable. SOC2 badge sits upper middle. 99.9% uptime sits lower middle. Bank grade security sits lower left. Pen test date sits upper right.
Only one claim on the whole page landed in the provable, exact corner. Most of it lived in the take-our-word zone.

The old process asked whether a page looked serious. The new one asks whether a page's claims could survive Vashti asking one specific follow-up question about each of them.

I used to treat a badge wall as most of the score on a vendor review, because for years it covered most of the real risk on the vendors we bought from. It took a new hire's plain question, the kind she asked because she hadn't yet learned to skip it, to show me the badges had stopped covering the part of the risk that actually mattered now.

AUDIT, five checks for a page that wants your trustNot a vibe check. AUDIT is what forces a marketing page to answer for its own claims.

Hand sketched icon list titled The five AUDIT checks. Five items: a question mark box icon labeled Ask who paid for it, a document icon labeled Uncover the eval set, a gauge icon labeled Demand the version pin, a box icon labeled Isolate what's missing, a scale icon labeled Test it yourself.
Five checks, in order. Halcyon Ridge's page fails four of the five before you even open a follow-up email.
A
Ask who paid for it.
Every claim on the page came from Halcyon Ridge itself, with no named independent auditor anywhere.
Separates a vendor's own copy from an outside party's finding.
U
Uncover the eval set.
No named test set behind the 98% figure. Not a public benchmark, not a described internal set, nothing.
A score with no eval set behind it is a number, not evidence.
D
Demand the version pin.
The same 14-month-old pin quoted in five other credit unions' files, with no date attached to the page's own claim at all.
The hardest check, and the one that exposes a claim nobody can reproduce.
I
Isolate what's missing.
No failure rate, no error category, no incident history anywhere on the page.
What's absent from a report is usually more informative than what's on it.
T
Test it yourself.
Forty real loan files, run before signing, landed at 81%, seventeen points under the vendor's own figure.
Replicating a claim on real data turns a hope into a fact, before it reaches a signature.
Days since the version pin last changed, tracked across six vendor questionnaires
420d 210d 0 Q1 Q2 Q3 Q4 Q5 Q6
A rising line here means the opposite of progress: the same old model version is still what's being sold, eighteen months on.

The recap, one line per letter: ask who paid for it means every figure traced back to Halcyon Ridge alone, uncover the eval set means no named set existed behind the 98%, demand the version pin means a 14-month-old pin nobody had refreshed, isolate what's missing means no failure rate anywhere on the page, and test it yourself means Vashti's own 81% replication was the real finding.

And if you want to be sure it really works, try it somewhere elseSame five checks, a municipal permits office instead of a credit union.

Milnrace sells an AI copilot that reviews building-permit applications for cities, flagging incomplete or non-compliant submissions before a human clerk looks at them. Reuben Kastanidis handles vendor procurement for a mid-size city's permits office.

Mapped onto AUDIT: ask who paid for it finds every claim on Milnrace's trust page sourced to Milnrace's own case-study team, with a named client city quoted but no independent reviewer anywhere. Uncover the eval set finds a real one this time, a described set of 500 real historical permit applications, though Milnrace won't say which city they came from. Demand the version pin finds an actual date, refreshed twice in the last year, a genuine point in Milnrace's favor. Isolate what's missing finds no false-negative rate anywhere, only the flag-accuracy number, meaning nobody knows how many bad applications the tool waves through uncaught. Test it yourself means Reuben runs 25 of his own city's past applications through it, before signing, specifically checking how many known-bad ones it missed rather than just how many flags matched.

Hand sketched icon list titled The five AUDIT checks, reused for the permits office example. Five items: ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself.
Same five checks. This vendor passes two of them and still fails on the one number that matters most: what it misses.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "no eval set, no version pin, no failure rate, so I'd test it myself before signing," and stop.
Cost: there's no time or budget to run your own forty-file test before a renewal deadline. Say so honestly, and ask for the vendor's own false-negative rate in writing instead, since a written number you can hold them to still beats a page with none at all.
The model gets better, for real: if Halcyon Ridge's next version genuinely improves, that's still not a reason to skip the eval-set question, a better model with no way to verify it is just a more convincing unverifiable claim.

Where people run it wrong.
They score a trust page on how many badges it has instead of on what those badges actually certify.
They accept a percentage with no eval set as if a number alone were proof.
They stop at reading the page instead of testing the product on their own real data before signing.

How to use it live. When someone hands you a trust page, ask yourself the version-pin question first, out loud if you can: which build, on what date. If the page can't answer that in one line, nothing else on it needs to be argued with.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "examine this trust and security page as a product artifact"?
Tap to flip
ANSWER
AUDIT: ask who paid for it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Vashti Orenthal, who does vendor risk review for Northscale Credit Union and can usually spot a claim that won't survive a follow-up question.
3 · THE HABIT
What did vendor review stop doing because it worked, and why did that stop being safe?
Tap to flip
ANSWER
Scoring mostly on badge count. It worked while vendors were plain software with no model inside making judgment calls a badge was never built to certify.
4 · WHAT'S MISSING
Name the three things Halcyon Ridge's trust page never names.
Tap to flip
ANSWER
A named eval set, a version pin with a date, and a failure rate. All three are absent from the page.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Treating badge count as most of the vendor-review score, a rule that made sense before AI vendors started making judgment calls no badge certifies.
6 · THE NUMBER
Fill in the blank: the vendor's page claims ___% accuracy, and Vashti's own 40-file replication found ___%.
Tap to flip
ANSWER
98% claimed, 81% replicated. A seventeen-point gap on a page with no eval set to explain it.
7 · THE REPLAY
Same vendor pitch, redesigned review process. What changes?
Tap to flip
ANSWER
Vashti asks for the version pin and eval set before the first call ends, and runs her own 40-file test before any contract review starts, instead of after.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what does it get right that Halcyon Ridge doesn't?
Tap to flip
ANSWER
Milnrace's permit-review copilot. It names a real eval set and a refreshed version pin, but still hides its false-negative rate.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Vashti's own replication on 40 real loan files found the copilot's flags matched a human reviewer on ___ of them.
Show hint
Look at the grouped-bar chart comparing the claim to the replication.
Show answer
81. The vendor's own page claimed 98%, a seventeen-point gap Vashti only found by testing it herself.
Multiple choice
2. Why does this answer treat the missing failure rate as more telling than the badge wall?
  • A. Badges are always fake and should never be trusted.
  • B. What a report leaves out is usually more informative than what it shows, and no failure rate appears anywhere on the page.
  • C. Failure rates are a legal requirement for AI vendors.
  • D. Badges only apply to non-AI software.
Show hint
Look at the "isolate what's missing" step.
Show answer
B. The AUDIT method treats a silent gap as data in itself, not as a neutral absence.
True or false
3. True or false: this answer says the SOC 2 and ISO badges on Halcyon Ridge's page are worthless and should be ignored entirely.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The badges still prove a real security process exists, they just can't certify whether the model's judgment is accurate.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Scoring vendor review mostly on badge count. It made sense while most vendors were plain software with no model making judgment calls a badge wasn't built to certify.
Short answer, where it wouldn't matter
5. Pick a product you use yourself. What's one badge, certificate, or claim on its site you've always taken at face value without checking who verified it?
Show hint
Think about a security badge, a "99.9% uptime" claim, or an app-store rating you never questioned.
Show answer
Model answer: Many people point to app-store star ratings, which can be inflated by review timing or prompts, the same shape as a badge with no version pin behind it.
Short answer, apply the number
6. If Milnrace's false-negative rate turned out to be 12%, meaning 12 in 100 bad applications get waved through uncaught, would that change Reuben's decision to sign? Why or why not?
Show hint
Think about what a false negative costs a permits office, compared to what a false positive costs.
Show answer
Model answer: Likely yes, since a false negative here means a non-compliant building gets approved, a much costlier error than a compliant one getting an extra look. That number alone might be worth negotiating a lower price or a slower rollout over.
Before you close the answer
Why this works
Tests whether you treat a vendor's marketing page as evidence or recognize it as an argument that has to earn its own proof, and whether you know which specific questions expose the gap.
Follow-up traps
"Isn't asking for the eval set unrealistic? Vendors won't share proprietary test data." Response: they don't have to share the data itself, just describe what kind of set it is and how large, enough to judge whether the number means anything.

"What if your own 40-file test is too small a sample to trust either?" Response: true, which is exactly why it's a floor, not a final answer, forty real files beats zero, and a genuinely serious vendor should welcome a larger joint test before a multi-year contract.
If pressed
Halcyon Ridge's actual model update cadence turned out to be roughly quarterly, but the trust page's headline claim had never been refreshed to match, meaning the 98% figure predates at least three silent model updates by the time Vashti read it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more