Artifact critiqueAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #9

How would you document an AI system for an external audit?

BOUND the product is Quaymark, an AI cargo-risk-scoring system at Solmara Port Authority

Solmara Port Authority uses Quaymark to score incoming shipping containers for customs inspection risk, flagging the ones worth a human officer's closer look. Renzo Bastian leads compliance and product for Quaymark, and keeps every past audit's request list pinned above his desk.

The direct answer
Documentation for an external audit breaks into five parts: data lineage, eval results, model version history, known limitations, and human-review points. Score each one honestly on a maturity scale, not a checkbox, because "we have a document" and "we can answer a specific question in minutes" are different claims. Right now Quaymark scores 7 out of 10, and the one category sitting at zero, tying a specific past decision to the exact model version that made it, is the single gap most likely to sink an audit fastest.
Do this, in order
  1. Build the version-pin log first, before anything else on the list.Why: it's the one gap that turns "explain this decision" into a two-week research project instead of a five-minute lookup.
  2. Score documentation maturity on a real scale, not a yes-or-no.Why: a "we have an eval doc" checkbox hides whether that doc can actually answer the auditor's next question.
  3. Run a mock audit before the real one.Why: it's the only way to find out whether your documentation survives a skeptical outsider's first question, not your own team's familiar one.
  4. Write the known-limitations doc honestly, including the failures.Why: an auditor trusts a document that admits what the system gets wrong more than one that claims it doesn't.
  5. Don't chase perfect documentation everywhere at once.Why: maturity is a spectrum, and spending equal effort on every category wastes time the weakest one needs most.

How to answer this, stage by stage

Nobody is grading whether you can list five document types from memory. They're grading whether you can show which one is missing and why that's the one that matters.

Stage 1
Scope it to one real system
Say it like this
"I'll answer this for Quaymark, a real cargo-risk-scoring system, since 'document an AI system' means something different depending on what's actually at stake if it's wrong."
Why this works
Grounds an abstract audit-prep question in one real, consequential system.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, the categories. Own the numbers, real evidence. Use a range, maturity isn't binary. Nail the sanity check. Direction, the gap that matters most."
Why this works
Signals a structured readiness assessment, not a vague promise to "have good documentation."
Stage 3
Break down the categories
Say it like this
"Five categories: where the training data came from, what the eval results actually show, which model version made which decision, what the system is known to get wrong, and where a human is supposed to step in."
Why this works
States the full scope before claiming any of it is done.
Stage 4
Own the numbers
Say it like this
"Our eval results are solid: run on 3,200 flagged containers, dated, versioned against model v4.2. Our version history is the weak one. We can't yet say which exact version scored a specific container from three months ago."
Why this works
Names real evidence for the strong category and a real gap for the weak one, not a blanket claim either way.
Stage 5
Give the range
Say it like this
"Overall, I'd score our audit readiness at 7 out of 10 right now, not a pass-fail. Two categories are complete, two are partial, one is effectively empty."
Why this works
A single yes-or-no claim about "being documented" implies a confidence the honest picture doesn't support.
Stage 6
Sanity-check it
Say it like this
"If a skeptical outside auditor asked, right now, which model version flagged a specific container from March, could we answer in five minutes? Today, honestly, no. That's the test that actually matters, not whether a document exists somewhere."
Why this works
Compares the documentation against the exact question an auditor is most likely to actually ask.
Stage 7
Name the direction, and close
Say it like this
"The missing version-pin log is the single gap most likely to sink this audit fastest. Fix that first, and the rest of the documentation actually means something."
Why this works
Names the one assumption worth fixing, ready for whatever the interviewer pushes on next.

Let's learn

Quaymark reads a container's manifest, origin, and shipping history, and scores how worth a closer customs look it is, before it ever reaches a human inspector's queue.

Before any real audit pressure, Quaymark's documentation lived wherever it was easiest to write it: eval results in a shared spreadsheet, model updates announced in an engineering Slack channel, limitations known only informally, by whoever had been on the team longest.

Knowledge spark: what's a version-pin log? A record tying one specific decision, made on one specific date, to the exact model version that made it, plus the eval results that version had passed at the time. Without it, "which model flagged this container" isn't a lookup, it's an investigation, because model updates and specific decisions were never actually linked together in the first place.

Now, the turn: the real risk isn't that Quaymark's documentation is thin in places. It's that "we have a document about this" and "we can answer a specific question about this" quietly became the same claim, when they had never actually meant the same thing.

Documentation maturity, by category, out of 2 points each
0 2 Data lineage 2 Eval results 2 Version history 0 Known limitations 1 Human review 2
Seven out of ten looks decent from a distance. Zero on version history is the number an auditor actually cares about.
The decision I would take back We announced model updates in an engineering Slack channel and left it there, since that was the fastest way to keep the team informed and nobody outside engineering ever asked for more. That made sense while Quaymark ran one model version for most of a year. It stopped making sense once updates started shipping every few weeks, because a Slack message that scrolls out of view is not a record anyone can search six months later.

What I would leave alone: Quaymark's data-lineage documentation, already scoring a full 2, doesn't need more investment right now. It's honestly the strongest category, and spending more effort polishing it further would come straight out of the time the version-history gap actually needs.

Seven out of ten looked like a passing grade. It was really one missing record away from being unable to answer the single question every audit actually asks first.

The lesson: documentation for an audit isn't measured by how many folders exist. It's measured by how fast a real question, asked cold, gets a real answer.

Now here is the same thing as a story

The short version above is what you'd say defending Quaymark's readiness to Solmara's board ahead of the audit. Read this one for how the gap actually got found.

Renzo Bastian has led compliance for Quaymark for three years. He can usually tell how a request will go just from how specific the first question is.

A shipper filed a formal appeal after a container got flagged and delayed for inspection, on solid grounds: their cargo had cleared customs at four other ports without incident that same month. The appeal's first question was simple: which version of Quaymark scored this container, and on what evidence.

Hand sketched icon list titled The five documentation categories. Five items: a document icon labeled Data lineage, a gauge icon labeled Eval results, a box icon labeled Model version history, a question mark box icon labeled Known limitations, a person icon labeled Human review points.
Five categories. Renzo could answer four of them from memory. The fifth took a search party.

Nobody could say immediately. Quaymark had shipped three model updates in the two months before that container was flagged, and none of them had been tied, in any searchable record, to the specific decisions made under each one.

Hand sketched flow diagram titled Tracing one flagged decision, the old way. Four boxes: Auditor asks, Search the wiki, Reconstruct by hand highlighted, Answer two weeks later.
Three steps that shouldn't have existed at all. The honest answer to a simple question took two full weeks to assemble.

It took Renzo's team two weeks of cross-referencing deployment logs, engineer memories, and a half-updated changelog to reconstruct which version had actually scored that container.

Hand sketched quadrant titled Sorting categories by maturity and importance. Axes maturity today from absent to complete, and importance to auditor from low to high. Version history sits at absent, high importance. Limitations doc sits mid-range. Eval results and data lineage sit high on both.
One category sat in the single worst corner: the thing an auditor cares about most, and the thing the team had documented least.

The shipper's appeal was eventually resolved in their favor, the flag had been a false positive from a since-corrected model bug. But the two-week delay to even answer the question was itself flagged, separately, as a governance concern by Solmara's own internal review board, months before any external audit ever arrived.

Hand sketched decision tree titled Is this documentation gap audit-fatal. Root Documentation gap found, branching to three leaves: no version pin leads to Audit fatal, no failure-case log leads to Serious recoverable, minor formatting only leads to Cosmetic.
The missing version pin took the worst branch. Everything else the team was missing was recoverable by comparison.

Renzo's team spent the following year building a version-pin log directly into Quaymark's deployment pipeline, so every scoring decision automatically recorded its model version, timestamp, and linked eval snapshot the moment it happened.

Hand sketched labeled parts diagram titled What a version-pin record needs. Center icon a document labeled Version-pin log, with four callouts: Decision timestamp, Model version ID, Linked eval snapshot, Commit hash.
Four fields, captured automatically, not four fields someone has to remember to fill in after the fact.
Hand sketched timeline titled The readiness plan. Four milestones: Version-pin log built Q1 highlighted, Eval doc formalized Q2, Limitations doc written Q3, Mock audit passed Q4.
A full year, in this order, because the version-pin log was the one thing everything else depended on being trustworthy.

The old process asked engineers to remember which version did what, well after the fact, if anyone ever asked. The new one records it automatically, the moment a decision is made, whether anyone asks or not.

Time to answer "which version scored this decision," by quarter
14d 7d 0 Before Q1 Q2 Q4 14d minutes
The same question, asked at four different points in the year, went from a two-week investigation to a five-minute lookup.

I let model updates live in an engineering Slack channel because it was the fastest way to keep the team moving, and for most of a year, updates were rare enough that nobody needed more. It took a shipper's appeal, and a genuinely embarrassing two-week wait for a simple factual question, to see that "fast to write" and "fast to retrieve later" had never been the same requirement.

BOUND, the readiness broken openNot a single grade. BOUND is what forces every category in that score to say where the evidence actually is.

B
Break it down. The categories.
Data lineage, eval results, model version history, known limitations, human-review points.
States the full scope of what "documented" actually has to cover before scoring any of it.
O
Own numbers. Each category, evidenced.
Eval run on 3,200 flagged containers, dated, versioned against model v4.2. Version history: zero decisions currently traceable to a specific model version.
Every claim has a stated source, or a stated absence, not a vague assurance either way.
U
Use a range. Maturity, not pass-fail.
Seven out of ten overall. Two categories complete, two partial, one effectively absent.
A single "we're documented" claim implies a confidence the real spread doesn't support.
N
Nail the sanity check.
Could a skeptical outside auditor get an answer to "which version scored this decision" in five minutes? Today, no.
Tests the documentation against the exact question most likely to actually get asked.
D
Direction. What sinks the audit fastest.
The missing version-pin log, the single gap most likely to turn a routine question into a two-week investigation.
The hardest step, and the one that turns a readiness score into an actual plan.

The recap, one line per letter: break it down is the five categories, own numbers is real evidence for the strong ones and a named gap for the weak one, use a range is the seven-out-of-ten spread, nail the sanity check is the five-minute lookup test, and direction is the missing version-pin log as the fastest way to sink the audit.

And if you want to be sure it really works, try it somewhere elseSame five letters, a university admissions office instead of a port authority. This time the missing category isn't version history at all.

Amsel College uses an AI tool to pre-screen scholarship applications before a human committee reviews the shortlist. Colette Furlan runs that admissions product, and prepared for her own external accreditation review the same way Renzo prepared for his.

Break it down: the same five categories, but Amsel's weakest one turned out to be known limitations, not version history. Own numbers: their model version log was solid, every scoring run tagged and timestamped from day one, but nobody had ever written down which applicant profiles the model consistently under-scored, first-generation applicants with nontraditional transcripts, until a committee member noticed the pattern by hand. Use a range: Amsel scored 8 out of 10 overall, stronger than Quaymark's readiness, with the limitations category sitting at a bare 1. Nail the sanity check: could an accreditor ask "what kinds of applicants does this system score worst" and get a real answer? Before the fix, no, only an impression from one committee member's memory. Direction: writing the known-limitations document, backed by an actual audit of scoring patterns by applicant type, was the single fix that mattered most, since version history was never actually the risk there.

Hand sketched comparison diagram titled Tracing a single flagged decision. Left panel, a question mark box icon labeled Before, caption two weeks to trace. Right panel, a document icon labeled After, caption minutes to trace.
The same before-and-after shape showed up at Amsel too, just anchored to a different missing category than Quaymark's.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "five categories, scored on a real scale, and the version-pin gap is what sinks an audit fastest," and stop.
Cost: there's no engineering budget this quarter for an automated version-pin pipeline. Say so honestly, and start with a manual log, one spreadsheet row per model deployment, since even a manual record beats a Slack channel nobody can search.
The model gets better, for real: if Quaymark's scoring accuracy improves, that says nothing about whether a specific past decision can still be traced to its model version. A better model with no version history is still a documentation gap.

Where people run it wrong.
They treat "a document exists" as the same claim as "a specific question can be answered fast," when the two are not the same thing at all.
They score readiness as pass-fail instead of a real maturity spread, hiding exactly which category needs the most attention.
They wait for the real external audit to find the gap instead of running a mock audit first, with someone playing the skeptical outsider on purpose.

How to use it live. When asked how to document an AI system for an audit, don't list the five categories and stop. Say which one is weakest today, and why that specific one is the one most likely to sink the audit first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how would you document an AI system for an external audit"?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. Direction names which documentation gap would sink the audit fastest.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Renzo Bastian, who leads compliance and product for Quaymark, Solmara Port Authority's cargo-risk-scoring system.
3 · THE FIVE CATEGORIES
What are the five documentation categories an external audit checks?
Tap to flip
ANSWER
Data lineage, eval results, model version history, known limitations, and human-review points.
4 · THE GAP
What was Quaymark's weakest documentation category?
Tap to flip
ANSWER
Model version history, scoring a zero. No specific past decision could be reliably tied to the exact model version that made it.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Announcing model updates only in an engineering Slack channel, which made sense while updates were rare but stopped working once they shipped every few weeks.
6 · THE NUMBER
Fill in the blank: Quaymark's overall documentation readiness scored ___ out of 10.
Tap to flip
ANSWER
7 out of 10. Two categories complete, two partial, and one, model version history, effectively absent.
7 · THE REPLAY
Same kind of shipper appeal, redesigned documentation. What changes?
Tap to flip
ANSWER
The version-pin log answers "which model scored this container" in minutes, instead of the two weeks it took the first time.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's its weakest category instead of version history?
Tap to flip
ANSWER
Amsel College's scholarship pre-screening tool. There, the weakest category is known limitations, an undocumented pattern of under-scoring first-generation applicants.

Check yourself Score: 0 / 0

Multiple choice
1. Why does this answer treat "we have an eval results document" and "we can answer a specific auditor question fast" as two separate claims?
  • A. Because eval documents are always inaccurate.
  • B. Because a document existing doesn't guarantee it can answer a specific, cold question quickly, which is the real test an audit applies.
  • C. Because auditors never read documents in advance.
  • D. Because eval results are unrelated to model versions.
Show hint
Look at the "nail the sanity check" step.
Show answer
B. The five-minute lookup test is what separates real readiness from a folder of documents nobody has actually tried to search under pressure.
True or false
2. True or false: this answer recommends spending equal additional effort on all five documentation categories.
  • True
  • False
Show hint
Look at "what I would leave alone" and the direction step.
Show answer
False. Data lineage, already scoring a 2, needs no more investment. Effort should go to the weakest category, version history, first.
Fill in the blank
3. Fill in the blank: before the fix, tracing which model version scored a specific decision took about ___ days.
Show hint
Look at the line chart of time to answer, by quarter.
Show answer
14 days. Down to under a day by Q2, and to minutes by the time the Q4 mock audit ran.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Announcing model updates only in an engineering Slack channel. It made sense while Quaymark ran one model version for most of a year, with no urgent need to search updates later.
Short answer, where it wouldn't matter
5. Name a documentation category at Quaymark where more investment right now genuinely isn't needed.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Data lineage. It already scores a full 2 out of 2, so polishing it further would take time away from the version-history gap that actually needs it.
Short answer, apply it yourself
6. Think of a decision your team or organization made recently. If someone asked cold, months from now, exactly why that decision was made, how long would it honestly take to answer?
Show hint
Think about whether the reasoning lives in a searchable document or only in people's memory of the conversation.
Show answer
Model answer: Most people find at least one decision whose real reasoning lives only in memory or a scrolled-past chat message, the same gap that cost Quaymark two weeks.
Before you close the answer
Why this works
Tests whether you can turn "document the system" into a real, evidenced readiness assessment, and whether you can name the single weakest link instead of describing documentation as uniformly solid.
Follow-up traps
"Isn't a 7 out of 10 already a passing grade for an audit?" Response: no single number passes or fails on its own. A 7 built on one category sitting at zero is a very different risk than a 7 spread evenly, and the auditor's actual question exposes exactly which is true.

"Couldn't you just write the missing documentation retroactively before the audit?" Response: only partially. You can reconstruct some history by hand, the way Renzo's team did during the appeal, but a retroactive record is slower to produce and easier to challenge than one captured automatically at the moment each decision was made.
If pressed
Quaymark's real mock audit, run each quarter, assigns one team member to play the skeptical outside auditor and pick a decision at random from the past ninety days, specifically to test the version-pin log's five-minute claim against a case nobody prepared in advance.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more