ConceptAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #4

What questions will your legal team ask about training data, and how do you prepare?

ORDER the product is Colophon, an AI tool that drafts catalog descriptions for scanned manuscripts

Thessaly University Library holds four centuries of donated letters, diaries, and estate papers. Colophon is the AI tool that reads a freshly scanned manuscript and drafts the catalog description a human archivist used to write by hand. Bram Oosterhuis leads the digital collections team, and keeps a paper index card, one per donated collection, in a box on his desk.

The direct answer
Legal will ask five things about every document Colophon trains on: where it came from, whether the library actually has the right to use it this way, whether it names a living person, who holds the copyright, and whether you can produce the answer to all four on demand. Prepare by building the provenance record before legal ever asks, starting with the collections where a donor restriction is most likely, not by waiting for the question and scrambling for an answer under a launch deadline.
Do this, in order
  1. Get the licensing sign-off on donor-restricted collections first.Why: it's the hardest thing to undo once a model has trained on something it shouldn't have.
  2. Run a provenance audit before you need one for a specific complaint.Why: everything else, including the licensing check, depends on actually knowing where each document came from.
  3. Sample-check a small slice before committing to a full inventory.Why: it's cheap, and it tells you fast whether the problem is small or much bigger than expected.
  4. Leave a full re-scan of the whole archive for later, if ever.Why: it's the most expensive item on the list and the least likely to be the actual bottleneck.
  5. Keep the provenance record updated as new collections arrive, not as a one-time project.Why: legal's questions don't stop after the first launch, and neither should the record that answers them.

How to answer this, stage by stage

Nobody is grading whether you can list every legal question word for word. They're grading whether you can show which prep task actually has to happen first.

Stage 1
Scope it to one real archive
Say it like this
"I'll answer this for Colophon, a tool trained on scanned manuscripts at a university library, since 'training data questions' land very differently depending on what the data actually is."
Why this works
Grounds an abstract legal-prep question in one real, physical collection.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what we're protecting. Reversibility, which gap is hardest to undo. Dependency, what blocks what. Evidence, what's cheap to check first. Rank, the actual order."
Why this works
Signals this is a prioritization problem, not a memorized list of legal questions.
Stage 3
Reframe the question
Say it like this
"Legal isn't really testing whether I know their five questions. They're testing whether the answers already exist somewhere, or whether I'd be inventing them under a deadline."
Why this works
Moves the answer from "guess the quiz" to "show the actual readiness."
Stage 4
Give the one decision
Say it like this
"I'd build the provenance record before legal asks for it, starting with the donor-restricted collections, since those are the ones most likely to have a real, undoable problem hiding in them."
Why this works
This is deliverable 0, spoken as a concrete first move instead of a vague promise to "be prepared."
Stage 5
Prove it with a failure
Say it like this
"A curator once asked, almost in passing, whether we were allowed to train on the Quillenbeck family's letters. Nobody could answer immediately, because 'digitized' and 'cleared for reuse' had never been tracked as two separate things."
Why this works
A real, specific gap beats a general claim that provenance "matters."
Stage 6
Say what you'd measure
Say it like this
"I'd track what share of the archive still has an unknown rights status. If that number isn't shrinking every month, the provenance work has stalled, even if nobody's complained yet."
Why this works
Shows the prep work gets tracked over time, not treated as done after one sprint.
Stage 7
Close on the one line
Say it like this
"Answer the questions before they're asked, starting with the collection where being wrong costs the most to undo."
Why this works
Leaves the interviewer with the ranking logic, not just a list of legal questions.

Let's learn

Colophon reads a freshly scanned manuscript, a letter, a diary page, an estate ledger, and drafts the short catalog description an archivist used to write from scratch, one document at a time.

Before Colophon existed, "digitized" was the only status the library tracked on any document. A letter scanned in 2019 and a letter scanned last month sat in the same digital shelf, with no separate note on whether either one was actually cleared for a new use like training a model.

Knowledge spark: what's a provenance record? A short file, one per document or collection, answering where it came from, who has rights over it, and any conditions the donor attached. Libraries have always kept some version of this for physical loans. Training an AI model on the same material just means the record now has to answer a new question too: can this specific use happen at all.

Now, with Colophon live, every newly digitized document flows straight into training data, on the same day it's scanned, unless someone actively stops it.

The digitized archive, by rights status
0% 70% Public domain 61% Unknown 18% Donor-restricted 12% In copyright 9%
Nearly a fifth of the archive, 18 percent, had no confirmed rights status at all when Colophon launched. That slice is where a real problem was most likely to be hiding.

The turn: the unknown 18 percent isn't really the danger. The real danger is that "unknown" and "fine" had always been treated as the same thing, because nobody had ever asked the question out loud until a curator did, almost in passing.

The decision I would take back We tracked a single status field, "digitized," for every scanned document, because at the time that was the only question anyone asked: had it been scanned yet or not. That made sense while scanned material only ever went back into the reading room. It stopped making sense the moment scanned material started training a model that generates public-facing text, and "digitized" quietly started standing in for "cleared," when it had never meant that at all.

What I would leave alone: Thessaly's collection of pre-1850 university administrative records needs no special provenance review. Public institutional records that old carry essentially no rights risk, and treating them like the donor-restricted collections would waste real audit time on the safest slice of the archive.

The eighteen percent unknown was never the real risk. The real risk was that nobody had ever separated "we scanned it" from "we're allowed to use it this way."

The lesson: a legal question about training data is really a question about whether your records already separate what you have from what you're allowed to do with it. If they don't, every answer legal asks for has to be built from scratch, under whatever deadline happens to be live that week.

Now here is the same thing as a story

The short version above is what you'd say defending Colophon's launch readiness to the university's general counsel. Read this one for how the gap actually got found.

Bram Oosterhuis has led Thessaly's digital collections team for five years. He can usually tell, just from the box a donation arrived in, whether a collection is going to have unusual restrictions attached.

Colophon had been drafting catalog descriptions for eight months, quietly, on whatever the scanning team fed it that week. Nobody had drawn a line between "material we've digitized" and "material we've confirmed we can reuse," because for eight months nothing had gone wrong.

Hand sketched icon list titled What legal will ask about the training data. Five items: a document icon labeled Where did this come from, a scale icon labeled Do we have the right to use it, a person icon labeled Does it name living people, a question mark box icon labeled Who holds the copyright, a box icon labeled Can we prove it on demand.
Five plain questions. Bram realized he could answer all five confidently for maybe a third of the archive.

Then, at a routine cataloguing meeting, a rare-books curator asked a question that wasn't meant to be dramatic at all: "Wait, are we allowed to train on the Quillenbeck family's letters? Didn't they restrict access to the correspondence about the divorce?"

Hand sketched decision tree titled Does this document need a rights review. Root Scanned document, branching to four leaves: public domain confirmed leads to Clear to use, donor-restricted leads to Needs sign-off, in copyright unclear holder leads to Full review, names a living person leads to Redact or review.
The Quillenbeck letters took the second branch. Nobody had ever actually walked them down it.

Nobody in the room knew for certain. The donor agreement was in a physical folder in a different building, and nobody could say, without checking, whether Colophon had already trained on those specific letters.

Hand sketched labeled parts diagram titled What's in a provenance record. Center icon a document labeled Provenance record, with four callouts: Source collection, Rights status, Donor restrictions, Date digitized.
Three of these four fields already existed for most collections. Rights status, the one that actually mattered, was the one nobody had filled in.

Bram spent the next week running a sample audit, fifty documents pulled at random, just to see how big the real gap was before committing the whole team to a full inventory.

Hand sketched quadrant titled Sorting the prep tasks. Axes cost to do now from cheap to expensive, and risk if skipped from low to high. Licensing sign-off sits high cost, high risk. Provenance audit sits lower cost, high risk. Sample check sits cheapest, moderate risk. Full re-scan sits most expensive, low risk.
Full re-scanning, the most expensive option, turned out to be the least useful one. The provenance audit, far cheaper, sat right where the real risk was.

The sample confirmed it: about one in six documents had no rights status recorded at all, and the Quillenbeck letters were among them, sitting quietly in Colophon's training set for the last eight months.

Hand sketched flow diagram titled What unblocks what. Four boxes in sequence: Data inventory, Rights status known highlighted, Licensing cleared, Legal signs off.
Legal sign-off was always the last box. It just couldn't move until the second one, rights status, actually got filled in first.

Bram pulled the Quillenbeck letters out of the active training set that same week, brought the family's original 1987 donor agreement to the university's counsel, and got a narrow written clearance for the correspondence that wasn't restricted, three weeks later.

Hand sketched comparison diagram titled Reversible or not. Left panel, a box icon labeled Pre-launch check, caption cheap, redo anytime. Right panel, a question mark box icon labeled Post-launch takedown, caption public, hard to undo.
Pulling the letters out before anyone downstream had used Colophon's output cost a week. Doing it after a public catalog entry existed would have cost far more.

The old approach treated every scanned document as reusable the moment it was scanned. The new one asks the rights question first, and only feeds a document into training once that question has an actual answer on file.

I built the single "digitized" status field because it answered the only question anyone was asking at the time. It took one curator's offhand question at a meeting that was supposed to be about something else entirely to see that the field had been quietly doing double duty for years, standing in for a promise it never actually made.

ORDER, ranked by what breaks firstNot a checklist. ORDER is what forces you to say which prep task actually has to happen before the others.

O
Outcome. What we're protecting.
Launching Colophon without a rights problem discovered after the fact, once its output is already public.
Without a stated outcome, any ranking is just opinion.
R
Reversibility. The hard step.
Training on a donor-restricted letter is nearly impossible to fully undo once the model's output is public. A pre-launch check is cheap to redo as often as needed.
The asymmetry that decides the whole ranking.
D
Dependency. What blocks what.
Legal sign-off can't happen until rights status is known, and rights status can't be known without a real inventory first.
Shows the order isn't a preference, it's forced by what depends on what.
E
Evidence. What's cheap to learn first.
A fifty-document sample audit, done in a week, showed the real size of the gap before committing to a full inventory.
A cheap check that sized the problem before spending real budget on it.
R
Rank. The actual order.
Sample audit first, then a full provenance pass on donor-restricted collections, then licensing sign-off, with a full re-scan left off the list entirely.
Defends the top pick in one line: cheapest first, hardest-to-undo second.
Documents with unresolved rights status, by audit week
18% 9% 0% Wk 1 Wk 2 Wk 3 Wk 4 Wk 6 3%
The unresolved slice shrank steadily once the audit was actually staffed and running, not because the archive got smaller but because someone was finally checking.

The recap, one line per letter: outcome is launching without a rights problem discovered after the fact, reversibility is a pre-launch check versus a public takedown, dependency is legal sign-off waiting on rights status waiting on inventory, evidence is the fifty-document sample that sized the real gap, and rank is sample first, restricted collections second, full re-scan never.

And if you want to be sure it really works, try it somewhere elseSame five letters, a small-town newspaper's photo morgue instead of a university archive. This time the hardest-to-undo item isn't a document at all.

The Cascade Weekly Ledger, a small regional newspaper, is training an AI captioning tool on eighty years of its own archived photographs to speed up digitizing back issues. Farida Nasser edits the paper and is handling the legal prep herself, since there's no in-house counsel.

Outcome: publish digitized back issues with accurate captions, without the paper being sued by a freelance photographer whose work it never actually owned outright. Reversibility: identifying which decades used staff photographers, whose work the paper owns, versus freelancers, whose contracts often only licensed one-time print use, is nearly impossible to sort out after captions built from those photos are already published online. A quick contract-file check, by contrast, is cheap and repeatable. Dependency: knowing which decade a photo comes from unlocks knowing which contract template applied then, which unlocks knowing whether the paper can use it this way at all. Evidence: pulling twenty photo captions at random from the 1990s freelance-heavy years showed that nearly half lacked any surviving contract on file. Rank: check the freelance-heavy decades first, since that's where an unreversible mistake is most likely, and leave the staff-photographer decades, clearly owned outright, for a lighter later pass.

Hand sketched timeline titled The actual plan, timed. Four milestones: Sample audit week 1, Licensing check week 2, Rights clearance week 3, Legal sign-off week 4 highlighted.
Four weeks, in this order, for the newspaper's photo archive too. Skip the sample audit and every later week gets guessed at instead of planned.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "sample-check first to size the gap, then fix the hardest-to-undo problem before the cheap one," and stop.
Cost: there's no budget for a full-time archivist to run the provenance audit. Say so honestly, and start with the sample check alone, done by whoever's available, since even a partial answer beats no answer when legal asks.
The model gets better, for real: if Colophon's descriptions get more accurate, that changes nothing about the rights question underneath. A better model trained on the same unresolved letters is still trained on unresolved letters.

Where people run it wrong.
They wait for legal to ask the questions instead of building the record before anyone asks.
They treat "we've always had this data" as the same thing as "we've confirmed we can use it this way."
They start with the most expensive fix, a full re-scan or a full manual review, instead of a cheap sample that shows whether the problem is even that big.

How to use it live. When asked what legal will ask about training data, don't recite the questions from memory. Say which one is hardest to undo if the answer turns out wrong, and start there.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what will legal ask about training data, and how do you prepare"?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Reversibility is the step that decides which prep task jumps to the front of the line.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Bram Oosterhuis, who leads Thessaly University Library's digital collections team and keeps a paper index card per donated collection.
3 · THE FIVE QUESTIONS
What are the five things legal will ask about the training data?
Tap to flip
ANSWER
Where it came from, whether you have the right to use it, whether it names a living person, who holds the copyright, and whether you can prove all four on demand.
4 · THE GAP
What did the sample audit find?
Tap to flip
ANSWER
About one in six documents had no recorded rights status at all, including the Quillenbeck family's restricted letters, which had been in Colophon's training set for eight months.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Tracking only a single "digitized" status field for every document, letting it quietly stand in for "cleared to use," when it had never actually meant that.
6 · THE NUMBER
Fill in the blank: about ___ percent of the archive had no confirmed rights status when Colophon launched.
Tap to flip
ANSWER
18 percent. That share fell to about 3 percent within six weeks once the provenance audit was actually staffed and running.
7 · THE REPLAY
Same kind of restricted letters, redesigned process. What changes?
Tap to flip
ANSWER
A new collection's rights status gets checked at intake, before it ever reaches Colophon's training set, instead of eight months after the fact.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's hardest to undo there?
Tap to flip
ANSWER
The Cascade Weekly Ledger's photo-captioning tool. There, the hardest-to-undo problem is publishing captions built from freelancer photos the paper never actually owned outright.

Check yourself Score: 0 / 0

Multiple choice
1. Why does this answer rank the licensing sign-off on restricted collections above a full re-scan of the archive?
  • A. Because a full re-scan is technically impossible.
  • B. Because a restricted document already trained on is far harder to undo than a re-scan is cheap and low-risk to skip.
  • C. Because donors always complain louder than copyright holders.
  • D. Because re-scanning costs nothing at all.
Show hint
Look at the quadrant sorting prep tasks by cost and risk.
Show answer
B. The ranking follows reversibility, not raw cost. A cheap, low-risk task like re-scanning loses to an expensive but high-risk one every time.
True or false
2. True or false: the sample audit was meant to replace the full provenance inventory entirely.
  • True
  • False
Show hint
Look at the "evidence" step of ORDER.
Show answer
False. The sample was a cheap way to size the problem before committing to the full inventory, not a substitute for doing it.
Fill in the blank
3. Fill in the blank: the unresolved rights-status share fell from 18 percent to about ___ percent within six weeks of the audit starting.
Show hint
Look at the line chart of unresolved documents by audit week.
Show answer
3 percent. A steady decline, not a single fix, as the audit worked through the collection week by week.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Tracking only a single "digitized" field for every document. It made sense while digitized material only ever went back into the reading room, with no separate reuse question to answer.
Short answer, where it wouldn't matter
5. Name a collection at Thessaly where this provenance concern genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The pre-1850 university administrative records. Public institutional records that old carry essentially no rights risk, so a full review would waste audit time better spent elsewhere.
Short answer, apply it yourself
6. Think of a dataset you've worked with, even a spreadsheet at a old job. If someone asked where every row of it actually came from, could you answer confidently?
Show hint
Think about whether "we've always had this" and "we know exactly where this came from" are actually the same thing in that dataset.
Show answer
Model answer: Most people find at least one part of a dataset they've used where the honest answer is "I assumed it was fine," the same gap the Quillenbeck letters sat in for eight months.
Before you close the answer
Why this works
Tests whether you can turn a legal-prep question into a real prioritization exercise, ranked by what's hardest to undo, rather than a memorized list of things a lawyer might say.
Follow-up traps
"Isn't a full re-scan the safest option, even if it's expensive?" Response: expensive doesn't mean highest-risk. The re-scan sat in the low-risk corner of the quadrant, since most of the archive's rights status wasn't actually in doubt, only a specific slice was.

"What if the sample audit misses a problem the full inventory would have caught?" Response: the sample was never meant to replace the full inventory, only to size it fast enough to know how urgently to staff it.
If pressed
Thessaly's actual provenance record also logs which specific model version a document was used to train, not just whether it was used, so if a rights problem surfaces later, the team can say precisely which model outputs might be affected, not just that "some documents" were involved.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more