Artifact critiqueAdvancedDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #14

How do you communicate that user data is or is not used for training?

AUDIT the artifact is Quillhart's data-training disclosure, reviewed by Baldric Onwuka, privacy counsel at Hollowick

Hollowick is a mid-size company that sends customer data to Quillhart, a B2B analytics platform, for reporting and insights. Baldric Onwuka is Hollowick's privacy counsel and signed off on Quillhart eighteen months ago, on the strength of one line in its policy: "we do not use customer data to train our AI models."

The direct answer
Say exactly what "training" covers, fine-tuning, embeddings, a shared benchmarking model, not just the base model, and name every feature and outside subprocessor the promise does or doesn't reach. Give customers a visible, dated setting they can check today, not a paragraph written three months ago with no changelog. A promise that never defines "training" protects the vendor's wording, not the customer's data.
Do this, in order
  1. Define what "training" actually means, in the policy itself.Why: a shared benchmarking model can dodge a narrow definition while still touching your data.
  2. Name every feature and subprocessor the promise covers, and which it doesn't.Why: one paragraph can't tell a customer whether their data reaches a third party the vendor routes requests through.
  3. Date every version of the policy, and summarize what changed.Why: an undated promise can't be checked against a quiet later change.
  4. Give a real, visible toggle a customer can check today, not just a paragraph.Why: a setting you can see is proof. A sentence you remember reading once is not.
  5. Don't publish full internal model architecture or pipeline details.Why: customers need a clear scope and a toggle, not a technical paper nobody outside engineering can use anyway.

How to answer this, stage by stage

This isn't a writing exercise about clearer sentences. It's about whether a customer can check the promise against something real.

Stage 1
Scope it to one real disclosure
Say it like this
"I'll answer this against Quillhart's actual policy line: 'we do not use customer data to train our AI models.'"
Why this works
Grounds a general "communicate this well" question in one sentence you can actually test.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who benefits from the wording, uncover what 'training' actually means, demand the version and date, isolate what's missing, and test it against the contract."
Why this works
Signals a real checklist, not a gut feeling that the sentence sounds vague.
Stage 3
Ask who benefits from the wording
Say it like this
"Quillhart's own legal team wrote this line. A vague word like 'training' costs them nothing to leave undefined, and it keeps their options open."
Why this works
Names the incentive plainly, without assuming the vendor set out to deceive anyone.
Stage 4
Uncover what "training" actually means
Say it like this
"The promise covers the base model. It says nothing about a shared benchmarking model that ingests aggregated customer data, which most people would still call training."
Why this works
Shows the gap between the plain reading of a word and its narrowest legal meaning.
Stage 5
Demand the version and the date
Say it like this
"This policy changed three months ago with no visible changelog. Nothing tells a customer whether data collected before that change is covered the same way."
Why this works
Names the exact gap an undated promise leaves open.
Stage 6
Isolate what's missing
Say it like this
"There's no per-feature breakdown, no subprocessor list, and no toggle a customer can check today. Just one paragraph and a promise to trust it."
Why this works
Lists the concrete gaps rather than a general complaint about clarity.
Stage 7
Test it against the contract
Say it like this
"I invoked our audit-rights clause and requested the actual subprocessor list and training documentation, instead of trusting the plain-language summary a second time."
Why this works
Shows a real, binding way to check the claim, not a request for reassurance.
Stage 8
Close on the one line
Say it like this
"A data-training promise with an undefined word at its center isn't dishonest, necessarily. It's just unfinished. Finish it with a definition, a date, and a toggle, or it isn't really a promise yet."
Why this works
Restates the standard plainly, ready for whatever gets pushed on next.

Let's learn

Say we build a data-training promise, and ask what the word training is actually allowed to mean before anyone signs anything on the strength of it.

Quillhart is a B2B analytics platform. Companies like Hollowick send it customer data, and Quillhart returns reports and AI-generated insights on top of it.

Knowledge spark: what's a subprocessor? A third party a vendor routes some of your data through to provide part of its service, a hosting provider, a separate model API, an analytics partner. A customer's data can reach a subprocessor even when the main vendor's own promise says nothing about it.

Baldric signed off on Quillhart eighteen months ago after reading its policy summary once. For most of that time, nobody at Hollowick asked a follow-up question, and the one-line promise sat quietly unread in a folder nobody reopened.

B2B SaaS vendors that define "training" precisely in their data policy
50 25 0 Define clearly: 12 Vague or undefined: 38
Quillhart's policy sits in the larger group. Vague is the industry default, not Quillhart's unique failing.

The turn: whether Quillhart actually trains a shared model on customer data was never really the question Baldric could answer from the policy alone. The real problem is that "training" was never defined, so there was no way to tell the difference between a genuinely narrow promise and a technically-true one that still allowed the thing customers would object to.

The decision I would take back Quillhart's product and legal teams decided years ago not to build a customer-facing dashboard showing data-usage status per account, since a plain-English paragraph was cheaper to write and nobody was asking detailed questions yet. That made sense while nobody asked. It stopped making sense the moment a competitor's failure made every customer start asking at once.

What I would leave alone: Quillhart doesn't need to publish its internal model architecture or training pipeline. Customers don't need a research paper. They need a clear scope statement and a toggle they can check, which is a much smaller ask than full technical transparency.

An undefined word at the center of a promise isn't a lie. It's just a promise nobody has finished writing yet.

The lesson: a data-training promise that can't be checked against a definition, a date, and a subprocessor list isn't really protecting the customer's data. It's protecting the vendor's ability to answer "yes" to two very different questions with the exact same sentence.

Hand sketched metaphor scene titled Two definitions of training, one word. Left, a box icon labeled Narrow, caption base model pretraining only. Right, a gauge icon labeled Broad, caption any model touching your data.
Quillhart's promise only ever answered for the box on the left. Customers assumed it covered the one on the right.

Now here is the same thing as a story

The short version above is what you'd say in a vendor-risk review. Read this one for how Baldric actually found the gap.

Baldric Onwuka has reviewed vendor contracts for Hollowick for nine years, long enough to have a reflex for which clauses matter and which are boilerplate. Quillhart's data-training line read, on first pass, like the second kind.

He signed off, filed the review, and moved on to the next vendor. Quillhart's insights kept arriving, useful and on schedule, and nobody at Hollowick had a reason to reopen that file for a year and a half.

Hand sketched timeline titled Baldric's timeline. Four milestones: signs with Quillhart trusts the plain summary, a rival's scandal breaks aggregated data trained a shared model highlighted, rereads the policy closely training left undefined, requests the subprocessor list via contract audit rights.
Nothing about Quillhart changed the week Baldric started worrying. A different company's failure did all the work.

Then a different analytics vendor, one Hollowick had briefly evaluated and passed on, made the news. It had used "anonymized" customer data to train a shared benchmarking model, and that model had leaked identifiable patterns back out to a completely different customer months later.

Baldric reread Quillhart's policy that same afternoon, properly this time, past the plain-language summary and into the actual defined terms. "Training" appeared nowhere with a definition attached. Nothing ruled out a shared benchmarking model built from aggregated data across Quillhart's customer base, the same shape of thing that had just gone wrong somewhere else entirely.

Hand sketched flow diagram titled Where the data actually goes, unlabeled today. Four boxes: customer uploads, Quillhart platform, an undefined step highlighted with a question mark, shared benchmarking model.
One box in this chain had no label at all. That's not a small gap. It's the exact shape of the failure that had just happened elsewhere.

What that cost, in the time it took to find out: nearly a month of internal review at Hollowick, legal time nobody had budgeted, all to answer a question the original policy should have made unnecessary to ask.

Hand sketched comparison diagram titled The promise vs what Baldric found. Left panel, a document icon labeled Policy says, caption we do not train on your data. Right panel, a question mark box icon labeled Reality, caption a shared model still ingests aggregates.
Both of these could be true at once, which is exactly the problem with a promise that never defines its own key word.

Quillhart's own account team, when Baldric finally asked directly, gave a verbal reassurance within the day: "we'd never do that to your data specifically." Baldric didn't accept it. A verbal reassurance doesn't survive an account manager leaving, or a policy quietly changing again in another eighteen months.

Hand sketched labeled parts diagram titled What one honest promise needs. Center document icon labeled The promise, with four callouts: defined scope, per-feature breakdown, subprocessor list, visible dated toggle.
Quillhart's actual policy had none of these four. All four were addable without touching a single line of model code.

Instead, he invoked Hollowick's contractual audit-rights clause and formally requested Quillhart's subprocessor list and internal training documentation. Six weeks later, the answer came back in writing: no shared benchmarking model existed yet, but nothing in the current contract would have prevented one, and Quillhart agreed to add a defined-scope clause and a customer-visible toggle within the quarter.

I trusted a one-line summary for eighteen months because the sentence sounded reassuring and nobody had a reason to push on it. It took a different company's failure, in the exact shape this contract had never actually ruled out, to see that a promise this important needs to survive being checked, not just sound good the first time you read it.

AUDIT, applied to a promise instead of a claimNot a compliance checklist. AUDIT is what separates a defined scope from a reassuring sentence.

A
Ask who benefits from the wording.
Quillhart's own legal team, who pay nothing to leave "training" undefined and keep their options open.
Names the incentive without assuming bad intent.
U
Uncover what "training" means. The hard step.
The promise covers base-model pretraining. It says nothing about a shared benchmarking model built from aggregated data.
This is the actual gap the whole promise rests on.
D
Demand the version and the date.
Changed three months prior, no changelog, no statement on whether older data is grandfathered differently.
An undated promise can't be checked against its own later changes.
I
Isolate what's missing.
No per-feature breakdown, no subprocessor list, no visible toggle, just one paragraph and a request to trust it.
What's absent here is louder than what's stated.
T
Test it against the contract.
Formal audit-rights request for the subprocessor list and training documentation, not a verbal reassurance.
Replaces a reassuring sentence with a binding, checkable answer.
Monthly support tickets asking "do you train on our data"
45 22 0 3 4 5 41 38 22 Scandal month
The ambiguity cost nothing for three months, then all at once. That's exactly what an absent, checkable state looks like from the outside.

The recap, one line per letter: ask who benefits is Quillhart's own legal team facing no cost for vagueness, uncover what training means is the gap between base-model and shared-model, demand the version is the undated three-month-old change, isolate what's missing is the absent toggle and subprocessor list, and test it against the contract is the formal audit request that got a real answer.

And if you want to be sure it really works, try it somewhere elseSame five letters, a consumer photo-editing app instead of a B2B analytics platform. A different industry, and this time the promise sits inside a terms-of-service page a person never rereads.

Corvid Editor is a consumer AI photo-editing app. Yevgenia Kovalenko, a professional photographer, has used it for two years and recently noticed a new "AI enhance" feature she never explicitly opted into.

Mapped onto AUDIT: ask who benefits from the wording is Corvid's legal team, who wrote "we use uploads to improve our AI features" broadly enough to cover almost anything. Uncover what training means is realizing that sentence could cover fine-tuning a style-transfer model on her specific uploaded photos, not just aggregate quality metrics. Demand the version and date is discovering the terms changed silently when "AI enhance" launched, with no re-consent prompt for existing users like her. Isolate what's missing is no per-photo opt-out and no distinction between photos she'd shared publicly and private client edits she never intended to leave her account. Test it against the contract is filing a formal data-subject access request demanding to know whether any specific uploaded photo appears in training data or model outputs.

Hand sketched icon list titled Yevgenia's own AUDIT trail. Four items: terms changed with no re-consent prompt, no per-photo opt-out offered, no public vs private distinction made, filed a data-subject access request.
None of these four required Corvid to change how its model actually works. All four were about what customers could check.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "define training, date the policy, name the subprocessors, and give a real toggle, or the promise can't be checked," and stop.
Cost: there's no budget to build a full customer-facing data dashboard this quarter. Say so, and start with the cheapest fix: define "training" precisely in the policy text itself, since that costs almost nothing and closes the biggest gap.
The model gets better, for real: even if Quillhart's actual models genuinely never touch customer data, that's still not a reason to leave the promise undefined. A true claim with no way to check it reads exactly the same as a false one.

Where people run it wrong.
They treat a reassuring sentence as equivalent to a defined, dated, checkable scope, when the two protect completely different things.
They accept a verbal reassurance from a vendor's account team instead of a binding answer in writing.
They wait for a scandal, their own or someone else's, before ever rereading a promise they signed off on once.

How to use it live. When asked how to communicate a data-training promise, ask yourself first: could a careful customer check this sentence against a definition, a date, and a name? If any of those three is missing, the sentence is reassurance, not disclosure.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how do you communicate that user data is or isn't used for training"?
Tap to flip
ANSWER
AUDIT: ask who benefits, uncover the definition, demand the version, isolate what's missing, test it against the contract.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Baldric Onwuka, Hollowick's privacy counsel, who signed off on Quillhart eighteen months before rereading its policy closely.
3 · THE HABIT
What had Baldric stopped doing for eighteen months?
Tap to flip
ANSWER
Rereading Quillhart's actual policy terms. He trusted the plain-language summary from his first review and never reopened the file.
4 · THE GAP
What did "training" fail to cover in Quillhart's promise?
Tap to flip
ANSWER
A shared benchmarking model built from aggregated customer data, which most people would call training even though the policy's narrow wording might not.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Not building a customer-facing dashboard showing data-usage status per account, since a plain paragraph was cheaper while nobody was asking.
6 · THE NUMBER
Fill in the blank: monthly support tickets asking about data training jumped to ___ the month a rival's scandal broke, from a baseline near 3 to 5.
Tap to flip
ANSWER
41. The cost of the ambiguity had been building invisibly for months before it became visible all at once.
7 · THE REPLAY
Same vendor review, redesigned promise. What changes?
Tap to flip
ANSWER
Quillhart's policy defines "training," names its subprocessors, and gives Hollowick a checkable toggle, closing the six-week review gap before it's ever needed again.
8 · CROSS PRODUCT TRANSFER
Section 4 runs AUDIT again on a different product. Which one, and what's the equivalent of the audit-rights request?
Tap to flip
ANSWER
Corvid Editor, a photo-editing app. There, it's Yevgenia's formal data-subject access request about her own specific uploaded photos.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer concludes that Quillhart was definitely lying about not training on customer data.
  • True
  • False
Show hint
Look at the audit-rights result: no shared model existed yet.
Show answer
False. The finding was that nothing in the contract ruled a shared model out, not that Quillhart had already built one.
Multiple choice
2. Why does this answer say Quillhart's promise "protects the vendor's wording, not the customer's data"?
  • A. Because Quillhart's lawyers admitted to writing a misleading sentence.
  • B. Because "training" was never defined, letting the same sentence be true under both a narrow and a broad reading.
  • C. Because Quillhart charges extra for a clearer policy.
  • D. Because the policy was written in a foreign language.
Show hint
Look at the "uncover what training means" stage.
Show answer
B. An undefined key term lets one sentence answer two very different questions the exact same way.
Fill in the blank
3. Fill in the blank: of the 50 B2B SaaS vendors reviewed, only ___ define what "training" means precisely in their data policy.
Show hint
Look at the bar chart comparing vendors.
Show answer
12. 38 of the 50 leave it vague or undefined, and Quillhart was one of them.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Not building a customer-facing data-usage dashboard, since a plain-English paragraph was cheaper while nobody asked detailed questions.
Short answer, apply it yourself
5. Pick an app on your phone that says it does or doesn't use your data for training. What's the one word in that sentence you'd want defined before you trusted it?
Show hint
Think about words like "improve," "personalize," or "training" itself.
Show answer
Model answer: Most people land on whatever verb the policy uses instead of "train" (improve, personalize, enhance), since that's usually the word doing the most undefined work.
Before you close the answer
Why this works
Tests whether you can tell a reassuring sentence apart from a checkable promise, and whether you know what specifically makes a data-usage claim verifiable instead of just well-written.
Follow-up traps
"Isn't demanding a full subprocessor list excessive for most customers?" Response: the ask isn't a subprocessor list for every casual user, it's a defined scope and a visible toggle in the policy itself; the formal audit request is a fallback for exactly the case where those are missing.

"What if Quillhart genuinely never intended to build a shared model?" Response: then defining the term costs them nothing and proves the point; the risk was never that Quillhart was lying, it's that the promise couldn't distinguish "never" from "not yet, and nothing stops it."
If pressed
Hollowick's contract audit-rights clause specifically allows a request for training-data documentation once per calendar year without additional fees, a term Baldric had negotiated in for an unrelated reason years earlier and only thought to use once this specific question came up.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more