How do you communicate that user data is or is not used for training?
Hollowick is a mid-size company that sends customer data to Quillhart, a B2B analytics platform, for reporting and insights. Baldric Onwuka is Hollowick's privacy counsel and signed off on Quillhart eighteen months ago, on the strength of one line in its policy: "we do not use customer data to train our AI models."
- Define what "training" actually means, in the policy itself.Why: a shared benchmarking model can dodge a narrow definition while still touching your data.
- Name every feature and subprocessor the promise covers, and which it doesn't.Why: one paragraph can't tell a customer whether their data reaches a third party the vendor routes requests through.
- Date every version of the policy, and summarize what changed.Why: an undated promise can't be checked against a quiet later change.
- Give a real, visible toggle a customer can check today, not just a paragraph.Why: a setting you can see is proof. A sentence you remember reading once is not.
- Don't publish full internal model architecture or pipeline details.Why: customers need a clear scope and a toggle, not a technical paper nobody outside engineering can use anyway.
How to answer this, stage by stage
This isn't a writing exercise about clearer sentences. It's about whether a customer can check the promise against something real.
Let's learn
Say we build a data-training promise, and ask what the word training is actually allowed to mean before anyone signs anything on the strength of it.
Quillhart is a B2B analytics platform. Companies like Hollowick send it customer data, and Quillhart returns reports and AI-generated insights on top of it.
Baldric signed off on Quillhart eighteen months ago after reading its policy summary once. For most of that time, nobody at Hollowick asked a follow-up question, and the one-line promise sat quietly unread in a folder nobody reopened.
The turn: whether Quillhart actually trains a shared model on customer data was never really the question Baldric could answer from the policy alone. The real problem is that "training" was never defined, so there was no way to tell the difference between a genuinely narrow promise and a technically-true one that still allowed the thing customers would object to.
What I would leave alone: Quillhart doesn't need to publish its internal model architecture or training pipeline. Customers don't need a research paper. They need a clear scope statement and a toggle they can check, which is a much smaller ask than full technical transparency.
The lesson: a data-training promise that can't be checked against a definition, a date, and a subprocessor list isn't really protecting the customer's data. It's protecting the vendor's ability to answer "yes" to two very different questions with the exact same sentence.
Now here is the same thing as a story
The short version above is what you'd say in a vendor-risk review. Read this one for how Baldric actually found the gap.
Baldric Onwuka has reviewed vendor contracts for Hollowick for nine years, long enough to have a reflex for which clauses matter and which are boilerplate. Quillhart's data-training line read, on first pass, like the second kind.
He signed off, filed the review, and moved on to the next vendor. Quillhart's insights kept arriving, useful and on schedule, and nobody at Hollowick had a reason to reopen that file for a year and a half.
Then a different analytics vendor, one Hollowick had briefly evaluated and passed on, made the news. It had used "anonymized" customer data to train a shared benchmarking model, and that model had leaked identifiable patterns back out to a completely different customer months later.
Baldric reread Quillhart's policy that same afternoon, properly this time, past the plain-language summary and into the actual defined terms. "Training" appeared nowhere with a definition attached. Nothing ruled out a shared benchmarking model built from aggregated data across Quillhart's customer base, the same shape of thing that had just gone wrong somewhere else entirely.
What that cost, in the time it took to find out: nearly a month of internal review at Hollowick, legal time nobody had budgeted, all to answer a question the original policy should have made unnecessary to ask.
Quillhart's own account team, when Baldric finally asked directly, gave a verbal reassurance within the day: "we'd never do that to your data specifically." Baldric didn't accept it. A verbal reassurance doesn't survive an account manager leaving, or a policy quietly changing again in another eighteen months.
Instead, he invoked Hollowick's contractual audit-rights clause and formally requested Quillhart's subprocessor list and internal training documentation. Six weeks later, the answer came back in writing: no shared benchmarking model existed yet, but nothing in the current contract would have prevented one, and Quillhart agreed to add a defined-scope clause and a customer-visible toggle within the quarter.
I trusted a one-line summary for eighteen months because the sentence sounded reassuring and nobody had a reason to push on it. It took a different company's failure, in the exact shape this contract had never actually ruled out, to see that a promise this important needs to survive being checked, not just sound good the first time you read it.
AUDIT, applied to a promise instead of a claimNot a compliance checklist. AUDIT is what separates a defined scope from a reassuring sentence.
The recap, one line per letter: ask who benefits is Quillhart's own legal team facing no cost for vagueness, uncover what training means is the gap between base-model and shared-model, demand the version is the undated three-month-old change, isolate what's missing is the absent toggle and subprocessor list, and test it against the contract is the formal audit request that got a real answer.
And if you want to be sure it really works, try it somewhere elseSame five letters, a consumer photo-editing app instead of a B2B analytics platform. A different industry, and this time the promise sits inside a terms-of-service page a person never rereads.
Corvid Editor is a consumer AI photo-editing app. Yevgenia Kovalenko, a professional photographer, has used it for two years and recently noticed a new "AI enhance" feature she never explicitly opted into.
Mapped onto AUDIT: ask who benefits from the wording is Corvid's legal team, who wrote "we use uploads to improve our AI features" broadly enough to cover almost anything. Uncover what training means is realizing that sentence could cover fine-tuning a style-transfer model on her specific uploaded photos, not just aggregate quality metrics. Demand the version and date is discovering the terms changed silently when "AI enhance" launched, with no re-consent prompt for existing users like her. Isolate what's missing is no per-photo opt-out and no distinction between photos she'd shared publicly and private client edits she never intended to leave her account. Test it against the contract is filing a formal data-subject access request demanding to know whether any specific uploaded photo appears in training data or model outputs.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "define training, date the policy, name the subprocessors, and give a real toggle, or the promise can't be checked," and stop.
Cost: there's no budget to build a full customer-facing data dashboard this quarter. Say so, and start with the cheapest fix: define "training" precisely in the policy text itself, since that costs almost nothing and closes the biggest gap.
The model gets better, for real: even if Quillhart's actual models genuinely never touch customer data, that's still not a reason to leave the promise undefined. A true claim with no way to check it reads exactly the same as a false one.
Where people run it wrong.
They treat a reassuring sentence as equivalent to a defined, dated, checkable scope, when the two protect completely different things.
They accept a verbal reassurance from a vendor's account team instead of a binding answer in writing.
They wait for a scandal, their own or someone else's, before ever rereading a promise they signed off on once.
How to use it live. When asked how to communicate a data-training promise, ask yourself first: could a careful customer check this sentence against a definition, a date, and a name? If any of those three is missing, the sentence is reassurance, not disclosure.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if Quillhart genuinely never intended to build a shared model?" Response: then defining the term costs them nothing and proves the point; the risk was never that Quillhart was lying, it's that the promise couldn't distinguish "never" from "not yet, and nothing stops it."
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Trust, transparency and explainability in UX
- #1 What does a user need to see to trust an AI recommendation?
- #2 Explain the difference between explainability and transparency in a product context.
- #3 How do citations change user behaviour, and what happens when they are wrong?
- #4 Design the disclosure that tells a user they are talking to an AI.
- #5 When does showing the model's reasoning help, and when does it reduce trust?
- #6 Critique a design that surfaces a chain of thought to end users.