ConceptAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #5
Describe the copyright exposure of a generative feature.
BOUND the product is StyleCast, a generative mockup tool at Selvage Studio, a made-to-measure tailoring service
Selvage Studio builds custom garments to measure. StyleCast is the feature that takes a customer's photo and a style reference and renders a mockup of the finished piece before a single stitch is cut. Coen Vermaas is the product lead who has to put a real number on the copyright exposure for the executive team, not just a feeling about it.
The direct answer
The exposure runs from about 150,000 dollars a year on the low end to 1.6 million on the high end, and the single assumption that swings it most is what share of the training photos were scraped from runway coverage instead of properly licensed. Cut the scraped share from 55 percent to 15 percent and expected yearly exposure falls by more than half, which is a cheaper fix than waiting for the first claim to land.
Do this, in order
Re-license the training corpus toward stock and consented photos first.Why: it's the single assumption that swings the whole exposure estimate the most.
Add a similarity check between any generated mockup and known recent designs before it reaches a customer.Why: it catches the specific failure mode, a near-identical output, before it ever ships.
Write the indemnification and takedown terms into the customer contract now, not after a claim.Why: it caps Selvage's own liability and gives a clear, fast response when a designer does object.
Carry the range, not a single number, into any budget or insurance conversation.Why: a single confident figure hides how much the estimate depends on one unverified assumption.
Leave the customer-upload slice of the corpus alone.Why: those photos already carry the customer's own consent, and it's the smallest share of real risk in the whole corpus.
How to answer this, stage by stage
Nobody is grading whether your number lands exactly right. They're grading whether you can show the arithmetic and name the one input worth fixing first.
Stage 1
Scope it to one real feature
Say it like this
"I'll answer this for StyleCast, the generative mockup feature at a made-to-measure tailoring company, since 'copyright exposure' means something different for every kind of generated output."
Why this works
Grounds an abstract legal question in one concrete feature with a real training corpus.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, the equation. Own the numbers, my assumptions. Use a range. Nail the sanity check. Direction, what swings it most."
Why this works
Signals a real estimate with shown work, not a guess dressed up as a legal opinion.
Stage 3
Break down the equation
Say it like this
"Expected yearly exposure equals the number of valid claims we'd expect in a year, times the average cost to settle one, plus whatever we owe customers under our own indemnification terms."
Why this works
States the equation before touching a single number, so the estimate isn't a guess with a confident tone.
Stage 4
Own the numbers
Say it like this
"I'd assume one to four valid claims a year, since 220,000 of our 400,000 training images are scraped runway photos from roughly 600 collections, and each collection carries some nonzero chance of a match. I'd assume 150,000 to 400,000 dollars per claim, based on published fashion-copyright settlement ranges."
Why this works
Every number has a stated source, not a figure pulled from nowhere.
Stage 5
Give the range
Say it like this
"Low end, one claim at 150,000. High end, four claims at 400,000 each, 1.6 million total. That's the honest spread, not a single number."
Why this works
A single number implies confidence nobody actually has about a legal outcome that hasn't happened yet.
Stage 6
Sanity-check it
Say it like this
"Our current errors and omissions policy caps out at 500,000 dollars. The high end of this estimate blows straight past our own coverage, which means the range isn't academic, it's a real budget conversation."
Why this works
Compares the estimate to something the executive team already has a number for.
Stage 7
Name the direction, and close
Say it like this
"The scraped-photo share of the training corpus is the one number that moves this the most. Cut it, and the whole exposure range comes down with it."
Why this works
Names the single assumption worth attacking, ready for whatever gets pushed on next.
Let's learn
StyleCast reads a customer's photo and a chosen style reference, then renders a mockup of the finished garment, so a customer can see roughly what they're ordering before any fabric is cut.
Before StyleCast shipped, Selvage's design team built training images the fast, free way: they scraped publicly posted runway photography, since it was abundant, current, and nobody stopped them.
Knowledge spark: why does a runway photo have an owner at all?
A runway show happens in public, but the photograph of it doesn't become public property just because thousands of people saw the show. The photographer usually owns the image. The designer often owns rights over the garment's distinctive silhouette too. "Anyone could see it" and "anyone can use it" are two completely different legal facts.
Selvage's design team treated these two as the same thing. They are not.
Now, with StyleCast live, every mockup blends a customer's own photo with patterns the model learned from that scraped corpus, including silhouettes lifted from specific, still-in-copyright designer collections.
More than half the corpus came from the one source with no confirmed right to reuse it this way.
The turn: the exposure isn't really about how many mockups get generated. It's about how many of those mockups land close enough to a specific, identifiable, still-in-copyright design that a designer could reasonably say "that's mine."
Estimated annual copyright exposure: low, typical, high case
The high case alone is more than three times Selvage's current errors and omissions coverage of 500,000 dollars.
The decision I would take back
We assumed a runway photo posted publicly on a fashion blog was safe to scrape and train on, since it was already visible to anyone, and building the corpus that way was faster and cheaper than negotiating licenses. That felt reasonable while the model was an internal prototype nobody outside the company had seen. It stopped being reasonable the moment StyleCast started generating customer-facing mockups that a designer could actually see and recognize.
What I would leave alone: the 10 percent of the corpus made up of customers' own uploaded photos needs no rework at all. Those images came with the customer's own consent for their own garment, and reviewing them the way we're reviewing the runway slice would waste effort on the safest part of the corpus.
The risk was never that thousands of people saw the runway show. The risk was mistaking "everyone could see it" for "anyone could keep it."
The lesson: a copyright exposure estimate isn't a legal exercise bolted onto a finished feature. It's a direct readout of one design decision, made back when the training corpus was built, about which shortcuts felt harmless at the time.
Now here is the same thing as a story
The short version above is what you'd say defending this estimate to Selvage's board. Read this one for how the near miss actually happened.
Coen Vermaas has led Selvage's product team for four years, mostly working on fit prediction and sizing, before StyleCast became his to own as well.
StyleCast launched to good reviews. Customers loved seeing a rendered mockup before committing to a custom order, and the design team kept feeding the model fresh runway photography every season, since it was the easiest way to keep the style references current.
The largest slice of the corpus sat in exactly the corner with the least clarity about who actually owned it.
Then, three months in, a customer service rep forwarded Coen a mockup a customer had shared on social media, tagging it with a designer's name and the caption "wait, did they just copy this dress?"
Side by side, the mockup and the original runway gown were close enough that even a customer noticed without being told to look.
Nobody had built a step to check a generated mockup against recent, identifiable designs before it reached a customer. The pipeline went straight from generation to display, with no gate in between.
Four steps ran the way they were supposed to. The missing fifth step was the one that would have caught this before a customer ever saw it.
Coen pulled the design team together and sorted every training image by where it actually came from, for the first time treating "licensed," "scraped," and "customer-consented" as three separate categories instead of one undifferentiated pile.
Sorting the corpus this way turned an abstract worry into a concrete removal list.
Legal helped write new contract language while the removal list was still being built, so customers ordering during the cleanup would already be covered.
Four plain terms, written before the next mockup shipped, not after a designer's lawyer called.
The old corpus treated any publicly visible photo as fair game the moment it could be downloaded. The new one asks a specific rights question about every image before it ever trains anything.
I approved scraping runway photography because it was the fastest way to keep the style library current, and at the time nothing about the model felt customer-facing enough to matter. It took one tagged social post, from a customer who never even knew there was a legal question at stake, to see that "current" and "cleared" had never been the same requirement.
BOUND, the exposure broken openNot a single guess. BOUND is what forces every dollar in that number to say where it came from.
B
Break it down. The equation.
Expected exposure equals expected valid claims per year, times average settlement cost per claim, plus contract-driven indemnification cost.
States the equation out loud before touching a single number.
O
Own numbers. Each assumption, sourced.
One to four claims a year, drawn from 600 distinct scraped collections. 150,000 to 400,000 dollars per claim, from published settlement ranges.
Every figure has a stated reason, not a number pulled from nowhere.
U
Use a range. Low and high, not one guess.
150,000 dollars low case. 1.6 million dollars high case. A typical case near 650,000.
A single number implies a confidence nobody actually has about an unresolved legal question.
N
Nail the sanity check.
The high case alone is more than three times Selvage's current errors and omissions coverage of 500,000 dollars.
Compares the estimate to a number the executive team already tracks closely.
D
Direction. What swings the estimate most.
The scraped-photo share of the corpus. Cutting it from 55 percent to 15 percent brings expected claims down from about 2.5 a year to about 0.7.
The hardest step, and the one that turns an estimate into an actual plan.
Expected infringement claims per year, as scraped-photo share falls
Nothing about the model changes on this chart, only the corpus behind it. Same feature, same customers, a smaller share of unresolved risk.
The recap, one line per letter: break it down is claims per year times cost per claim, own numbers is one to four claims at 150 to 400 thousand each, use a range is the 150,000 to 1.6 million spread, nail the sanity check is the high case against a 500,000 dollar insurance cap, and direction is the scraped-photo share as the number worth fixing first.
And if you want to be sure it really works, try it somewhere elseSame five letters, a home-renovation visualizer instead of a garment mockup. This time the exposure is measured in complaints, not dollars.
Birchcroft Renovations uses an AI feature that renders a homeowner's kitchen with a new layout, pulled partly from photos of finished projects posted by other contractors and design blogs. Wren Halloway runs product there, and the same estimation question looks different once you're not talking about designer gowns.
Break it down: expected exposure equals expected takedown demands per year from contractors whose finished work appears recognizably in a rendered mockup, times the cost of legal review and response per demand, since litigation here is rare but demand letters are common. Own numbers: an estimated two to six demand letters a year, given how much of the training set comes from contractor portfolio sites with ambiguous reuse terms, at roughly 3,000 to 8,000 dollars in legal review time per letter. Use a range: 6,000 dollars low case, up to 48,000 dollars high case, far smaller in absolute terms than Selvage's exposure, but a real cost against a much thinner renovation-industry margin. Nail the sanity check: the high case is still less than one month of the legal team's retainer, so it's real but not yet a board-level number. Direction: the single biggest lever is tagging which portfolio photos came with an explicit reuse permission versus which were simply scraped from a public site, the same underlying assumption as Selvage's, wearing a different industry's numbers.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "150,000 to 1.6 million a year, and the scraped-photo share is what swings it," and stop.
Cost: there's no budget this quarter to relicense the whole corpus at once. Say so honestly, and start by relicensing just the most recent two seasons, since recent designs are the ones most likely to still be actively enforced.
The model gets better, for real: if StyleCast's rendering quality improves, that doesn't lower the exposure, it raises it. A sharper, more accurate mockup is more likely to look recognizably close to the original design, not less.
Where people run it wrong.
They treat "publicly visible" and "free to use" as the same fact, when they are legally unrelated.
They quote a single exposure number with no range, which hides how much the whole estimate leans on one unverified assumption.
They wait for a claim to arrive before building the similarity check that should have been part of the pipeline from day one.
How to use it live. When asked to describe copyright exposure, don't start with the legal theory. Start with where the training data actually came from, and let the arithmetic follow from that.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "describe the copyright exposure of a generative feature"?
Tap to flip
ANSWER
BOUND: break it down, own numbers, use a range, nail the sanity check, direction. Direction names which single assumption swings the estimate most.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Coen Vermaas, the product lead at Selvage Studio who had to put a real number on StyleCast's copyright exposure for the executive team.
3 · THE EQUATION
What's the equation behind this exposure estimate?
Tap to flip
ANSWER
Expected exposure equals expected valid claims per year, times average settlement cost per claim, plus contract-driven indemnification cost.
4 · THE BIGGEST LEVER
Which single assumption swings the exposure estimate the most?
Tap to flip
ANSWER
The share of the training corpus that's scraped runway photography instead of properly licensed, 55 percent at the time of the estimate.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Scraping publicly posted runway photos for the training corpus because they were visible and free to download, treating "visible" as the same thing as "free to use."
6 · THE NUMBER
Fill in the blank: expected annual exposure runs from about 150,000 dollars low case to ___ dollars high case.
Tap to flip
ANSWER
1.6 million dollars. More than three times Selvage's own errors and omissions insurance cap of 500,000.
7 · THE REPLAY
Same kind of near-identical mockup, redesigned pipeline. What changes?
Tap to flip
ANSWER
A similarity check catches it before it reaches a customer, instead of a customer's own social post surfacing it after the fact.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the exposure measured in there?
Tap to flip
ANSWER
Birchcroft Renovations' kitchen visualizer. There, exposure is measured in contractor demand letters, 6,000 to 48,000 dollars a year, not designer litigation.
Check yourself Score: 0 / 0
Multiple choice
1. Why does a photo of a public runway show still carry copyright risk for StyleCast?
A. Because runway shows are always ticketed events.
B. Because being visible to a public audience doesn't make the photograph or the design free to reuse.
C. Because only customer-uploaded photos carry any legal risk.
D. Because fashion designs can never be copyrighted.
Show hint
Look at the metaphor scene contrasting "on the runway" with "in the public domain."
Show answer
B. Public visibility and legal permission to reuse are two separate facts. The photographer and often the designer retain rights regardless of how many people saw the show.
True or false
2. True or false: this answer recommends removing every scraped runway photo from the training corpus immediately.
True
False
Show hint
Look at the decision tree for which training images need removal.
Show answer
False. Only scraped images with no credited artist and no rights confirmation get removed outright. Credited scraped images get flagged for review first, not automatic deletion.
Fill in the blank
3. Fill in the blank: Selvage's current errors and omissions insurance policy caps out at ___ dollars.
Show hint
Look at the "nail the sanity check" step.
Show answer
500,000 dollars. The high-case exposure estimate of 1.6 million is more than three times that coverage.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Scraping publicly posted runway photography to build the training corpus. It made sense while the model was an internal prototype nobody outside the company had seen.
Short answer, where it wouldn't matter
5. Name a slice of Selvage's training corpus where this copyright concern genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The customer-uploaded photo slice, about 10 percent of the corpus. Those images already carry the customer's own consent for their own garment.
Short answer, apply it yourself
6. Pick an AI tool you know of that generates images or text from something scraped off the internet. What's one assumption about that training data you'd want to check before trusting the tool's copyright position?
Show hint
Think about whether "publicly available" and "licensed for this use" have actually been checked as two separate questions.
Show answer
Model answer: Most people land on some version of "was this actually licensed, or just visible," the same distinction that sat at the center of Selvage's exposure estimate.
Before you close the answer
Why this works
Tests whether you can turn a vague legal worry into a real, sourced estimate with a range, and whether you can name the one assumption that would actually move the number instead of gesturing at "some risk."
Follow-up traps
"Couldn't you just stop using runway photos entirely, right now?" Response: possible, but a sudden full removal would gut the style library customers currently rely on. Relicensing the most recent, most exposed seasons first gets most of the risk reduction without breaking the feature overnight.
"Isn't a similarity check just a band-aid over bad training data?" Response: it's a genuine second layer, not a substitute for fixing the corpus. Even a fully licensed corpus can still generate something that looks too close to a specific design by coincidence, so the check stays even after the corpus is cleaned up.
If pressed
Selvage's real settlement-range assumption came from three previously reported fashion-industry copyright settlements between 2019 and 2023, not a single data point, since one past case would be too thin a base to build a whole estimate on.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.