ConceptFoundationalDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #7
What visual patterns signal that content was AI-generated?
SPARK design the signal before the tool gets good enough to hide behind it
Aldermarsh Community College's bridge-year writing program uses Penmark, an AI tool that drafts margin comments on student essays for instructors to review before sending. Thandiwe Marechera is the program's Writing Director. She has taught composition for fourteen years.
The direct answer
Don't mark that a model was involved once. Mark its current state, all the time: a comment sits in one of three visible states, unreviewed AI draft, human-edited AI draft, or fully human, each with its own color and weight, and the state only changes when a real edit event happens. A one-time watermark tells a reader what made the words. A state tag tells them whether anyone has actually looked yet.
Design it in this order
Show a persistent state tag, not a one-time badge.Why: a badge that appears once and disappears can't tell a reader anything about right now, only about the moment of creation.
Tie the tag's color and weight to whether a person has actually changed the text.Why: a label that says "AI-assisted" forever, even after a teacher rewrites half of it, stops meaning anything within a week.
Keep a visible edit trail behind the tag, not just a static word.Why: "reviewed" is a claim; a diff a reader can open is proof.
Never let an untouched draft render in the same weight as a signed final.Why: identical rendering is what let an unread comment pass as a read one in the first place.
Skip the tag entirely on fully human work.Why: tagging everything, even work with no model in it, trains readers to ignore the tag.
How to answer this, stage by stage
Nobody is grading whether you can name five UI patterns off the top of your head. They're grading whether the pattern you pick can survive someone trying to lean on it wrongly.
Stage 1
Pin it to one real screen
Say it like this
"I'll answer this for one real case: Penmark, an AI tool that drafts margin comments on student essays, and Thandiwe Marechera, a writing director reviewing them before they go out."
Why this works
Stops the answer from turning into a list of patterns with no product behind any of them.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how the work gets done today. Payoff, the habit I want the design to build. Anchor, the one design decision that answers this. Risk, what breaks the first time it's wrong. Keep out, what I won't build yet."
Why this works
Shows the interviewer you have a method for designing a signal, not just a memorized list of icons.
Stage 3
Reframe what the question is really asking
Say it like this
"This isn't really 'how do I mark that a model touched it.' It's 'how do I keep showing whether a person has actually looked,' because that fact changes by the hour and a one-time badge can't keep up."
Why this works
Separates a strong answer from a list of static icons a student memorized.
Stage 4
Give the anchor, before any reasoning
Say it like this
"Every AI-touched comment carries a state tag, not a one-time badge: unreviewed draft in amber, human-edited in full color, plus a diff a reader can open to see exactly what changed."
Why this works
This is the direct answer, and it names an actual design, not a category of design.
Stage 5
Prove it against the compressed near miss
Say it like this
"A student teacher shadowing Thandiwe once picked two comments at random and couldn't tell which one she'd actually read. When they checked, forty of her last sixty comments had gone out unread, sitting in the exact same box as the ones she'd rewritten by hand."
Why this works
Turns "add a visual signal" into a concrete failure the anchor was built to stop.
Stage 6
Say what you're deliberately not building yet
Say it like this
"I wouldn't tag a spelling correction or a comma fix. Nobody assumes a red squiggly underline required a person's judgment, so tagging it just trains readers to stop reading tags at all."
Why this works
Shows judgment about where the pattern matters instead of a rule applied everywhere out of caution.
Stage 7
Close on the one line
Say it like this
"Don't signal what made it. Signal whether anyone has actually looked yet, and keep that signal alive as long as the comment exists, not just at the moment it was written."
Why this works
Restates the anchor in one breath, ready for a live follow-up.
Let's learn
Here's what a real classroom looks like once an AI drafting tool has been running quietly for a semester.
Before Penmark, Thandiwe hand-wrote margin comments on roughly 150 essays a semester, about 12 minutes each, close to 30 hours of writing feedback by hand. Penmark drafts a first-pass comment in under 10 seconds. Once she started reviewing and editing those drafts instead of writing from scratch, her time per essay fell to about 4 minutes.
Five ways a screen can show its work. Penmark, at launch, used none of them.
Here's the turn: the extra minutes Penmark saved were never the real story. The real problem showed up the moment an unread AI draft and a hand-rewritten comment started rendering in the exact same box, same font, same weight, with nothing on the page to tell them apart.
Students who could correctly tell an AI-drafted comment from Thandiwe's own, before and after the state tag
Before the tag, most students were guessing. After it shipped, nearly all of them could tell a reviewed comment from an unreviewed one on sight.
A pencil draft and an inked page hold the same words sometimes. Only one of them should look finished.
At its worst, a comment box that can't tell a reader whether a person looked at it doesn't just waste a parent's trust once. It quietly tells every student in the program that the personal attention they were promised might not have happened for any of them, and there's no way to check.
The decision I would take back
When Penmark launched, the team decided any AI-drafted comment would render in the exact same font, size, and color as Thandiwe's own, "to keep the page looking clean and like one voice." That made sense when adoption was optional and only a handful of comments a week came from the model. It stopped making sense once most comments did.
What I would leave alone: a spelling fix or a comma correction doesn't need this same tagging treatment. Nobody assumes a red squiggly underline required a person's judgment, and tagging every small mechanical fix would just teach readers to stop noticing tags where they actually matter.
The lesson: a comment box that looks the same whether a person or a model made it isn't simpler. It's a hidden decision about who gets the credit, made without anyone choosing it on purpose.
Now here is the same thing as a story
The short version above is what you'd say defending this design to Aldermarsh's department chair. Read this one for how close the actual near miss came.
For fourteen years, Thandiwe's margin comments were the thing students kept. Not graded feedback, real sentences, written in her own hand in the gap beside a paragraph that needed one. She could tell a genuinely confused essay from a merely lazy one by the second page.
Penmark arrived in August, and the first semester was good. She'd open a draft comment, read it fully, rewrite half of it in her own words, and send it. The tool saved her real time. Nothing about the page looked any different than it always had.
Knowledge spark: why do AI writing-feedback drafts even need review?
A model drafting margin comments hasn't read the whole semester the way a teacher has. It can miss that a student always struggles with the same kind of sentence, or praise a paragraph that quietly repeats last week's mistake. The draft is a starting point, not a finished thought, which is exactly why whether someone reviewed it matters.
By October, she was reading every draft fully. By January, she skimmed the ones that looked reasonable at a glance. By March, she'd taken to approving a batch of ten or twelve at a time between classes, trusting that Penmark's drafts were, by now, usually fine.
Both panels used to render as one identical box. The redesign is what finally tells them apart.
The trigger was small. A student teacher named Mateo, shadowing her for the week, pulled up two essays side by side and asked which comment was hers and which was Penmark's draft. She looked for a long moment and admitted she genuinely couldn't say without opening her own memory of writing it.
She didn't lose the essays. She lost the ability to tell her own words from the tool's, because the page had never asked her to keep them apart.
That evening she pulled the last sixty comments she'd sent and checked each one against her own edit history. Forty of them had gone out with zero changes from Penmark's first draft, unread in any real sense, sitting in the exact same box as comments she'd rewritten from scratch.
Six good months hid a habit that had already changed. The question in week 25 is what made it visible.
Aldermarsh didn't pull Penmark. They rebuilt the comment box: a small amber tag for anything Penmark drafted and Thandiwe hadn't yet touched, switching to full color the moment she made a real edit, with a one-click diff showing exactly what changed. Batch-approving without opening a comment now leaves it visibly amber, forever, until someone actually looks.
SPARK, in one screenNot a UI style guide. SPARK is what decides which signal earns a place on the page at all.
S
Situation. How the work gets done today, without the design decision yet.
Thandiwe hand-writing every margin comment, 12 minutes an essay, 150 essays a semester.
Grounds the anchor in a real workflow instead of an abstract feature list.
P
Payoff. The habit the design should build.
Students trusting that a marked comment reflects a real person's actual judgment, not just a plausible sentence.
Names what the design is actually for, past just saving time.
A
Anchor. The one design decision everything else hangs on.
A persistent state tag, amber until a real edit happens, full color after, with an openable diff. This is the answer to the question.
The hardest step, and the one a reader should be able to point at on the actual screen.
R
Risk. What breaks the first time the design is wrong.
A teacher batch-approves ten comments without reading them, same as before. Now those ten sit visibly amber instead of blending in.
Proves the anchor survives the exact failure that almost happened.
Four small facts turn a vague tag into something a reader can actually check.
K
Keep out. What stays off the page on day one.
No tag on spelling or punctuation fixes. Tagging mechanical corrections trains readers to stop noticing the tag where it counts.
Shows restraint instead of a rule applied everywhere out of caution.
Three states, three signals. The branch on the right is the one worth protecting: no tag where none is owed.
The recap, one line per letter: situation is Thandiwe writing every comment by hand, payoff is students trusting that a marked comment reflects real judgment, anchor is the persistent amber-to-full-color state tag with an openable diff, risk is a batch-approved comment staying visibly unreviewed instead of blending in, and keep out is leaving mechanical fixes untagged so the tag still means something where it matters.
And if you want to be sure it really works, try it somewhere elseSame five letters, a logistics company's customer support inbox instead of a classroom. A different anchor risk breaks the second story.
Ferro Logistics runs customer support for freight delay claims. Kenji Osei, Support Operations Lead, oversees a team of twelve agents using ReplyDraft, an AI tool that drafts replies to delay complaints for an agent to send or edit. Mapped onto SPARK: situation is agents typing every reply by hand, about 6 minutes each, across roughly 300 tickets a day; payoff is customers trusting that a reply actually addresses their specific shipment, not a template; anchor is the same persistent state tag idea, an amber marker on any reply sent without an agent edit, full color once an agent has changed at least one sentence.
The risk here is different from Thandiwe's. It isn't a teacher losing track of her own hand. It's a customer, already annoyed about a late shipment, receiving a reply that reads as generic and assuming nobody read their complaint at all, whether or not that's true. The old decision being taken back: ReplyDraft's replies were designed to match the company's support voice exactly, on purpose, so customers wouldn't feel like they were talking to a bot. That made sense when reply volume was low enough that agents edited nearly everything. It stopped making sense once volume tripled and unedited sends became the majority.
The same five signals apply to a support reply as to a margin comment. Only the stakes changed.
Weekly "did a person actually read this" follow-up tickets, before and after the state tag shipped
The drop starts right at the tag's launch line, not before it. That's what makes it the tag's effect and not a seasonal dip.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a persistent state tag, amber until a person edits it, full color after, with a diff behind it," and stop.
Cost: there's no engineering time this quarter for a diff view. Say so honestly, and ship the two-state color tag alone first, since even a plain amber-or-not signal beats no signal at all.
The model gets better, for real: if Penmark's drafts get good enough that teachers barely need to edit them, that's still not a reason to remove the tag. It's a reason to watch the unreviewed-and-sent rate even more closely, since a better model makes it easier to stop looking at all.
Where people run it wrong.
They ship a one-time watermark that says "AI-assisted" and call the trust problem solved.
They tag everything, including mechanical fixes, until readers learn to ignore every tag on the page.
They let the tag's color depend on the model's own confidence score instead of on whether a person actually made a change.
How to use it live. When someone asks what visual pattern signals AI-generated content, ask yourself first: does this signal answer "what made it" or "has anyone looked yet"? Only the second question is the one worth designing for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "design a feature for X" question like this one?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. It designs the anchor against the risk before building it, instead of running FLIPS backward.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Thandiwe Marechera, Writing Director at Aldermarsh Community College, who has taught composition for fourteen years.
3 · THE PAYOFF
What habit should the design build in the reader?
Tap to flip
ANSWER
Trusting that a marked comment reflects a real person's actual judgment, not just a plausible-sounding sentence.
4 · THE ANCHOR
What's the one concrete design decision this answer commits to?
Tap to flip
ANSWER
A persistent state tag: amber until a person makes a real edit, full color after, with an openable diff showing what changed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Rendering any AI-drafted comment in the exact same font, size, and color as Thandiwe's own, to keep the page looking like one voice.
6 · THE NUMBER
Fill in the blank: when Thandiwe checked her last 60 sent comments, ___ of them had gone out with zero edits from Penmark's draft.
Tap to flip
ANSWER
40 out of 60. Two out of three comments had gone out unread in any real sense, indistinguishable on screen from the ones she'd rewritten.
7 · THE REPLAY
Same batch-approval habit, new tag design. What changes?
Tap to flip
ANSWER
The unread comments still get approved in a batch, but they now stay visibly amber forever until someone actually opens and edits them, instead of blending in as final.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's different about its risk?
Tap to flip
ANSWER
Ferro Logistics' ReplyDraft customer support tool. There, the risk isn't a teacher losing track of her own hand, it's a customer assuming nobody read their complaint at all, whether or not that's true.
Check yourself Score: 0 / 0
Multiple choice
1. Why does a persistent state tag work better here than a one-time "AI-generated" watermark?
A. Watermarks are harder to build than state tags.
B. Whether a person has reviewed the text changes after the moment of creation, and a one-time badge can't reflect that.
C. State tags are required by most writing platforms.
D. Students prefer colorful tags to plain text badges.
Show hint
Look at the direct answer.
Show answer
B. A comment can be edited after it's created, so only a live state, not a one-time stamp, can show whether a person actually looked.
True or false
2. True or false: this answer recommends tagging every correction Penmark makes, including spelling and punctuation fixes.
True
False
Show hint
Look at "keep out" and "what I would leave alone."
Show answer
False. Mechanical fixes stay untagged on purpose, so the tag keeps meaning something where a real judgment call was made.
Fill in the blank
3. Fill in the blank: after Ferro Logistics shipped its own state tag in week 3, weekly "did a person actually read this" tickets fell from 38 to about ___ by week 8.
Show hint
Look at the line chart in Section 4.
Show answer
9 tickets a week. The drop begins right at the tag's launch line, which is what makes it the tag's effect.
Short answer, apply it yourself
4. Think of a tool you use that mixes AI-drafted and human-written text in the same space. Can you actually tell which is which, and how?
Show hint
Ask whether the signal you're using tells you "what made it" or "has anyone looked yet."
Show answer
Model answer: Most people can't reliably tell, which is exactly the gap a persistent state tag is built to close.
Short answer, why no middle setting
5. Why wouldn't a single, one-time "AI-assisted" label placed at the top of the page solve this problem?
Show hint
Look at the reframe in Stage 3.
Show answer
Model answer: Because the real question isn't what made the text, it's whether a person has reviewed it since, and that fact changes after the page loads.
Short answer, where it wouldn't matter
6. Name a kind of correction on this same page where this state-tag treatment genuinely isn't needed.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Spelling and punctuation corrections. Nobody assumes those required a person's judgment, so tagging them just adds noise.
Before you close the answer
Why this works
Tests whether you can design a signal that survives changing over time, not just describe a static icon, and whether you know the difference between marking what made something and marking whether anyone has checked it since.
Follow-up traps
"What if a teacher edits one word just to change the tag's color?" Response: set a minimum real-edit threshold, like a changed sentence rather than a changed character, so the tag can't be gamed with a single keystroke.
"Doesn't this just add more UI clutter to an already busy screen?" Response: it replaces an existing, misleading signal, an identical-looking box, rather than adding a new one; the clutter risk is real but it's a tradeoff against a screen that currently lies.
If pressed
The redesigned Penmark also logs every state change with a timestamp, not just the current state, so a department chair auditing a complaint can see exactly when a comment moved from amber to reviewed, not just where it sits today.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.