CaseAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #5

Describe how you would turn user edits into a quality signal.

SPARK the product is Larkspell Assist, which drafts a freelancer's outreach message for a job post

Larkspell is a freelance marketplace. Larkspell Assist drafts a personalized outreach message the moment a freelancer applies to a job post. Rasheed Iqbal manages trust and quality for Assist, and checks its numbers from his phone most mornings.

The direct answer
Don't just track whether a draft got edited. Track what changed, broken into which part, the opening line, the skill mentioned, the price, the call to action, and how much. Then weight an edit pattern that repeats across many different freelancers as a real quality signal, and treat one freelancer's habitual personal tweak as noise, not proof the draft was flawed.
Do this, in order
  1. Break the edit signal into structured parts, not one yes-or-no flag.Why: "23% of drafts get edited" can't tell you what to actually fix.
  2. Weight edits that repeat across many different freelancers as the real signal.Why: a pattern across strangers usually points at the draft. One person's habit usually doesn't.
  3. Route high-severity, high-repeat patterns into a weekly human review before touching the drafting prompt.Why: an automatic pipeline reacting to every edit risks chasing noise or drifting toward one vocal freelancer's taste.
  4. Measure edit severity by what changed, not just how many characters moved.Why: a two-word tweak and a full rewrite can look identical on a naive character count.
  5. Never score or flag individual freelancers by how often they edit.Why: editing often could just mean someone cares about their applications, not that the draft failed.
  6. Leave the actual sending flow untouched.Why: this is about what signal an edit produces, not how or when a freelancer sends a message.

How to answer this, stage by stage

Nobody is grading whether you'd log edits at all. They're grading whether you can tell a real pattern apart from one person's habit.

Stage 1
Scope it to one real feature
Say it like this
"I'll design this for Larkspell Assist, which drafts a personalized outreach message when a freelancer applies to a job."
Why this works
Turns an open design prompt into one concrete feature the interviewer can push on.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how edits get lost today. Payoff, the habit I want. Anchor, the actual design. Risk, what breaks it. Keep out, what I won't build yet."
Why this works
Signals a method before a single implementation detail gets said.
Stage 3
Ground it in today, without the signal
Say it like this
"Right now, a freelancer edits the AI draft before sending it, or doesn't, and that edit just vanishes. Nobody on the Assist team ever sees what changed or why."
Why this works
Shows the gap is real and current, not a hypothetical.
Stage 4
Give the anchor
Say it like this
"Break every edit into structured parts, which section changed and how much, then weight a pattern that repeats across many different freelancers far higher than one freelancer's personal tweak repeated every time."
Why this works
This is the direct answer, stated as a concrete, arguable design decision.
Stage 5
Prove the anchor survives its own risk
Say it like this
"The first time one freelancer edits every single draft the same personal way, the anchor doesn't chase it as a quality signal, since it only counts a pattern once it shows up across a spread of different freelancers, not one person's habit."
Why this works
Answers the real follow-up: what happens the one time the naive version of this idea would go wrong.
Stage 6
Close on the one line
Say it like this
"An edit isn't a vote of no confidence. It's data, but only once you can tell a personal habit apart from a real pattern."
Why this works
Restates the direct answer in one breath, ready for whatever gets pushed on next.

Let's learn

The draft message sits on a phone screen, one thumb away from being sent exactly as the AI wrote it, or rewritten first.

Larkspell is a freelance marketplace, and Larkspell Assist drafts a personalized outreach message the moment a freelancer applies to a job post.

The decision I would take back When we shipped Assist, we decided not to log the pre-edit draft alongside the sent message at all, since it felt like extra storage and complexity for a feature still in beta. That made sense when Assist handled a few hundred applications a day and the team could eyeball a handful by hand. It stopped making sense once tens of thousands of edits a day were happening completely unrecorded.

Before Assist, a freelancer spent close to twenty minutes writing a pitch to a job post by hand. Assist drops a full draft in under ten seconds, and a freelancer can send it as-is or edit first.

Share of Assist drafts edited, by which part changed
50% 25% 0% 41% Opening line 34% Skill mentioned 12% Call to action 9% Price or rate
A yes-or-no edit flag would have blended all four of these into one meaningless number. Broken apart, "skill mentioned" is the one worth chasing.

At its worst: every edit gets treated as a vote against the draft, fed straight into a model that retrains in real time, and the drafting tool quietly overfits to a handful of the most active freelancers' personal writing style, while the real skill-mismatch problem, the one hurting many different freelancers, gets diluted among thousands of purely cosmetic tweaks and never gets fixed at all.

What I would leave alone: a freelancer who's used Assist for months and clearly has a signature style, a specific greeting, a specific sign-off, doesn't need their own repeated personal edit feeding into the shared quality signal at all. That's their voice, not a flaw in the model.

An edit was never proof the draft was wrong. It only becomes proof once the same fix shows up across strangers who never talked to each other.

The lesson: the useful part of an edit was never the fact that it happened. It's whether a stranger, somewhere else, made the exact same fix without ever seeing the first person's draft.

Now here is the same thing as a story

The short version above is what you'd say defending this design to Larkspell's product council. Read this one for how the pattern actually got found.

For months, Larkspell Assist looked like a clean win. For months after that, it looked like nothing, because nothing about it was ever actually measured.

Hand sketched flow diagram titled Today, without Assist. Four boxes: freelancer reads job post, writes pitch by hand highlighted, 20 minutes gone, sends application.
The second box used to take twenty minutes. Assist collapsed it to seconds, and nobody watched what happened after.

Rasheed Iqbal checks Assist's numbers most mornings from his phone, thumbing through a dashboard during his commute before he's even sat down at his desk. Adoption climbed fast. Freelancers applied to more jobs per week, since drafting no longer cost them real time.

There was no single bad day. Over months, win-rate for Assist-drafted proposals, whether a freelancer actually got hired, sat a few points below hand-written ones, without ever cratering hard enough to trip an alarm. It was just persistently, quietly worse, and nobody could say why, because the edited text itself was never kept anywhere.

Knowledge spark: why does a raw edit count fail as a quality signal on its own? A raw count can't tell a systemic gap in the draft from one person's personal habit. If one freelancer edits every draft the same way, out of preference, and that gets counted the same as ninety strangers independently fixing the same real mistake, the two look identical on a chart that only counts "edited: yes."

Rasheed started informally asking a few freelancers what they usually changed. Several, without prompting, mentioned the same thing: "I always fix the skill Assist highlights, it's usually not the one that actually matters for this job."

Hand sketched comparison diagram titled The day it's wrong. Left panel, a person icon labeled One freelancer, caption same tweak every time. Right panel, a circle icon labeled Ninety freelancers, caption same fix converging.
One person's habit and ninety strangers converging on the same fix look nothing alike, once you're actually counting who they are.

That offhand pattern became the seed of a structured edit-capture system: every edit gets tagged by section, and the tool tracks how many distinct freelancers made the same kind of fix, not just how many times it happened.

Distinct freelancers making the "wrong skill highlighted" edit, week by week
100 50 0 20 freelancers: systemic threshold Wk1: 4 Wk2: 11 Wk3: 26 Wk4: 58 Wk5: 94
By week three, enough strangers had made the identical fix that it stopped looking like anyone's personal taste.
Hand sketched labeled parts diagram titled The anchor, close up. Center document icon labeled Edited draft, with four callouts: opening line changed, skill mentioned changed, price changed, call to action changed.
Four parts, and only one of them, skill mentioned, turned out to be the systemic problem.

The pattern traced back to how Assist parsed job posts with several required skills listed: it grabbed the first skill mentioned in the post, not the one most relevant to that specific freelancer's profile. Reordering how the retrieval step ranked skills closed most of the gap within a few weeks, and win-rate for Assist-drafted proposals caught up to hand-written ones by the following month.

The old design asked win-rate alone to explain itself, with no view into what freelancers were actually fixing. The new one asks the edit itself what it means, by checking whether it repeats across people who've never spoken to each other.

I didn't build edit tracking into version one because it felt like a nice-to-have we could add anytime. It took months of a quietly underperforming number, with no single bad day I could point to, to see that anytime never has a deadline unless something forces it to.

SPARK, and the edit that finally countedNot a debate over whether to log edits. SPARK is what tells you which edits are actually worth trusting.

S
Situation. How this happens today, without the signal.
A freelancer edits the AI draft, or doesn't, and that edit vanishes. Nobody on the Assist team ever sees what changed or why.
Grounds the design in a real, current gap, not a hypothetical improvement.
P
Payoff. The habit this should build.
Treating a pattern that repeats across many strangers as a real signal, while ignoring one person's habitual personal tweak.
Names the real behavior change the design is trying to produce, not just "collect more data."
A
Anchor. The one decision everything hangs on.
Structured, section-level edit capture, weighted by how many distinct freelancers make the same fix, not by raw edit count.
This is the hardest step and the direct answer: spread across people, not volume, is what makes an edit a signal.
R
Risk. What breaks the first time it's wrong.
One freelancer editing every draft the same personal way. The anchor survives it, since it only counts a pattern once it spreads across different people.
Proves the anchor was designed against its most likely failure, not just its ideal case.
K
Keep out. What we won't build, day one.
No real-time auto-retraining on every edit, no scoring individual freelancers by edit rate, no fully automated prompt rewrite without human review.
Shows judgment about scope, not a wish list of everything the signal could eventually power.
Hand sketched quadrant titled Sorting edits by spread and size. Axes how many freelancers from few to many, and how big the change from small wording to full rewrite. Personal sign-off tweak sits bottom left, few and small. Skill-mismatch fix sits top right, many and medium. Random typo fix sits bottom left near sign-off. Full pitch rewrite sits top left, few and large.
Only the top right corner, spread across many freelancers and a real structural change, is worth a weekly review.
Hand sketched icon list titled What we won't build, day one. Three items: a question mark box icon labeled No real-time auto-retrain on edits, a box icon labeled No per-freelancer edit scoring, a document icon labeled No automated prompt rewrite unreviewed.
Each of these is a real feature the signal could eventually power. None of them is safe to ship on day one.

The recap, one line per letter: situation is an edit that vanishes with no record, payoff is trusting a cross-freelancer pattern instead of one person's habit, anchor is section-level edit capture weighted by spread, risk is the single freelancer's personal tweak the anchor has to ignore, and keep out is holding off on automatic retraining and individual scoring.

And if you want to be sure it really works, try it somewhere elseSame five letters, a public library's catalog records instead of a freelance pitch. A completely different field, the same anchor.

Bellhaven Public Library uses an AI tool to draft the catalog description for a newly acquired book, before a cataloger reviews and publishes it to the library's public search system. Callista Wren is the cataloging systems librarian who owns that tool.

Mapped onto SPARK: situation is a cataloger today, editing the AI-drafted record silently before publishing, with no record kept of what changed. Payoff is trusting a pattern that repeats across many different catalogers and many different titles, instead of one cataloger's personal formatting preference. The anchor is structurally the same idea, aimed at library records instead of pitches: capture edits by section, subject headings, summary, age rating, and weight a fix by how many distinct catalogers make it across how many distinct titles, not by raw edit volume. Risk is a veteran cataloger who always reformats the summary a certain way out of personal style, the anchor ignores it because it never spreads to anyone else's records. Keep out is no automatic retraining in real time, no scoring individual catalogers, and no automated changes to controlled-vocabulary subject headings without a professional librarian's review, since those follow strict standards no model should touch unsupervised.

Hand sketched decision tree titled Is this catalog edit a real signal? Root, cataloger edits an AI draft record. Branches: one cataloger same style always leads to personal ignore, many catalogers converging leads to systemic flag for review, large structural rewrite leads to flag regardless of spread, typo fix only leads to ignore.
Two of the four branches lead to "ignore." That's by design, not a gap in the signal.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "capture what changed in an edit, then trust the pattern only once it repeats across strangers, not one person's habit," and stop.
Cost: there's no time to build the full structured capture this sprint. Say so honestly, and start by just logging the pre-edit draft alongside the sent message, since even that alone beats losing every edit completely.
The model gets better, for real: if Assist's overall drafting quality improves, that's still not a reason to drop edit tracking, a better average model can still carry one narrow, systemic gap that only shows up once enough strangers hit it.

Where people run it wrong.
They treat every edit as equally meaningful, instead of separating personal style from a real structural gap.
They wire raw edit volume straight into automatic retraining, risking a model that drifts toward whoever edits the most.
They score individual users by how often they edit, punishing engagement instead of measuring the draft itself.

How to use it live. When someone asks how to turn edits into a signal, ask yourself one question first: would a stranger, with no knowledge of the first person's edit, make the exact same fix. If the answer is yes across enough people, you've found a real signal. If not, you've found someone's personal style.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "describe how you'd turn user edits into a quality signal"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. The anchor here is a section-level edit capture weighted by spread across users.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Rasheed Iqbal, who manages trust and quality for Larkspell Assist and checks its numbers from his phone most mornings.
3 · THE SITUATION
How does this happen today, without the signal?
Tap to flip
ANSWER
A freelancer edits the AI draft, or doesn't, and the edit vanishes. Nobody on the Assist team ever sees what changed or why.
4 · THE ANCHOR
What makes an edit count as a real quality signal here?
Tap to flip
ANSWER
The same fix repeating across many different freelancers who've never spoken to each other, not just the raw fact that a draft got edited.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping Assist without logging the pre-edit draft at all, since it made sense only while volume was small enough to eyeball by hand.
6 · THE NUMBER
Fill in the blank: by week five, ___ distinct freelancers had made the same "wrong skill highlighted" edit.
Tap to flip
ANSWER
94 distinct freelancers. It crossed the team's 20-freelancer systemic threshold between week two and week three.
7 · THE REPLAY
Same skill-mismatch pattern, redesigned system. What changes?
Tap to flip
ANSWER
The structured signal flags the pattern automatically once it crosses the freelancer-spread threshold, and win-rate closes the gap with hand-written pitches within weeks instead of staying quietly worse indefinitely.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the anchor there?
Tap to flip
ANSWER
Bellhaven Public Library's AI catalog drafts. Same anchor: section-level edit capture weighted by spread across catalogers and titles, not raw edit count.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the "skill mentioned" section changed in ___% of edited Assist drafts, the second most common edit after the opening line.
Show hint
Look at the bar chart of edits by section.
Show answer
34 percent. It turned out to be the one category hiding a real systemic problem, not just a stylistic preference.
Multiple choice
2. Why does the anchor weight edits by how many distinct freelancers make the same fix, instead of by raw edit volume?
  • A. Because raw edit volume is too expensive to calculate.
  • B. Because one freelancer editing every draft the same way, out of personal habit, would otherwise look identical to a real, widespread problem.
  • C. Because freelancers are only allowed to edit a draft once.
  • D. Because distinct-freelancer counts are required for legal compliance.
Show hint
Look at "the day it's wrong" comparison.
Show answer
B. Spread across strangers is what separates a real pattern from one person's taste. Raw volume alone can't tell the two apart.
True or false
3. True or false: a freelancer who edits nearly every draft they receive should be flagged as someone Assist is failing.
  • True
  • False
Show hint
Look at the priority list's fifth item.
Show answer
False. Editing often could just mean someone cares about their applications. Individual freelancers are never scored or flagged by their own edit rate.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Shipping Assist without logging pre-edit drafts, since it felt like unneeded complexity for a small beta. It stopped making sense once tens of thousands of edits a day were happening with no record at all.
Short answer, where it wouldn't matter
5. Name a freelancer on Larkspell whose repeated edit shouldn't feed the shared quality signal.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A freelancer with months of history and a clear signature style, a specific greeting or sign-off they always add. That's their voice, not a flaw in the draft.
Short answer, apply it yourself
6. Pick a product you use yourself. What's one edit you always make to something it generates or suggests for you?
Show hint
Think of an autocomplete suggestion, a template, or a suggested reply you always rewrite the same way.
Show answer
Model answer: Many people always rewrite a suggested reply's greeting, or always correct the same kind of autocomplete guess, without ever reporting it as a bug.
Before you close the answer
Why this works
Tests whether you can turn a messy, everyday behavior, editing a draft, into a real signal without drowning it in personal style noise or individual-level scoring risks.
Follow-up traps
"Couldn't you just ask freelancers to rate the draft instead of inferring quality from edits?" Response: almost nobody bothers to rate, while nearly everyone who has a problem edits the draft one way or another, so edits are the far more complete signal.

"What if two different freelancers make similar edits for completely unrelated reasons?" Response: that's exactly why edits are tagged by structured category and reviewed by a person weekly, not auto-applied, a human check catches a coincidence a raw count wouldn't.
If pressed
Larkspell's real severity score also weighs how semantically different the edited text is from the original, not just which section changed, since a full rewrite of the opening line matters more than swapping one adjective in it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more