ConceptIntermediateDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #3

What implicit signals tell you an output was bad?

LEAD the product is Tallowmark's SnapList, which drafts a resale listing straight from a seller's photos

Tallowmark is a resale marketplace where sellers list secondhand furniture, electronics, and clothing. SnapList drafts each listing's description straight from the seller's uploaded photos. Beckett Osei is the trust and quality PM who watches what SnapList ships.

The direct answer
Watch the rate of buyers messaging a seller with a question the listing should have already answered. That number climbs two to three weeks before returns citing "not as described" ever move, because it catches a vague or incomplete AI description while the item is still sitting in someone's cart, not after it's already shipped back.
Do this, in order
  1. Track buyer pre-purchase questions as the leading signal, not the return rate.Why: it moves weeks before returns do, while the fix is still cheap.
  2. Pair the question rate with the lagging return rate, so you can prove the leading signal actually predicts real cost.Why: a leading indicator nobody's checked against reality is just a hunch with a chart.
  3. Watch for sellers gaming question rate down with longer, vaguer text that just discourages asking.Why: a metric that can be satisfied without fixing the real problem isn't measuring the real problem anymore.
  4. Set a real threshold: hold an AI draft for seller review once its question rate clears the hand-written baseline.Why: a signal nobody acts on is a dashboard decoration, not a quality system.
  5. Watch photo re-zoom and return reason codes as secondary tells, not replacements.Why: they add detail, but question rate is still the one that moves first.
  6. Leave the photo-to-text drafting model itself alone.Why: this is about which signal you watch, not how the description gets written.

How to answer this, stage by stage

Nobody is grading whether you can list implicit signals. They're grading whether you picked the one that actually arrives early.

Stage 1
Scope it to one real system
Say it like this
"I'll answer this for Tallowmark's SnapList, which drafts a resale listing's description straight from a seller's photos."
Why this works
Grounds a general question about "bad output" in one concrete drafting system.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the outcome that actually matters. Early signal, what moves before it. Abuse, how the signal gets gamed. Decision, what I'd do at each threshold."
Why this works
Signals a method before naming any specific signal, so it doesn't sound like a grab-bag list.
Stage 3
Name the real outcome
Say it like this
"What actually matters isn't whether a description sounds polished, it's whether the buyer keeps the item without a fight."
Why this works
Without this, any signal you pick is just a guess at what "bad" means.
Stage 4
Give the early signal
Say it like this
"The tell is the rate of buyers messaging a seller with a question the listing should have already answered. It climbs two to three weeks before returns citing 'not as described' ever move."
Why this works
This is the direct answer to the literal question, stated as one checkable claim.
Stage 5
Name how it gets gamed
Say it like this
"A seller could pad the description until nobody bothers asking a question, and the question rate would drop while the real vagueness problem never actually left."
Why this works
Every metric has a cheap way to be satisfied without doing the real work, naming it shows you'd catch it.
Stage 6
Give the decision, and close
Say it like this
"When question rate on an AI draft clears the hand-written baseline, I hold it for the seller to review before it goes live. A leading signal nobody acts on is just a chart, not a fix."
Why this works
Restates the direct answer with a real action attached, ready for a follow-up push.

Let's learn

What happens the first time an AI-written listing gets something wrong, quietly, and nobody complains?

Tallowmark runs SnapList, a tool that drafts a resale listing's description straight from a seller's uploaded photos, no typing required.

Knowledge spark: what's a leading versus a lagging signal? A lagging signal shows you the damage after it's already happened, a return, a dispute, a bad review. A leading signal moves first, while the damage is still small and cheap to stop. The trick is finding the one that reliably shows up early, not just any number that happens to be available.

Before SnapList, a seller spent about twelve minutes writing a listing description by hand. SnapList cuts that to under a minute, and adoption climbed fast because sellers could list ten times as many items in the same evening.

The extra speed was never the danger. The danger was a draft that sounded confident and clean while quietly leaving out the one detail a buyer actually needed to trust it.
Buyer question rate vs return rate, vintage furniture listings, week by week
20% 10% 0% Wk 1 Wk 2 Wk 3 Wk 4 Wk 5 Returns: 11% Questions: 15%
Question rate (solid) started climbing in week two. Return rate (dashed) didn't move until week five, three weeks later, and by then the item had already shipped.

At its worst: a whole category, vintage furniture, drifts toward vague material descriptions, calling a veneer piece "wood" without saying which kind, for months. Tallowmark only finds out once a wave of returns lands during a big promotional push, exactly when pulling every affected listing costs the most.

The decision I would take back We built SnapList's only quality dashboard around return rate and dispute rate, because finance already tracked those numbers and it was cheap to reuse the same pipeline. That made sense when SnapList handled a small slice of listings and returns were the only cost anyone had actually budgeted for.

What I would leave alone: for a seller with years of clean listings and a strong track record, per-listing question-rate monitoring matters less. Their own history is already a stronger signal than watching every new draft they publish.

The lesson: a bad AI output doesn't announce itself with a return. It announces itself quietly, in a question a buyer shouldn't have needed to ask, weeks before the return ever shows up on anyone's dashboard.

Now here is the same thing as a story

The short version above is what you'd say defending this alert to Tallowmark's operations lead. Read this one for how the near miss actually got caught.

The trust and quality floor at Tallowmark gets quiet around two in the afternoon, the slow stretch before evening listings spike.

SnapList launched to good months. Sellers loved the time it saved, returns stayed flat, and Beckett Osei's Monday dashboard always came back looking calm.

Hand sketched flow diagram titled From a vague photo to a buyer's question. Five boxes: photo uploaded, AI drafts description, buyer reads listing, buyer asks a question highlighted, signal logged.
The fourth box is the one nobody at Tallowmark had ever wired up to anything.

The habit thinned in three beats. First, Beckett glanced at the dashboard quickly each Monday, since it always looked fine. Then he stopped drilling into any single category, since nothing on the aggregate view ever stood out. Then he stopped checking mid-week entirely, since Monday's number always matched what he expected to see.

One Friday afternoon, ahead of a Fall Furniture promotion scheduled to send ten times the usual traffic into that category on Monday, Beckett flipped through a printed sample sheet he still liked to annotate by hand, an old habit from before the dashboards existed. Several listings called a veneer piece "solid wood."

Hand sketched comparison diagram titled Two clocks, only one rings on time. Left panel, a gauge icon labeled Return rate, caption rings in week 5. Right panel, a gauge icon labeled Question rate, caption rings in week 2.
Both clocks tell the truth eventually. Only one of them tells it while there's still time to act.

Return rate for that category hadn't moved yet, the way it always looks fine right before it isn't. But when Beckett pulled the raw message logs, buyer questions asking about wood type had already climbed for two straight weeks, a signal nobody had ever plotted before.

Hand sketched timeline titled Three weeks nobody was looking. Four milestones: drafts miss detail week 1, question rate climbs week 2 highlighted, near miss caught week 3, return rate moves week 5.
By the time the return rate would have rung, the promotion would already be two days old.

In an early planning meeting, the team had agreed the quality dashboard should track what finance already tracked, return rate and dispute rate, since it was free to build off existing data. Nobody in that room asked what would show up on a chart three weeks before a return ever happened.

Hand sketched labeled parts diagram titled What counts as an implicit bad-output signal. Center box icon labeled Bad output, with four callouts: buyer asks a question, buyer zooms photos repeatedly, return cites mismatch, seller rewrites draft.
Four ways a bad draft shows itself. The return is the last one to arrive, not the first.
Returns citing "not as described," before and after the question-rate alert
340 170 0 340 Before the alert 95 After the alert
Returns citing "not as described" fell by more than half once the question-rate alert gave the team three extra weeks to fix a draft before it shipped.

With the leading indicator wired in as a real alert, the same drift is now caught automatically in week two, giving three weeks to fix the description prompt before a promotion, instead of Beckett noticing by chance while flipping through paper on a slow Friday afternoon.

The real fear was never the returns budget. It was almost letting a systemic gap ride straight into the one week when the most new buyers would ever see Tallowmark's furniture category for the first time, souring a first impression before the brand had a chance to earn it.

I built the quality dashboard off numbers finance already had because it was free and it looked thorough. It took a Friday afternoon with a printed sheet, a habit I'd almost stopped keeping, to see that the real leading signal had been sitting in our own message logs the whole time, never once plotted.

LEAD, before the number movesNot a list of "signs of trouble." LEAD is what tells you which one rings first.

L
Link. The outcome that actually matters.
Whether a buyer keeps the item without a fight, not whether the description reads smoothly.
Without naming this first, any signal you pick is a guess at what "bad" means.
E
Early signal. What moves before the outcome does.
Buyer question rate on an AI-drafted listing. It climbs two to three weeks before returns citing "not as described" ever move.
The hardest step and the direct answer to the question itself.
A
Abuse. How the metric gets gamed.
A seller pads a description until nobody bothers asking, and the question rate drops while the real vagueness never left.
Every metric has a cheap way to be satisfied without fixing the real problem.
D
Decision. What you'd actually do at each level.
Question rate above the hand-written baseline holds the draft for seller review. At or below, it publishes straight through.
A metric nobody acts on is a dashboard decoration.
Hand sketched quadrant titled Sorting listing problems. Axes how fast it's noticed from slow to fast, and how costly if missed from cheap to expensive. Typo sits bottom right, fast and cheap. Vague material sits top left, slow and expensive. Wrong price sits middle right, fast and expensive. Missing dimension sits bottom left, slow and cheap.
The top left corner, slow to notice and expensive if missed, is exactly where a leading signal earns its keep.

The recap, one line per letter: link is whether a buyer keeps the item without a fight, early signal is the buyer question rate that moves weeks ahead of returns, abuse is padding a description until nobody asks, and decision is the hold-for-review threshold set at the hand-written baseline.

And if you want to be sure it really works, try it somewhere elseSame four letters, an HVAC dispatch service instead of a resale marketplace. A different trade, the same leading signal.

Pallister Field Services sends technicians to homes and dispatches an AI tool that drafts the diagnostic summary a customer receives after a repair visit. Farrow Okafor manages quality for that tool.

Mapped onto LEAD: link is whether the diagnosis holds without a costly return visit, not whether the summary reads professionally. Early signal is the customer callback-request rate asking to clarify or dispute a diagnosis summary, which climbs a week or two before the "no-fix confirmed" return-visit rate ever moves. Abuse: a technician could tell a confused customer to just call the office directly instead of using the tracked callback channel, which would hide the real signal without fixing anything. Decision: when callback rate on an AI-drafted diagnosis for a given technician clears the hand-written baseline, a senior technician reviews the report before it's finalized and sent.

Hand sketched icon list titled The dispatcher's leading-signal checklist. Three items: a gauge icon labeled Callback request rate, a circle icon labeled Technician re-visit rate, a question mark box icon labeled Customer confusion calls.
Same shape, a different trade. The callback rate is HVAC's version of the buyer's question.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "watch buyer questions, not returns, since questions move first," and stop.
Cost: there's no time to build a full alerting pipeline this sprint. Say so honestly, and start with a weekly manual pull of the question-rate number by category, since a thin signal checked by hand still beats no signal at all.
The model gets better, for real: if SnapList's photo-to-text accuracy improves overall, that's still not a reason to drop the question-rate watch, a better model on average can still drift badly on one specific category nobody's checking closely.

Where people run it wrong.
They build the quality dashboard around whatever data was already free to reuse, instead of asking what would move first.
They treat a single number as safe forever, without checking whether it can be gamed without fixing the real problem.
They wait for the lagging metric to move before reacting, by which point the cheap fix has already turned into an expensive one.

How to use it live. When someone asks what implicit signal tells you an output was bad, ask yourself one question first: what would move weeks before the obvious lagging number does. That's your answer, not whichever signal happens to already be on a dashboard.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "what implicit signal tells you an output was bad"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. The early signal step is the direct answer to a metric question like this one.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Beckett Osei, the trust and quality PM at Tallowmark who watches what SnapList's AI-drafted listings ship.
3 · THE LINK
What's the real outcome, not the model's own polish?
Tap to flip
ANSWER
Whether a buyer keeps the item without a fight. A smooth-sounding description that leaves out the wrong detail still fails this test.
4 · THE EARLY SIGNAL
What's the leading tell in this story?
Tap to flip
ANSWER
Buyer questions asking about something the listing should have already answered. It climbed weeks before the return rate ever moved.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building the quality dashboard only around return rate and dispute rate, because that data was already free to reuse from finance.
6 · THE NUMBER
Fill in the blank: returns citing "not as described" fell from 340 a month to ___ a month after the alert shipped.
Tap to flip
ANSWER
95 a month. The fix came three weeks earlier than the old return-rate dashboard would ever have caught it.
7 · THE REPLAY
Same drift, redesigned system. What changes?
Tap to flip
ANSWER
The question-rate alert catches the drift automatically in week two, instead of Beckett finding it by chance on a printed sample sheet on a slow Friday.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Pallister Field Services' AI-drafted HVAC diagnostics. Same shape: customer callback-request rate, moving ahead of the return-visit rate.

Check yourself Score: 0 / 0

True or false
1. True or false: return rate is the fastest way to see that SnapList's descriptions have gotten worse.
  • True
  • False
Show hint
Look at the line chart comparing question rate and return rate.
Show answer
False. Return rate is a lagging signal. Buyer question rate rises two to three weeks earlier, while the fix is still cheap.
Multiple choice
2. Why does a buyer question specifically count as a signal, rather than just normal marketplace chatter?
  • A. Because messaging a seller is against Tallowmark's terms of service.
  • B. Because it's a question about something the AI-drafted listing should have already answered, revealing a real gap in the description.
  • C. Because sellers get paid extra for answering questions quickly.
  • D. Because buyers only message sellers when they intend to return an item.
Show hint
Look at the abuse case and the "how it gets gamed" list.
Show answer
B. Not every question is a signal, only the ones asking about something a complete description should have covered, like material or condition.
Fill in the blank
3. Fill in the blank: buyer question rate on the vintage furniture listings rose to ___% by week four, three weeks before the return rate caught up.
Show hint
Look at the line chart's week 4 point.
Show answer
15 percent. Return rate for the same week was still only 6 percent, and didn't reach 11 percent until week five.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Building the quality dashboard only off return rate and dispute rate, since that data was free to reuse from finance. It made sense when SnapList handled a small slice of listings and returns were the only budgeted cost.
Short answer, where it wouldn't matter
5. Name a seller on Tallowmark where watching question rate closely matters less.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A seller with years of clean listings and a strong track record. Their own history is already a stronger signal than watching every new draft closely.
Short answer, apply it yourself
6. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think of something you stopped double-checking once it was reliable for a long stretch.
Show answer
Model answer: Many people stop reading a map app's turn instructions closely, or stop verifying an auto-fill address, once it's been right often enough in a row.
Before you close the answer
Why this works
Tests whether you reach for a leading indicator instead of waiting for the lagging cost metric to move, and whether you can name how your own leading signal could be gamed.
Follow-up traps
"Isn't a rising question rate actually a good sign, since buyers are engaging with the listing?" Response: not when the question is about something the listing should have already answered. That's a gap in the draft, not curiosity.

"What if sellers just start answering questions faster? Doesn't that fix it?" Response: no, faster answers treat the symptom. The real fix closes the gap in the draft itself so the question never needs to be asked at all.
If pressed
Tallowmark's real alert only counts a question as a signal when it's about an attribute the AI draft's own schema was supposed to fill in, material, dimensions, condition, not any question at all, since some questions are just about shipping time and carry no quality signal whatsoever.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more