CaseAdvancedAI Opportunity & Model Strategy / Data strategy as product strategy / #11

Design a labelling strategy for a feature launching in eight weeks.

ORDERforty thousand deck photos, and the one species with no room for error

Picture this: you have eight weeks, forty thousand deck photos, and one species you cannot afford to get wrong. Drayton Bay Seafood Co. is launching CatchSight, a tool that identifies fish species from deck photos so a crew of sixty vessels can sort faster and stay inside quota and protection rules. Corwin Haas, the fleet's data lead, is the one who has to decide what gets labelled first, and how carefully, before the new season starts.

The direct answer
Rank the labelling decisions by how expensive they are to undo, not by how important each species feels. Lock the species taxonomy and edge-case rules in week one, since getting those wrong means relabeling much of the corpus. Label the common, high-volume species first to hit coverage targets. Route real review depth, up to full double-blind checks, only to the species where a wrong label carries outsized cost, like a legally protected one. A cheap-to-fix dial, like spot-check percentage, can start light and tighten the moment evidence says it needs to.
Do this, in order
  1. Rank every labelling decision by how expensive it is to undo, before spending a single day of the eight weeks.Why: undo-cost, not how important a species feels, is what should decide the order.
  2. Lock the species taxonomy and edge-case labeling rules in week one.Why: changing this after labelling is underway means relabeling much of the corpus, the single most expensive late mistake.
  3. Build a small gold or validation set with double-blind labelling before scaling to the full corpus.Why: it catches labeler confusion cheaply, on a handful of photos, instead of expensively across tens of thousands.
  4. Label the common, high-volume species first to hit coverage and speed targets.Why: this is the bulk of what the tool needs to work at all, and it's cheap to expand later.
  5. Route full double-blind review only to species where a wrong label carries outsized cost, like a legally protected one.Why: a rare, high-stakes species needs a different bar than a common one the model already sees thousands of times.
  6. Say plainly where a lighter spot-check is still the right call, like common species with thousands of examples and no regulatory stakes.Why: shows judgment instead of demanding the same rigor everywhere out of caution.

How to answer this, stage by stage

Nobody is scoring whether you know that labelling needs quality control. They're scoring whether you can rank it, out loud, by what's actually expensive to get wrong.

Stage 1
Scope it to one fleet, one deadline
Say it like this
"Let's ground this in Drayton Bay's CatchSight tool, and the actual eight weeks before the new fishing season, where every labelling choice either survives being redone or doesn't."
Why this works
Turns "design a labelling strategy" into something with a real clock on it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as ORDER. Outcome, what every labelling choice competes to protect. Reversibility, which decision is hardest to undo. Dependency, what's forced first. Evidence, what's cheap to learn early. Rank, the actual eight-week order."
Why this works
Signals a repeatable way to sequence a deadline, not a to-do list assembled by instinct.
Stage 3
Reframe the question
Say it like this
"The trap is treating every species and every QA choice as equally important. The real question is which decision, if wrong, means relabeling a huge chunk of forty thousand photos, versus which one you can safely tighten later once you see where it's needed."
Why this works
Separates a real sequencing answer from an even, undifferentiated labelling plan.
Stage 4
Give the rank, applied to two real decisions
Say it like this
"Locking the species taxonomy and edge-case rules happens in week one, non-negotiable, since getting that wrong means relabeling most of the corpus. Spot-check depth starts light, five percent, and only tightens for a specific species if the evidence says so, because that dial is cheap to turn later."
Why this works
This is the direct answer, applied concretely instead of stated as an abstract principle.
Stage 5
Prove it with the compressed near miss
Say it like this
"A routine five percent spot-check caught a labeler mismarking a protected juvenile species as a common lookalike in two of twenty photos checked, out of a corpus where that species barely had a hundred and eighty examples total. We moved to full double-blind review on all three protected species that same week, without touching the QA depth for the other thirty one."
Why this works
Compresses the whole story into the one moment that proves reversibility, not raw accuracy, was the real ranking test.
Stage 6
Say what you'd measure, and what you'd leave alone
Say it like this
"I'd track disagreement rate between the two blind reviewers on protected species specifically, not overall labelling accuracy, since that blended number would have hidden this the same way it almost did. I'd leave the common species' lighter spot-check alone, since nothing in the evidence said it needed tightening."
Why this works
Shows judgment about where the extra rigor earns its cost, and where it doesn't.
Stage 7
Close on the one line
Say it like this
"Lock what's expensive to undo first, taxonomy and edge-case rules, then let the cheap, reversible dials, like spot-check depth, start light and tighten only where the evidence actually points."
Why this works
Restates the direct answer in one breath, the way a live answer needs to land.

Let's learn

Hand sketched icon list titled ORDER the five letters. Five rows: Outcome, what every choice competes to protect, scale icon. Reversibility, which decision is hardest to undo, gauge icon, shown in a different color. Dependency, what's forced first, box icon. Evidence, what's cheap to learn early, document icon. Rank, state the order defend the top, question mark box icon.
The five letters, held up as one page. Reversibility is the step this question is really testing.

CatchSight reads a photo of a fish on deck and identifies its species, so a crew can sort quickly and stay inside quota and protection rules. Before it existed, an uncertain catch meant a four-minute radio call to an onboard observer, a few times per haul. CatchSight's target is under two seconds per photo, with anything low-confidence or near a protected species flagged for a person.

Hand sketched flow diagram titled What unblocks what, across the eight weeks, first step emphasized. Five steps left to right: Lock taxonomy. Build gold set. Label common species. Review protected species. Final reconciliation.
The first step is the one everything else depends on, and the one hardest to undo if it's wrong.
Spot-check coverage, before and after the near miss
100% 50% 0 5% 15% Common species (31) 5% 100% Protected species (3)
Both tiers started at the same blanket five percent. Only the protected species needed to move all the way to full review.

Here's the turn: the extra mistakes at that spot-check rate were never about the overall labelling accuracy looking bad, it looked fine, blended across thirty four species. Say plainly what mattered instead: two mislabeled photos out of a rare species with barely a hundred and eighty examples total is a very different kind of miss than two mislabeled photos out of a common species with thousands.

A five percent spot-check is not one policy. On a species with forty thousand photos it's thorough. On a species with a hundred and eighty, it barely looks at all.
The choice I would take back The plan set one blanket spot-check rate, five percent, for every species, because a single simple rule was easy to explain to twelve part-time labelers in a one-page instructions sheet. That made sense when every species had thousands of examples. It stopped making sense for the three species where five percent of a small number is barely any real coverage at all.

What I would leave alone: the common species' lighter spot-check, now fifteen percent, stayed exactly where it was after the scare, since nothing in the evidence ever suggested those needed full review too.

The lesson: a labelling strategy isn't one number, like percent reviewed. It's a set of decisions ranked by how much a mistake there would actually cost, and by how hard that decision is to undo once the clock is already running.

Now here is the same thing as a story

The short version above is what you'd say scoping this on a whiteboard. Read this one for how close the fix actually came to landing after launch instead of five weeks before it.

Corwin Haas has managed fleet data for Drayton Bay for seven years, and can usually tell from a single deck photo which species a fresh labeler is going to struggle with.

Hand sketched comparison titled Reversible, or not. Left panel, a gauge icon labeled Spot check depth, caption cheap to tighten later. Right panel, a box icon labeled Taxonomy and edge rules, caption expensive to redo mid stream, shown in a different color.
Two kinds of decision. Only one of them punishes you for getting it wrong on day one instead of day thirty.

Week one locked the taxonomy, thirty four target species, three of them legally protected, along with rules for edge cases like a partially obscured fish or a juvenile that looks unlike its adult form. A small gold set, five hundred photos, got double-blind labelled first, to catch confusion before scaling up. From week two, twelve part-time labelers, mostly off-season fleet observers, began working through the full forty-thousand-photo corpus, with a standard five percent spot-check on every labeler's work.

Knowledge spark: why would a common species need less scrutiny than a rare one? A model learns from repetition. A species with thousands of labelled examples can absorb a handful of mistakes without its overall accuracy moving much. A species with only a hundred or two examples has almost no room for error, since each mislabeled photo is a much bigger share of everything the model will ever see of that species.
Hand sketched quadrant titled Where review depth actually belongs. X axis how common, rare to common. Y axis cost if wrong, low stakes to high stakes. Common lookalike species placed common and low stakes. Protected juvenile species placed rare and high stakes. Common quota species placed common and medium high stakes.
Review depth was never about the species. It was always about this exact chart, rarity against real cost.

In week five, a routine spot-check on one labeler's batch turned up something worse than a typo. Two of twenty checked photos had a young silverfin grouper, one of the three protected species, mismarked as a common lookalike species almost identical at that age. The corpus held only about a hundred and eighty photos of that grouper in total. A five percent check across the whole batch had, by luck, caught the problem at all.

Hand sketched timeline titled Eight weeks, ranked, third milestone emphasized. Four milestones: Week 1, taxonomy locked, gold set built. Weeks 2 to 4, common species labelled at volume. Week 5, near miss forces full protected review, shown in a different color. Week 8, reconciled and launched.
Nobody planned to change the QA plan mid-stream. One near miss at week five made the case on its own.

Corwin moved that same week to full double-blind review for all three protected species, two independent labelers plus a marine biologist consultant adjudicating any disagreement, without touching the fifteen percent spot-check the common species had already settled into. Seven hundred ten photos across the three protected species needed re-checking under the new standard, a targeted week of extra work, not a redo of the whole forty-thousand-photo corpus.

Cumulative photos labelled, common species versus protected species
40,000 20,000 0 40,000 wk5, near miss wk6, all 710 reviewed Wk 1 Wk 5 Wk 8
Teal is the 40,000-photo common-species corpus. Coral is the 710-photo protected-species slice, small enough to fully re-review without missing launch.

When the plan was first written, someone said, "one spot-check rate for everyone keeps this simple for a twelve-person labeling team on a tight deadline," and it sounded reasonable, since a single clear rule really is easier to train twelve part-time labelers on than a shifting one.

Rerun the same eight weeks with review depth ranked by cost from day one: the three protected species get full double-blind review from week one, not week five, since the plan already knew they were rare and legally sensitive before a single photo was labelled. The near miss never becomes a scramble, because there was never a version of the plan where those three species got the same five percent as everything else.

What I'd tell myself, looking at those two mismarked photos: the corpus was never short forty thousand photos. It was short a plan that had already decided, before week one, which hundred and eighty photos could least afford five percent.

ORDER, the rank that found the one species that mattered mostNot a script for reviewing every label to the same depth out of caution. ORDER is what tells you exactly where the extra rigor actually earns its cost.

O
Outcome. What every choice competes to protect.
Accurate species identification, especially of protected species, without slowing the crew down or missing the season's start.
Without a shared outcome, ranking labelling decisions is just habit.
R
Reversibility. Which decision is hardest to undo.
The taxonomy and edge-case rules are nearly impossible to change mid-stream without relabeling much of the corpus. Spot-check depth is cheap to tighten at any point.
This is the hardest step, and the one the original blanket rule never actually asked.
D
Dependency. What's forced first.
Bulk labelling can't meaningfully start until the taxonomy and a small gold set exist; launch can't happen until all three protected species clear full review.
Some order is forced by what the plan actually needs, not by preference.
E
Evidence. What's cheap to learn early.
The routine five percent spot-check was itself the evidence step, and it barely caught the protected-species problem by luck.
Cheap evidence, gathered early enough, is what makes a labelling plan self-correcting instead of a fixed schedule.
R
Rank. State the order, defend the top.
Taxonomy and gold set first, common species at volume next, protected species at full review throughout, reconciliation last.
The ranking follows directly from undo-cost, not from which species feels most important.

The recap, one line per letter: outcome is accurate identification without slowing the crew or missing the season, reversibility is taxonomy locked hard against spot-check depth left soft, dependency is taxonomy and a gold set before bulk labelling, evidence is the spot-check that (barely) caught the near miss, and rank is taxonomy first, common species at volume, protected species at full review throughout, reconciliation last.

And if you want to be sure it really works, try it somewhere elseSame five letters, a garment factory instead of a fishing fleet. Different flip family entirely, the same review that quietly stopped mattering to anyone.

Priyasha Deol runs quality operations at Loomcraft Garment Works, where an AI tool flags fabric defects from line-camera photos before shipment. Mapped onto ORDER: outcome is catching real defects, especially safety-relevant stitching failures, without slowing the production line. Reversibility says the defect taxonomy is expensive to redo once labelling starts, while spot-check depth is cheap to adjust. Dependency needs the taxonomy locked before bulk labelling begins. Evidence favors a two-day pilot on one production line before committing all eight weeks broadly. Rank puts taxonomy first, common cosmetic-defect labelling next, and escalated double-blind review on safety-relevant defects throughout. The flip here is abandonment, not verification: once inspectors noticed the tool flagging common cosmetic flaws at a high rate, mostly harmless, they quietly stopped opening the flagged-photo queue at all after a few weeks, which meant the rare, real, safety-relevant defect it also correctly caught went unseen too, since nobody was watching whether the queue was still being opened.

Hand sketched decision tree titled Which defect photos get double checked. Root, garment defect photo. Four branches: common cosmetic flaw leads to single labeler, light spot check. Safety relevant stitching failure leads to double blind, full review, shown in a different color. New unclassified defect shape leads to escalate, update taxonomy. Low image quality leads to reshoot, do not label.
A different flip entirely: not a rare species barely surviving a five percent check, but a whole queue nobody was opening anymore.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "lock what's expensive to redo first, taxonomy and edge rules, then rank review depth by cost if wrong, not by species count," and stop.
Cost: there's no time to build a full active-learning pipeline before week eight. Say so honestly, and ship a manual escalation rule for rare, high-stakes categories first, the fancier tooling can come later.
The rare species turns out to be easier than assumed: if a "high-risk" species actually looks nothing like any common lookalike, that's good news worth confirming early, and it can safely stay at a lighter review tier.

Where people run it wrong.
They apply one QA depth to every category, so a rare, high-stakes case gets the same thin check as a common, low-stakes one.
They treat blended accuracy as proof the labelling is fine, missing a serious problem hiding in a category too small to move the average.
They tighten review everywhere after a scare, instead of tightening only where the evidence actually points.

How to use it live. The moment an interviewer asks for a labelling strategy under a deadline, ask yourself which decisions are expensive to undo and which are cheap. Lock the expensive ones first, and let the cheap ones start light and tighten only where real evidence says they need it.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Verification flip: a five percent spot-check moved to full, one hundred percent double-blind review, but only for the three protected species, not across the board.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Corwin Haas, who has managed fleet data at Drayton Bay Seafood Co. for seven years.
3 · THE HABIT
What did the original plan treat as simple to keep consistent for twelve labelers?
Tap to flip
ANSWER
One blanket spot-check rate, five percent, applied the same way to all thirty four species regardless of how rare or high-stakes each one was.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Five percent spot-checking every labeler's work, versus full double-blind review of every single photo, for the three protected species specifically.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting one blanket five percent spot-check rate for every species instead of ranking review depth by how rare and high-stakes each species actually was, from day one.
6 · THE NUMBER
Fill in the blank: the protected juvenile species had only about ___ photos in the entire forty-thousand-photo corpus.
Tap to flip
ANSWER
180 photos, which is why a five percent spot-check barely covered it at all.
7 · THE REPLAY
Same eight weeks, review depth ranked by cost from week one instead of week five. What changes?
Tap to flip
ANSWER
All three protected species get full double-blind review from the start, the near miss never becomes a scramble, and the common species' lighter spot-check never has to change at all.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Loomcraft Garment Works' fabric-defect detector. The flip is abandonment: inspectors quietly stopped opening the flagged-photo queue after too many harmless cosmetic flags, missing the rare safety-relevant defects too.

Check yourself Score: 0 / 0

Multiple choice
1. What actually determines how much review depth a species should get, according to this answer?
  • A. How many total species exist in the taxonomy.
  • B. How visually distinctive the species looks in a photo.
  • C. How rare the species is and how costly a wrong label would be.
  • D. How many of the twelve labelers have experience with that species.
Show hint
Look at the quadrant diagram and the direct answer.
Show answer
C. Rarity and cost-if-wrong together decide review depth, which is why the three protected species got full review while thirty one common species stayed at a light spot-check.
True or false
2. True or false: this answer recommends raising every species' review to full double-blind checking after the near miss.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Only the three protected species moved to full review. The thirty one common species stayed at a lighter fifteen percent spot-check, since nothing in the evidence said they needed more.
Fill in the blank
3. Fill in the blank: after the near miss, protected-species labelling moved from a five percent spot-check to ___ percent, full double-blind review.
Show hint
Look at the grouped bar chart, "spot-check coverage, before and after the near miss."
Show answer
100 percent. All 710 photos across the three protected species were re-reviewed under the new standard by week six.
Short answer, where it wouldn't matter
4. Name the part of this labelling plan where the original blanket five percent rule genuinely didn't need to change, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The common species' spot-check. With thousands of examples each, a small percentage of missed errors doesn't meaningfully move the model's accuracy, so the lighter check was already the right call.
Short answer, apply it yourself
5. Think of a project you've worked on with a hard deadline. What was one decision that would have been expensive to undo later, and did the plan treat it that way from the start?
Show hint
Look for a decision made early that would have forced redoing a lot of later work if it turned out wrong.
Show answer
Model answer: Choosing a database schema before importing years of historical data. Getting the schema wrong meant re-importing everything, while adding a new report later was cheap and reversible by comparison.
Short answer, work the number
6. If the protected species' corpus had been 720 photos instead of 180, would a five percent spot-check have been more likely to catch the mislabeling pattern, and roughly how many photos would that check cover?
Show hint
Five percent of 720 versus five percent of 180.
Show answer
Model answer: Yes, more likely. Five percent of 720 is about 36 photos, roughly four times the coverage of the actual 180-photo corpus, giving a real pattern more room to show up.
Before you close the answer
Why this works
Tests whether you'll rank a labelling plan by what's actually expensive to get wrong, or default to one blanket rule that feels fair but ignores which mistakes actually cost the most.
Follow-up traps
"Isn't full double-blind review for every species just the safer choice?" Response: it would blow the eight-week deadline for no real benefit on species where thousands of examples already absorb the occasional mistake without hurting accuracy.

"What if the near miss hadn't been caught by the five percent check at all?" Response: that's exactly the risk of a blanket rule on a rare category, which is why the redesigned plan puts protected species at full review from week one instead of waiting for a lucky catch.
If pressed
The double-blind process for protected species used a disagreement threshold, not a simple majority: any split between the two labelers routed automatically to the marine biologist consultant, rather than defaulting to whichever labeler answered first.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more