CaseIntermediateResponsible AI & Advanced Practice / AI product case study teardowns / #8

Compare two competing AI writing tools on their handling of user control.

PICK the products are Penlight and Draftwing, two AI writing tools Baptiste's translation agency has tried

Penlight and Draftwing both polish translated documents with an AI rewrite pass. Baptiste Ferron owns a small translation agency handling business and legal documents, and has run both tools across the same kind of client work.

The direct answer
Draftwing gets user control right by making one line the unit you accept or reject. Penlight makes the whole rewritten paragraph the unit, so a translator either accepts one risky reworded clause along with the good edits, or rejects the good edits to avoid the risky one. That single design choice, the size of what you're allowed to approve, is the entire gap between the two.
Do this, in order
  1. Make the unit of approval one line, not a whole paragraph.Why: this decision alone explains almost the entire control gap between the two tools.
  2. Optimize against a silently accepted risky edit, not against a rejected safe one.Why: rejecting a good trim costs seconds. Accepting a bad legal rewording can cost a client relationship.
  3. Flag high-stakes clause types, like legal terms and names, for closer review.Why: not every line carries the same risk if it's wrong.
  4. Leave low-stakes trims, like filler words, on a fast accept-all path.Why: nothing bad happens if one of those gets accepted without a second look.
  5. Watch review thoroughness as document volume rises, not just at a calm baseline.Why: an all-or-nothing tool's real risk only shows up once someone's actually busy.
  6. Reconsider line-level review if it ever measurably slows turnaround enough to cause skimming.Why: that's the one signal that would flip this pick back toward simpler batch review.

How to answer this, stage by stage

Six moves. Committing to a real position, with a real cost behind it, matters more here than listing every feature each tool has.

Stage 1
Scope it to two real tools and one real user
Say it like this
"I'll compare Penlight and Draftwing, two AI writing tools, through a translation agency owner who's used both on the same kind of client documents."
Why this works
Turns "handling of user control" into something you can point at directly.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position first, then impact, cost asymmetry, and what would flip my pick."
Why this works
Signals a real comparison with a stance, not a feature-by-feature list.
Stage 3
Take the position
Say it like this
"Draftwing gets control right because it lets you accept or reject one line at a time. Penlight forces you to accept or reject the whole rewritten paragraph as one block."
Why this works
One sentence, no hedging, exactly what PICK's opening step requires.
Stage 4
Name who feels each kind of error
Say it like this
"Rejecting a good filler-word trim costs a translator ten seconds to redo by hand. Accepting a subtly reworded legal clause, hidden inside an otherwise good paragraph, can cost the agency a client's trust entirely."
Why this works
Puts both sides of the tradeoff in real, comparable terms.
Stage 5
Name the cost asymmetry
Say it like this
"A rejected safe edit is cheap and visible, you just redo it yourself. A silently accepted risky edit is hidden and expensive, since it's bundled inside a paragraph you approved as a whole. Optimize against that one."
Why this works
This is the actual heart of the comparison, not a surface-level feature difference.
Stage 6
Give the kill criteria, and close
Say it like this
"If line-level review ever slowed translators down enough that they started skimming and accepting everything anyway, I'd rethink it. Right now it adds barely a minute, and it's catching real risky edits, so the pick holds."
Why this works
Shows the position would change given different evidence, not just held out of habit.

Let's learn

Say we build two AI writing tools, both of which take a translated paragraph and offer a polished rewrite. One of them lets you take back a single sentence. The other one doesn't.

Penlight and Draftwing both do this for Baptiste's agency: take a translated business or legal document and offer AI-polished phrasing across the page.

Knowledge spark: what's an "atomic" edit, and why does its size matter? An atomic edit is the smallest unit you're allowed to accept or reject on its own. If that unit is a whole paragraph, one bad sentence forces an all-or-nothing choice about everything around it. If the unit is one line, a bad sentence can be rejected without touching anything else.

With Penlight, every AI rewrite arrives as one paragraph-sized change. Translators either accept the whole thing or reject the whole thing and redo it manually.

Average review time and risky-edit slip-through rate, per document
7 min / 7 per 100 0 4.5 min 6.2 / 100 Penlight 5.8 min 0.9 / 100 Draftwing
Draftwing takes slightly longer per document, but its risky-edit slip-through rate is nearly seven times lower.

At its worst: a legal clause about payment terms, reworded by the AI in a way that subtly shifted its meaning, gets bundled inside an otherwise fine paragraph. A busy translator, trusting the good 90 percent of the rewrite, accepts the whole thing.

The cost asymmetry that matters here Rejecting a good edit costs seconds: retype the trimmed word yourself and move on. Accepting a bad edit silently costs far more: a mistranslated legal clause reaching a client can mean an apology, a redo, and real damage to the relationship. A tool built around user control should be built to make the second kind of mistake harder, even at the cost of a slightly longer review.

What I would leave alone: a trimmed filler word or a slightly smoothed transition sentence doesn't need the same scrutiny as a legal term. Reviewing every single line with equal intensity would slow translators down for a risk that mostly isn't there.

Penlight doesn't fail because its rewrites are bad. It fails because it never lets you take back just the one sentence that was.

The lesson: user control in an AI writing tool isn't about how good the suggestions are. It's about how small a mistake you're allowed to fix without undoing everything around it.

Now here is the same thing as a story

The short version above is what you'd say defending this comparison live. Read this one for how the gap actually showed up on a real client file.

Baptiste's agency translates contracts and business correspondence for clients who care, understandably, about exact wording. His translators are good at catching subtle shifts in meaning, the kind of thing a quick skim misses.

Hand sketched comparison diagram titled Fix one line or redo the page. Left panel a box icon labeled Penlight, caption whole paragraph rewritten. Right panel a document icon labeled Draftwing, caption just the flagged line.
Both tools rewrite the same paragraph. Only one of them lets you keep 90 percent of it and fix the rest.

On a rushed Friday, one translator ran a payment-terms clause through Penlight. The AI rewrite improved the sentence's flow but subtly softened a "shall" into a "may," changing an obligation into an option.

Hand sketched flow diagram titled Penlight's all or nothing accept. Four boxes: draft requested, whole paragraph rewritten, accept all or reject all highlighted, one error slips through.
Four steps, and the third one is where the whole gap lives: there's no smaller choice available than all or nothing.

The rest of the paragraph read better than the original. Rejecting the whole thing meant losing genuinely good edits. Accepting it meant trusting one changed word buried inside four good sentences. Under deadline pressure, the translator accepted it.

Hand sketched quadrant titled Sorting edits by how they land. Axes how often it happens from rare to common, and cost if trust is broken from low to high. Legal term reworded sits top right, common and high cost. Client name changed sits top left, less common but high cost. Filler word trimmed sits lower right, common and low cost. Tone softened sits middle.
Legal term rewording sits in the dangerous corner: it happens often enough to matter, and getting it wrong is genuinely expensive.

The client caught it two weeks later, during their own legal review, and it cost the agency an uncomfortable call and a rushed correction. Nobody at Penlight had done anything wrong technically. The rewrite itself was fluent and mostly accurate.

Hand sketched icon list titled What line level control actually buys. Three rows: a document icon labeled approve one clause at a time, a scale icon labeled reject only the risky line, a gauge icon labeled keep the rest untouched.
None of these three exist in Penlight's design. All three are the entire reason Draftwing's slip-through rate is so much lower.

Baptiste switched a pilot group of translators to Draftwing the next month. On an equivalent clause, misworded the same way, the tool flagged just that one line for review. The translator rejected only the risky sentence, kept the rest of the paragraph's improvements, and the whole review took under six minutes.

Review thoroughness as weekly document volume rises
100% 50% 0 Draftwing: 82% Penlight: 34% 20 docs/wk 80 docs/wk
At a calm pace both tools look fine. Under real volume, Penlight's all-or-nothing review collapses while Draftwing's barely moves.

Someone at Penlight decided, early on, that a paragraph was the natural unit for an AI rewrite, since sentences read better in context and a whole-paragraph rewrite produced smoother prose than patching line by line. That was a reasonable design call for casual writing, where a slightly off word rarely matters.

I would take that decision back the moment the product started handling legal and business documents, where one changed word genuinely changes what a sentence promises. The unit of an edit needs to match the unit of actual risk, not just the unit that reads most smoothly.

I built the paragraph-sized rewrite because it produced the most fluent-sounding prose in every early demo, and none of those demos involved a contract. It took watching a real client catch a real changed obligation, two weeks after the fact, to see that fluent and safe aren't the same test.

PICK, in one screenFour letters, and the third one, cost asymmetry, is what actually separates these two tools.

P
Position.
Draftwing's line-level accept or reject beats Penlight's whole-paragraph accept or reject. The size of the approvable unit is the whole gap.
Stated first, no hedging, exactly what PICK asks for.
I
Impact.
A rejected safe edit costs a translator ten seconds. An accepted risky edit, bundled inside a good paragraph, can cost a client relationship.
Names both sides in real, comparable units.
C
Cost asymmetry.
Rejecting a good edit is cheap and visible. Accepting a bad one, bundled with good ones, is hidden and expensive. Optimize against the hidden one.
The hardest step, and the one this whole comparison actually turns on.
K
Kill criteria.
If line-level review ever caused translators to start skimming and blanket-accepting anyway, that would flip the pick back toward smarter automatic grouping.
Shows the position would change given different evidence.

The recap, one line per letter: position is line-level control beating paragraph-level control, impact is ten seconds versus a client relationship, cost asymmetry is optimizing against the hidden risky edit, and kill criteria is watching for review fatigue at scale.

And if you want to be sure it really works, try it somewhere elseSame four letters, a factory quality report instead of a legal contract. This time the risky word is a tolerance spec, not a legal term.

A manufacturing plant uses an AI tool to draft quality-control reports from inspection notes, comparing two competing tools much like Penlight and Draftwing. One rewrites the whole report section at once; the other lets an inspector approve one spec line at a time.

Mapped onto PICK: position is the same, line-level control over section-level control, this time applied to a measurement tolerance instead of a legal clause. Impact: rejecting a good phrasing edit costs an inspector a few seconds to redo; accepting a silently altered tolerance number, "within 0.5mm" quietly smoothed to "within 0.6mm," can mean a defective part passes inspection. Cost asymmetry: the hidden and expensive error here isn't a client dispute, it's a part that fails in the field, so the line-level tool is worth its slightly slower review. Kill criteria: if inspectors started rubber-stamping every flagged line without truly reading it, the granular control would stop providing real safety and the design would need a different fix, like requiring a physical remeasurement above a certain size discrepancy.

Hand sketched labeled parts diagram titled A QA doc tool with the same control. Center document icon labeled QA rewrite tool, with four callouts: one spec line flagged, approve just that, rest stays locked, no full re-draft.
A factory floor instead of a law office, and the exact same underlying design call: let someone fix the one risky line without redoing everything around it.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "line-level control beats paragraph-level control, because the real risk is a bad edit hiding inside good ones" and stop.
Cost: no engineering time to rebuild Penlight's editor this quarter. Say so, and start with the cheapest version: highlighting which specific words changed within an accepted paragraph, so at least the risk is visible even without granular accept and reject.
The model gets better, for real: even if Penlight's rewrite quality improves substantially, the all-or-nothing unit remains the same risk. A better model just means the bad edit is rarer, not that it's ever fully caught before it's accepted.

Where people run it wrong.
They compare two writing tools on tone or fluency, missing that the real difference is structural: what size of edit you're allowed to control.
They assume more granular review is always slower, without checking whether it actually adds much time in practice.
They blame the translator or inspector for missing a bad edit, when the tool bundled it inside something good and gave them no smaller unit to work with.

How to use it live. When comparing two AI writing tools, ask yourself one question before anything else: if the AI gets one word wrong, what's the smallest thing I'm allowed to fix? That question reveals the real control gap faster than comparing either tool's suggestions.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "compare two competing products on X"?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. Built for A-or-B comparisons where you have to actually commit.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Baptiste Ferron, who owns a translation agency handling business and legal documents, and has used both Penlight and Draftwing.
3 · THE POSITION
What's the one design decision this whole comparison turns on?
Tap to flip
ANSWER
The size of the unit you're allowed to accept or reject: one line for Draftwing, a whole paragraph for Penlight.
4 · THE COST ASYMMETRY
Which error is cheap, and which is expensive, in this story?
Tap to flip
ANSWER
Rejecting a good edit costs seconds. Accepting a bad one bundled inside good edits can cost a client relationship entirely.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Making the paragraph the unit of an AI edit, reasonable for casual writing where a slightly off word rarely matters.
6 · THE NUMBER
Fill in the blank: Penlight lets about ___ risky edits per 100 documents slip through, versus 0.9 for Draftwing.
Tap to flip
ANSWER
6.2 per 100. Nearly seven times Draftwing's rate.
7 · THE REPLAY
Same softened payment clause, Draftwing's line-level design. What changes?
Tap to flip
ANSWER
The translator rejects only that one flagged line, keeps the rest of the paragraph's good edits, and finishes review in under six minutes.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which setting, and what's the risky detail there?
Tap to flip
ANSWER
A factory quality-control report tool. There, the risky detail is a measurement tolerance instead of a legal clause.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Draftwing's risky-edit slip-through rate is about ___ per 100 documents, versus 6.2 for Penlight.
Show hint
Look at the grouped bar chart.
Show answer
0.9 per 100. Roughly a seventh of Penlight's rate, at the cost of about 1.3 extra minutes of review time per document.
Multiple choice
2. Why does this answer say Penlight's whole-paragraph accept design is risky, even though its rewrites are usually accurate?
  • A. Because Penlight's AI model is fundamentally less accurate than Draftwing's.
  • B. Because one bad line, bundled inside several good ones, forces an all-or-nothing choice that hides the risk instead of isolating it.
  • C. Because Penlight is slower to generate a rewrite.
  • D. Because Penlight doesn't work on legal documents at all.
Show hint
Look at the cost asymmetry key point.
Show answer
B. The risk isn't accuracy, it's that one bad edit can't be isolated from the good ones surrounding it.
True or false
3. True or false: this answer recommends reviewing every single line with equal scrutiny, regardless of what it says.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Low-stakes edits like filler-word trims can stay on a fast path. Only high-stakes clause types need closer review.
Short answer, name the reversal
4. What old design decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at the story's paragraph about the "reasonable design call."
Show answer
Model answer: Making the paragraph the atomic unit of an edit, reasonable for casual writing where individual word choices rarely carry real risk.
Short answer, where it wouldn't matter
5. Name a kind of edit where Penlight's paragraph-level design genuinely isn't a problem.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A trimmed filler word or a smoothed transition sentence, neither of which carries real risk if accepted without a second look.
Short answer, apply it yourself
6. Pick an AI writing or editing tool you use. When it rewrites something, can you fix just the one part you don't like, or do you have to accept or reject the whole thing?
Show hint
Think about the last time an AI tool rewrote something and you only disagreed with part of it.
Show answer
Model answer: Many popular tools still work at the paragraph or suggestion level, forcing the same all-or-nothing choice this answer describes.
Before you close the answer
Why this works
Tests whether you can identify a real, structural difference in how two AI products handle control, instead of comparing them on tone, speed, or which one "feels" better to use.
Follow-up traps
"Isn't reviewing line by line just slower and less efficient overall?" Response: it adds about a minute per document, but it cuts the risky-edit slip-through rate by nearly seven times, which is a good trade for documents where a mistake is expensive.

"Couldn't Penlight just highlight the risky word instead of changing its whole approval unit?" Response: that would help, but it still forces an all-or-nothing accept underneath the highlight. The real fix is letting the smaller unit actually be approved separately, not just visible.
If pressed
Draftwing's line-level flags aren't uniform either: it weights legal and numeric terms more heavily than general prose when deciding which lines to flag for closer review, so the added review time concentrates on the genuinely risky lines rather than spreading evenly across the whole document.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more