ConceptFoundationalAI Opportunity & Model Strategy / When NOT to use AI / #19

Explain when a template beats generation.

PICKfour templates covered 1,000 real emails, generation still found a way to invent a fifth

Greetline builds FirstTouch, the welcome email a new customer gets the moment they sign up. Zohra Petrakis is the PM deciding whether FirstTouch should generate each email fresh with a model, or fill in one of a small set of tested templates. Milo Reznik, an engineer, ran the pilot that settled it.

The direct answer
A template wins when the real variation across cases is small and structured enough to write down in advance. Here, an audit of 1,000 real signups showed the welcome email only ever needed to vary by plan tier, company name, and one next-step link, four template variants covered it completely. Generation earns its cost only when the message genuinely needs to vary in ways no fixed set of templates could capture, and this wasn't that.
Do this, in order
  1. Audit what actually varies across real past cases before choosing either approach.Why: the real range of variation, not a guess about it, is the only thing that should decide between a template and generation.
  2. Count how many templates it would take to cover that real range well.Why: a small, countable number of variants is the kill criteria. If you can write them down, a template almost always wins.
  3. Price generation's real hallucination risk into the comparison, not just its cost.Why: a generated message can state something false about a specific customer's account in a way a template, built from real fields, structurally cannot.
  4. Compare the ongoing cost, not just the build cost, of each option.Why: a template is cheap to test once and trust forever. Generation adds a real, recurring inference and eval bill for content that never needed to vary that much.
  5. Reserve generation for the cases that genuinely fail the template test.Why: this isn't a case against generation everywhere, it's about matching each tool to the shape of variation it's actually facing.
  6. Re-audit if the real range of cases changes later.Why: a template set that covered 1,000 cases well might not cover the next 1,000 if the product or customer base changes.

How to answer this, stage by stage

Nobody is scoring whether you know what a template is. They're scoring whether you'd have caught the hallucination before a customer did.

Stage 1
Scope it to one real decision, not templates versus AI in general
Say it like this
"Let me give you a real case. A team is deciding whether to generate each customer's welcome email fresh, or fill in one of a small set of templates. An audit of 1,000 real signups is the thing that should actually decide it, not a preference for one approach or the other."
Why this works
Grounds the answer in something checkable instead of a general opinion about AI-generated content.
Stage 2
Say your structure out loud before any content
Say it like this
"I'll run this as PICK. Position, the real distinction between the two. Impact, what's lost picking the wrong one. Cost asymmetry, which mistake is cheap and which is expensive. Kill criteria, the one test that actually decides it."
Why this works
Signals a repeatable way to weigh a tradeoff instead of a personal preference dressed up as a decision.
Stage 3
State the position plainly
Say it like this
"A template wins when the real variation is small enough to write down in advance. Generation earns its cost when the message genuinely needs to vary in ways no fixed set of templates could cover. Those are two different shapes of problem, and you can tell which one you have by looking at real past cases."
Why this works
This is where a strong answer separates from "AI is usually better" or "templates are usually safer" as an unexamined default.
Stage 4
Give the one decision: run the kill-criteria test
Say it like this
"Here's what I'd actually do. Pull 1,000 real past cases and count how many templates it would take to cover them well. If that number is small, four, five, six, build the templates. Here, it was four. That's not a close call once you've actually counted."
Why this works
This is the direct answer, stated as an actual test you'd run against real data, not a general philosophy.
Stage 5
Prove it with the compressed failure
Say it like this
"This is exactly what the pilot showed. Milo ran both approaches against the same 200 real accounts. The templates made zero factual errors, because they only ever pull real account fields. The generated version told 14 customers, 7 percent, they had a feature their plan didn't actually include."
Why this works
This is where the story lives, compressed to the one number that actually proves the decision.
Stage 6
Name the AI-specific risk and the cost being traded
Say it like this
"The honest reason generation loses here isn't that it's a worse writer, the emails read fine. It's that generation can state something false about a specific customer's account, a real hallucination risk a template structurally can't have, since a template only ever pulls fields that are actually true. And it costs real, recurring money to run, on content that never needed the flexibility."
Why this works
This is the load bearing judgment. It wouldn't make sense to ask this about a feature with no model in it, since the risk is specifically that a model can say something false a template can't.
Stage 7
Say what you'd leave alone, then close on one line
Say it like this
"This isn't a case against generation everywhere at Greetline. A customer support reply that genuinely has to respond to whatever someone actually wrote needs real generation, no fixed template could cover that. This one welcome email just never needed that flexibility. Count the real range first, and let the count decide."
Why this works
Closes with real judgment instead of a blanket stance on AI-generated content, and restates the direct answer in one breath.

Let's learn

FirstTouch is the email a new customer gets the moment they sign up for Greetline, meant to feel personal and tell them exactly what to do next.

Hand sketched icon list titled What actually varies across 1,000 real welcome emails. Four rows. A gauge icon captioned plan tier, exactly four real values. A document icon captioned company name, pulled from account data. A person icon captioned sender name, one of three account reps. A box icon captioned one next-step link, tied to plan tier.
Four things, and only four things, actually changed across 1,000 real emails. Everything else was identical every time.

Before the audit, the team had assumed a generated, fully personalized welcome email would obviously feel better than a template, the same instinct that makes "personalized" sound like it must beat "templated" by default. Nobody had actually checked what real personalization would need to cover.

Factual errors per 200 emails, template versus generated
10% 5% 0 0% Template emails 7% Generated emails
Templates, real fields onlyGeneration, real hallucination risk
Templates can't state a false fact about a customer's account, they only ever pull data that's actually true. Generation, tested on the same accounts, genuinely can.

So Milo ran the real test: 200 accounts through both approaches. The template version pulled plan tier, company name, and the matching next-step link straight from account data, every time, correctly. The generated version wrote fresh, warm, personal-sounding prose, and on 14 of those 200 accounts, it described a feature the customer's actual plan didn't include.

The generated emails weren't worse writing. They were confidently, warmly wrong about something a template structurally cannot get wrong.

Here's the turn: the problem was never that generation writes badly. The turn is that "more personalized" and "more accurate" aren't the same axis, and this specific email never needed enough real variation to be worth the accuracy risk generation introduced.

Knowledge spark: why can't a template make this specific mistake? A template only ever fills in blanks from real account data, plan tier, company name, a link. It has no way to say something that isn't in that data. Generation writes new sentences from a pattern it learned, so it can produce something that sounds right and simply isn't true, especially about a specific customer's specific account.
Monthly cost as signup volume grows, template versus generation
$700 $350 $0 $80 $640 10k signups 30k 50k 80k
Template, near-zero costGeneration, climbing with volume
The template's cost doesn't move as Greetline grows. Generation's cost climbs in a straight line with every new signup, for an email nobody was asking to feel more personal in the first place.

At its worst, this cost showed up as a real support ticket: a Starter-tier customer emailed asking how to access the "advanced analytics dashboard" her welcome email had described, a feature that only exists on the Pro and Enterprise plans. The support team had to explain the email was simply wrong, an awkward first conversation with a brand-new customer.

The choice I would take back Assuming, without checking, that a generated welcome email would obviously beat a template because "personalized" sounds better than "templated." That assumption felt reasonable before anyone had actually audited what real personalization here would need to cover. It stopped holding up the moment the audit showed the real variation was four values, not an open-ended range.

What I would leave alone: Greetline's support reply drafting tool stays fully generative, every customer message is genuinely different, and no fixed set of templates could cover the real range of things people actually ask. The kill-criteria test says generation there, clearly.

The lesson: "personalized" is not automatically better than "templated." Count the real variation first. If it's small enough to write down, a template will be cheaper, faster, and structurally incapable of the one mistake that matters most.

Now here is the same thing as a story

The short version above is what you'd say out loud in the room. Read this one for what it actually felt like to watch a warmer-sounding email lose to a plainer one.

Zohra Petrakis had shipped onboarding flows for three years, and her instinct leaned, like most product people's, toward "more personal is better." When a generation-based rewrite of FirstTouch came up as a roadmap idea, it felt like an obvious upgrade, barely worth debating.

Milo, the engineer building it, asked one question before starting: "What actually varies across the emails we already send?" Nobody in the room had a real answer. It had just never been asked.

Hand sketched comparison titled The asymmetry, drawn. Left panel, a document icon labeled template's worst case, caption a slightly generic line, cheap, absorbed. Right panel, a question mark box icon labeled generation's worst case, caption a feature the customer doesn't actually have.
Not two equally-sized risks. One side's worst case is a shrug. The other's is a support ticket and a shaken new customer.

He pulled 1,000 real welcome emails sent over the past year and read through a sample by hand, looking for what genuinely changed message to message. Plan tier. Company name. Which of three account reps had signed the email. One next-step link, tied to the plan. That was it. Four things, four real values for the biggest one.

He built the generation pilot anyway, since the team wanted a real comparison, not a hunch. It read beautifully. Warm, specific-sounding, genuinely well written prose, for every single one of the 200 test accounts.

Hand sketched comparison titled Same customer, two emails. Left panel, a document icon labeled template email, caption pulls only real account fields. Right panel, a question mark box icon labeled generated email, caption invents an analytics feature she doesn't have.
The same real customer, the same real account, two completely different claims about what she actually has access to.

Then he checked the claims against real account data, line by line. Fourteen of the 200 generated emails, about 7 percent, described something the customer's plan simply didn't include. One of them was the "advanced analytics dashboard" line, sent to a Starter-tier customer whose plan had never had that feature.

We weren't comparing good writing to plain writing. We were comparing a system that can only tell the truth to one that can, confidently, tell you something false about your own account.

Zohra had assumed, walking in, that this would be a close call, a tradeoff between personality and safety worth debating. It wasn't close. Once the real variation was counted and the real error rate was measured, the decision made itself.

Hand sketched decision tree titled Template or generation. Root node, can a small set of templates cover the real range of cases. Three branches. Yes, variation is small and structured leads to build the templates. No, variation is genuinely open ended leads to generation earns its cost. Unsure, never actually audited leads to audit real cases first.
The fork that should run before any "personalized versus templated" debate even starts.

Back before the audit, assuming generation would obviously be better wasn't an unreasonable instinct. "Personalized" really does often beat "generic" for other kinds of content. It stopped being a safe assumption the moment someone actually counted what personalization here would need to cover, and found the honest answer was almost nothing.

Hand sketched labeled parts diagram titled Why templates stay cheap forever. A document icon at the center labeled Four Templates, with four labeled callouts around it: Built once, tested once. No drift to watch. No eval set to refresh. Cost stays near zero at scale.
None of these four things are true of the generated version, and none of them needed to not be true here.

Here's the replay: four templates, one per plan tier, tested once against real account data and shipped. Zero factual errors across the next 10,000 sends. No drift to monitor, no eval set to refresh, no recurring inference bill climbing as Greetline grows.

One version of this story spends real, recurring money making an email marginally warmer while quietly risking a false claim about a customer's own account. The other spends a single afternoon building four templates that get it right every time, forever, for free.

What I'd tell myself, reading that fourteenth wrong email: "personalized" was never actually the question. The question was always how much the message really needed to vary, and nobody had bothered to count until Milo did.

PICK: the test that decided FirstTouch's welcome emailNot a script for defending a template because it feels safer. PICK is what makes you count the real variation before either option gets picked on vibes.

P
Position. What's the real distinction between the two options?
A template wins when the real variation is small and structured enough to write down in advance. Generation earns its cost when the message genuinely needs to vary in ways no fixed template set could cover.
Not a preference between the two. A test of which shape the actual problem has.
I
Impact. What's lost picking generation where a template would do?
A real hallucination risk, 7 percent of generated emails stated something false about the customer's specific plan, plus real, ongoing inference cost for content that never needed to vary that much.
Not a hypothetical risk. Measured, at 14 wrong emails out of 200 real accounts.
C
Cost asymmetry. Which mistake is cheap, and which is expensive?
A template is cheap to build, test once, and trust forever for genuinely template-shaped content. Generation adds real, recurring cost and real hallucination risk for content that was never going to need that flexibility.
This is the direct answer's real mechanism: the two mistakes aren't the same size.
K
Kill criteria. What's the one test that actually decides it?
Can you write down a small number of template variants that cover the real range of cases well? Here, four templates covered 1,000 real accounts completely. That's the kill criteria, answered with a real audit, not a guess.
This is the direct answer to the question, turned into a test you could run on any similar decision.

The recap, one line per letter: position separates a small, structured range from a genuinely open-ended one, impact is the real hallucination risk and recurring cost of picking generation where a template would do, cost asymmetry says which mistake is cheap and which is expensive, and kill criteria is the countable test, how many templates would it take, that actually settles it.

And if you want to be sure it really works, try it somewhere elseSame four letters, a legal document generator instead of a welcome email. This time the real range genuinely was too open for a template.

Corwin Aldas runs product at Clauseform, a tool that drafts vendor contracts. A teammate proposed replacing the drafting step with five fixed contract templates, one per deal type, to avoid any generation risk at all. Mapped onto PICK: position asked what actually varies across real past contracts, and the audit found genuinely open-ended variation, custom payment terms, deal-specific liability clauses, negotiated exclusivity language, none of it reducible to five templates without either missing real clauses or forcing every deal into an awkward, wrong-shaped mold. Impact of forcing a template here would be worse than generation's risk: a contract missing a clause the deal actually needed is a real legal exposure, not a cosmetic mismatch. Cost asymmetry flips from the welcome email case: here, generation's ongoing cost and review overhead is worth paying, because the alternative, a template stretched over cases it was never built for, is the expensive mistake. Kill criteria: could five, or even fifteen, templates cover the real range well? The audit said no, so generation, with a human legal reviewer in the loop, was the right call.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "count how many templates it would take to cover the real cases, and let the count decide," and stop.
Cost: no time to audit real past cases before the meeting. Say so honestly, and propose the audit as the very next concrete step instead of guessing which way it would go.
The model got better, for real: say a future model's hallucination rate drops to near zero. Rerun the kill-criteria test anyway, because the real question was never generation's quality, it was whether the underlying variation was small enough to template in the first place.

Where people run it wrong.
They assume "personalized" or "generated" is automatically better than "templated," without ever counting the real variation.
They price generation's build cost but forget its recurring inference cost and its structural hallucination risk.
They force a template onto content with genuinely open-ended variation, producing an awkward, wrong-shaped result instead of admitting generation was the right tool.

How to use it live. The moment an interviewer asks whether something should be templated or generated, ask yourself first: how many templates would it actually take to cover the real range of cases. That question buys real thinking time, and it's usually exactly where the honest answer is hiding.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits deciding when a template beats generation?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It turns "which feels better" into a countable test against real past cases.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Zohra Petrakis, the PM deciding FirstTouch's approach at Greetline. Milo Reznik is the engineer who ran the real audit and pilot that settled the decision.
3 · THE TEST
What's the one concrete thing this answer says to actually do?
Tap to flip
ANSWER
Audit real past cases and count how many templates it would take to cover them well. A small, countable number means build the templates.
4 · THE RISK
What's the specific risk generation introduced that a template structurally can't have?
Tap to flip
ANSWER
Stating a false fact about a specific customer's account, like describing a feature their plan doesn't include. A template only pulls real data fields, so it can't make that mistake.
5 · THE OLD DECISION
What old assumption would this answer take back?
Tap to flip
ANSWER
Assuming a generated welcome email would obviously beat a template because "personalized" sounds better, without ever auditing what real personalization here would need to cover.
6 · THE NUMBER
Fill in the blank: templates made ___ factual errors out of 200 test emails. Generation made ___.
Tap to flip
ANSWER
Zero, versus 14, a 7 percent error rate. Templates can't state a false fact about an account. Generation, tested on the same real accounts, genuinely did.
7 · THE COST
How does the cost gap between the two options change as Greetline grows?
Tap to flip
ANSWER
The template's cost stays near zero regardless of volume. Generation's cost climbs in a straight line with every new signup, from about $80 a month at 10,000 signups to $640 a month at 80,000.
8 · CROSS PRODUCT TRANSFER
Section 4 runs PICK again on a different product. Which one, and which way does it flip?
Tap to flip
ANSWER
Clauseform's vendor contract drafting tool. It flips the other way: real variation across contracts was too open-ended for templates, so generation, with a human legal reviewer, was the right call.

Check yourself Score: 0 / 0

Short answer, apply it yourself
1. Think of a piece of automated or AI-written content you've received, an email, a notification, a summary. Would a template have covered its real range of cases, or did it genuinely need to vary?
Show hint
Think about how many genuinely different versions of that message you'd realistically need to cover most real cases.
Show answer
Model answer: A shipping notification email usually only varies by carrier, tracking number, and delivery estimate, a handful of template variants, not genuine open-ended content that needs generation.
Multiple choice
2. Why did the generated welcome email describe a feature the customer's plan didn't include?
  • A. The template had an outdated feature list.
  • B. Generation writes new sentences from a learned pattern, which can produce something plausible-sounding that isn't actually true.
  • C. The customer had recently downgraded their plan.
  • D. Milo configured the pilot incorrectly.
Show hint
Look at the knowledge spark in "Let's learn."
Show answer
B. A template only fills in real account fields and structurally can't state something false. Generation writes fresh prose from a pattern, which can sound right without actually being true.
Fill in the blank
3. Fill in the blank: the audit of 1,000 real welcome emails found the message only needed to vary by ___ things.
Show hint
Look at the icon list in "Let's learn."
Show answer
Four things. Plan tier, company name, sender, and one next-step link, small and structured enough to cover with four templates.
True or false
4. True or false: this answer argues that Greetline should stop using AI generation anywhere in its product.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The support reply drafting tool stays fully generative, since real customer messages genuinely vary too much for any fixed template set to cover.
Short answer, where it wouldn't matter
5. Name a feature at Greetline where the kill-criteria test would point toward generation instead of templates, and say why.
Show hint
Look at "what I would leave alone" and the cross-product example in Section 4.
Show answer
Model answer: Support reply drafting. Real customer messages vary too openly for a fixed set of templates to cover well, so generation earns its cost there, unlike the welcome email.
Short answer, work the number
6. If the audit had found 40 real template variants were needed instead of 4, would the direct answer still favor templates?
Show hint
Think about what "small enough to write down" actually means as a number gets larger.
Show answer
Model answer: Probably not as clearly. Forty variants starts to strain the kill criteria's "small and structured" test, and the maintenance burden of keeping forty templates correct could start to rival generation's own ongoing cost, worth re-checking case by case rather than assuming templates still win automatically.
Before you close the answer
Why this works
Tests whether you'll count real variation before choosing a tool, instead of defaulting to whichever option sounds more advanced, and whether you know templates and generation carry structurally different risk profiles, not just different costs.
Follow-up traps
"Couldn't you just add a fact-checking step to the generated emails?" Response: you could, but that adds real ongoing review cost to fix a problem four free templates never have in the first place, worth doing only if generation was already justified by real variation.

"Isn't a template less impressive to customers?" Response: the pilot's own generated emails were well written and still got the facts wrong on 7 percent of accounts, "impressive" and "accurate" turned out to be two different axes entirely.
If pressed
The four templates were reviewed and re-approved every time a plan tier's actual feature list changed, a lightweight, infrequent check, nothing like the ongoing drift monitoring a generation-based approach would have needed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more