CaseAdvancedDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #11
How would you teach prompt technique without calling it prompt engineering?
ORDER the product is ListCraft, an AI property-description generator used during leasing walkthroughs
Milltide Properties equips leasing agents with ListCraft, which turns quick walkthrough notes into a finished listing description. Foster Amadi carries a tablet from unit to unit and taps out notes between showings, usually with one hand while holding a clipboard in the other.
The direct answer
Build the teaching into the note-taking screen itself, not a help article. Start with the cheapest, most reversible move first: a suggestion chip that appears the moment a note is too thin, showing exactly what a stronger note looks like. Save the bigger, harder-to-undo moves, like the AI asking its own clarifying question, until you've proven the cheap one actually works.
Do this, in order
Ship inline suggestion chips first, on the note-taking screen itself.Why: it's the cheapest, fastest, most reversible way to teach, and it's where the thin input actually happens.
Hold off on the AI asking its own clarifying question until the chips prove out.Why: once agents get used to being asked, pulling that back feels like the tool got dumber, so it's the hardest of the four to undo.
Don't build the saved detail-template library until you have real data on what "good" looks like per property type.Why: it depends on examples the earlier two moments haven't generated yet.
Test the cheapest move on a slice of agents before committing a whole quarter to the bigger one.Why: you can learn whether teaching-by-suggestion even works before building the harder, stickier version.
Leave agents who already write detailed notes alone.Why: the chips barely fire for them anyway, and more prompting would just be noise.
How to answer this, stage by stageSix moves, in the order you'd actually say them.
Stage 1
Scope it to one real screen
Say it like this
"I'll answer this for ListCraft, a listing-description generator leasing agents use on a tablet during walkthroughs."
Why this works
Grounds "teach prompt technique" in a specific screen instead of a training curriculum.
Stage 2
Say your structure out loud
Say it like this
"I'll use ORDER. Outcome, what we're all competing to move. Reversibility, what's hardest to undo. Dependency, what unblocks what. Evidence, what we can learn cheaply. Rank, the actual order."
Why this works
Signals this is a prioritization call between real candidates, not a single idea presented as the only option.
Stage 3
Reframe the question
Say it like this
"'Teach prompt technique' really means: get agents to hand the model more specific, checkable detail, without ever making them read a manual about how to talk to AI."
Why this works
Separates the real goal from the trap of building a training module nobody will open.
Stage 4
Give the one decision
Say it like this
"Ship the suggestion chip first: the moment a note is thin, it shows what a stronger note would include, right there on the note screen, with one tap to accept it."
Why this works
Concrete, cheap, and the top of the actual rank order.
Stage 5
Name what's hardest to undo
Say it like this
"The riskier move is having the AI ask its own follow-up question. Once agents get used to that back-and-forth, removing it later feels like a downgrade, so I'd only build it after the cheaper chip has proven the idea works."
Why this works
This is the hard step in ORDER: reversibility, not just impact, decides the order.
Stage 6
Close on the one line
Say it like this
"So the teaching isn't a lesson. It's a habit built one small, reversible nudge at a time, starting with the cheapest one."
Why this works
Restates the decision and the reasoning in one breath.
Let's learn
ListCraft reads an agent's walkthrough notes and drafts a full listing description in about fifteen seconds, ready to paste into the leasing site.
Before it, Foster wrote every description himself, roughly twelve minutes per unit, working from memory and whatever he'd scribbled on a clipboard between showings.
Knowledge spark: what's a "detail score"?
A simple internal measure of how specific a note is: does it name a real feature, a number, a material, instead of a vague word like "nice" or "spacious." A higher score usually means a stronger finished description.
Milltide's first attempt at teaching agents to write better notes was a help-center article titled "Get better results from ListCraft." It listed five tips: be specific, mention square footage, name finishes, and so on.
Four real candidates for where the teaching could live. A help article isn't even on this list, because almost nobody opened it.
And the turn is this: the problem was never that agents didn't know what a good description needed. It's that nobody reads a tips article between showings with a clipboard in one hand.
The suggestion chip wins on both axes at once. That's rare, and it's exactly why it ranks first.
At its worst: Milltide skips straight to the AI asking its own clarifying question on every thin note, agents get used to the back-and-forth within weeks, and then a later redesign that tries to remove it for being too slow gets treated internally as a downgrade, even though the actual detail scores never improved enough to justify keeping it.
The choice I would take back
We wrote a help-center tips article instead of touching the note-taking screen itself, because writing an article was fast and didn't require any product changes. That made sense as a quick first move. It stopped making sense once we saw only six percent of agents ever opened the help center in their first month.
What I would leave alone: agents who already write dense, specific notes out of habit barely trigger the chip at all. Building more prompting on top of an already-good note would just be noise for them.
We weren't failing to explain prompt technique. We were explaining it in the one place nobody was standing when it mattered.
The lesson: teaching someone to talk to a model works when it happens exactly where the thin input is being typed, not in a document they have to go find first.
Now here is the same thing as a storyThe short version is above. Read this for how the actual rank order got picked.
Foster has leased units at Milltide for six years and can size up a unit's best-selling feature within thirty seconds of walking in.
His notes, though, stayed short: "2bd, nice light, good closets." Fast to type between showings, and it had always been enough for him to remember what to say to a prospective tenant later.
This one small chip is doing the entire teaching job, right where the thin note gets typed.
The chip fired on his very next unit: "Try adding: year built, one standout feature, like the closets you mentioned." Foster tapped it, and his note became "2bd, nice light, walk-in closets, updated 2019, corner unit."
Nothing about the later two moments makes sense before the first one runs and generates real data.
The description ListCraft drafted from that note read noticeably sharper, and Foster noticed it read sharper without anyone ever telling him he'd just done something called "prompting." He just tapped a suggestion that felt like his own idea.
The chip can be turned off tomorrow with nobody noticing. The habit of being asked a question can't.
Three weeks in, Milltide's product team debated adding the bigger move: ListCraft asking its own follow-up question out loud whenever a note stayed thin even after the chip. It was tempting, since early tests showed it taught even faster.
They held off. The chip alone had already moved the average detail score from 2.1 out of 10 to 5.4, without asking agents to change how they worked at all. Committing to the harder-to-undo conversational version before proving that gap mattered would have locked in a habit nobody could easily walk back.
I would take back the tips article. It felt like the responsible thing to ship first, cheap and low-risk. It took watching six percent of agents ever open it to see that "low-risk" and "actually reaching anyone" are not the same thing.
ORDER, the four candidates rankedFive letters. The second one decides more than the first.
O
Outcome. What's being competed for.
Agents supplying richer, more specific input on their own, without ever reading anything called a prompting guide.
Without a named outcome, ranking is just opinion.
R
Reversibility. What's hardest to undo.
Suggestion chips can be turned off with nobody noticing. The AI's own clarifying question, once agents rely on it, is the one that feels like a downgrade to remove.
The hardest step, and the one that actually sets the order.
D
Dependency. What unblocks what.
The saved detail-template library needs real examples of strong notes by property type, which only exist once the chips and the clarifying question have been running for a while.
Some order is forced by reality, not preference.
E
Evidence. What's cheap to learn first.
Run the suggestion chip on a slice of agents for two weeks and check the detail-score lift before committing a full quarter to the conversational version.
Buys real data before the expensive, harder-to-reverse bet.
R
Rank. The actual order, defended.
Chips first, clarifying question second, example gallery third, template library last, because it's the only one that needs the others to exist first.
A ranking that would hold up even if you had to defend just the top pick.
Average detail score, by teaching moment in place
The chip alone did more than double the score. The riskier move added more, which is exactly why it's worth testing, just not worth building first.
Detail score, week by week across the rollout
Each new teaching moment moves the line at its own marker. Nothing about the second jump would make sense without the first one already proven.
The recap, one line per letter: outcome is agents handing over richer detail on their own, reversibility is the chip against the clarifying-question habit, dependency is the template library needing data the earlier two moments generate, evidence is the two-week slice test, and rank is chips, then question, then gallery, then templates.
And if you want to be sure it really works, try it somewhere elseSame five letters, a waste-sorting facility instead of a leasing office.
Cascadia Waste Recovery uses an AI tool that reads a route worker's photo and short note about a flagged load and drafts a contamination report. Gideon Okafor supervises a sorting line and reviews flagged loads between truck arrivals.
Mapped onto ORDER: outcome is workers supplying a specific, checkable note ("wet cardboard, bin 4, smells of solvent") instead of a vague one ("bad load"). Reversibility: a quick example prompt shown on the flagging screen is easy to remove later; an AI voice assistant that asks a follow-up over a headset, once workers rely on it mid-shift, would be much harder to pull back without it feeling like a real loss. Dependency: a shared library of "what contamination looks like" photo examples needs real flagged-load data before it's worth building. Evidence: test the on-screen example prompt on one shift before building anything voice-based.
Different floor, same order of operations: cheapest, most reversible nudge first, voice-based follow-up only once it's earned.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the cheapest, most reversible nudge, right where the thin input happens, before anything bigger," and stop.
Cost: no budget for a full clarifying-question feature this quarter. Say so honestly, and ship the suggestion chip alone first, since it's the whole method's cheapest, highest-leverage move anyway.
The model gets better, for real: if ListCraft's drafting quality improves overall, thin notes still produce thinner descriptions than detailed ones. A better model doesn't remove the need to teach better input, it just makes the gap between a good note and a lazy one less visible day to day.
Where people run it wrong.
They write a help article and call it done, without checking whether anyone under deadline pressure will ever open it.
They build the flashiest teaching moment first, instead of the cheapest reversible one, and get stuck defending a hard-to-undo habit before they know it's worth keeping.
They treat "teach prompt technique" as a training problem instead of a product-design problem that lives on the exact screen where thin input happens.
How to use it live. When someone asks how you'd teach a skill without naming it, ask yourself one question first: which of my candidate teaching moments could I undo tomorrow with nobody noticing. Start there.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "how would you teach prompt technique without calling it prompt engineering"?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. It ranks four candidate teaching moments instead of picking just one.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Foster Amadi, a six-year leasing agent at Milltide Properties who types notes on a tablet between showings.
3 · THE OUTCOME
What are all four candidate teaching moments actually competing to move?
Tap to flip
ANSWER
Agents supplying richer, more specific walkthrough notes on their own, without ever reading a guide about how to talk to AI.
4 · THE HARD STEP
Which candidate is hardest to undo, and why?
Tap to flip
ANSWER
The AI asking its own clarifying question. Once agents rely on that back-and-forth, removing it later feels like the tool got worse.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Writing a help-center tips article as the first move, since it was fast to ship but only six percent of agents ever opened it.
6 · THE NUMBER
Fill in the blank: average detail score went from 2.3 with the article alone to ___ once suggestion chips shipped.
Tap to flip
ANSWER
5.4. Adding the clarifying question on top pushed it to 7.8.
7 · THE REPLAY
Same thin note from Foster, chip shipped. What changes?
Tap to flip
ANSWER
The chip suggests adding year built and a standout feature. He taps it, and the note goes from four words to a specific, detailed one in seconds.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what stays the same about the order?
Tap to flip
ANSWER
Cascadia Waste Recovery's contamination-report tool. The cheapest, most reversible nudge still ships before any voice-based follow-up feature.
Check yourself Score: 0 / 0
Multiple choice
1. Why does this answer rank suggestion chips ahead of the AI asking its own clarifying question?
A. Because the clarifying question doesn't teach anything.
B. Because chips are cheaper and easier to undo, while the clarifying-question habit is hard to remove once agents rely on it.
C. Because agents dislike being asked questions.
D. Because chips require no engineering work at all.
Show hint
Look at the reversibility step and the quadrant diagram.
Show answer
B. ORDER ranks by what's hardest to undo, not just by which one teaches the most.
True or false
2. True or false: this answer recommends building the saved detail-template library before the suggestion chips.
True
False
Show hint
Look at the dependency step.
Show answer
False. The template library depends on real example data the chips and clarifying question generate first, so it ranks last.
Fill in the blank
3. Fill in the blank: only about ___ percent of agents ever opened the original help-center tips article.
Show hint
Look at "the choice I would take back."
Show answer
Six percent. Which is why the teaching moved onto the note-taking screen itself instead.
Short answer, where it wouldn't matter
4. Name a kind of agent for whom the suggestion chip barely matters.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An agent who already writes detailed, specific notes out of habit. The chip rarely fires for them, since their input is already dense.
Short answer, name the reversal
5. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Writing a help-center article instead of touching the product. It made sense as a fast, low-risk first move, until the data showed almost nobody opened it.
Short answer, apply it yourself
6. Pick a tool you use yourself that quietly taught you how to use it better. What was the smallest, most reversible nudge it gave you?
Show hint
Think about an autocomplete suggestion, an example placeholder, or a tooltip that changed how you typed something.
Show answer
Model answer: Many people describe a search bar's greyed-out example text teaching them what a good query looks like, the same cheap, reversible nudge this answer ranks first.
Before you close the answer
Why this works
Tests whether you'll rank real candidates by what's hardest to undo, rather than by which one sounds most impressive, and whether you can find where the teaching actually needs to live.
Follow-up traps
"Isn't the clarifying question just a better version of the chip, so why not skip straight to it?" Response: it teaches faster, but it's much harder to walk back once agents depend on it, so it's worth testing only after the cheaper move has proven the underlying idea works.
"What if an agent just ignores the suggestion chip entirely?" Response: that's fine, and expected for agents who already write detailed notes. The chip is built to fire only on thin input, not to nag everyone equally.
If pressed
The real pilot only showed the chip when a note scored under 4 out of 10 on Milltide's own detail-score model, so agents who already wrote strong notes almost never saw it at all.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.