Artifact critiqueIntermediateDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #7

Critique an onboarding that starts with a blank chat box.

AUDIT the artifact is GrainForge's first screen, an AI tool that turns a text description into a furniture cut list

Here is what a blank chat box promises silently, before a single word gets typed into it: that anything you type is a fair question, and whatever comes back is a fair answer. Dax Okafor runs a small furniture workshop and uses GrainForge, an AI tool, to turn a written description into a cut list.

The direct answer
A blank chat box is a bad first screen, not because chat is a bad interface, but because it makes a silent claim: anything you ask is fair game, and it will always look equally confident answering. Fix it by replacing the blank box with one real example request and one plain sentence naming what the tool won't attempt, before the cursor ever gets a chance to blink at someone.
Do this, in order
  1. Replace the blank box with one real example request, shown before typing starts.Why: it teaches the shape of a good ask faster than any instructions page.
  2. Name what the tool won't attempt, in the same breath as what it will.Why: a blank box never says where its edge is, so nobody knows they've crossed it.
  3. Show a confidence signal on the output, not just a clean-looking answer.Why: an out-of-scope answer looks exactly as polished as an in-scope one.
  4. Give a one-tap way to flag "this doesn't look right."Why: without it, a wrong answer only gets caught after the material's already cut.
  5. Don't fix this by adding a long instructions page nobody will read first.Why: the fix has to live in the one screen people actually see before typing.

How to answer this, stage by stage

Six stages. Say each one plainly, and the critique holds up under a real follow-up.

Stage 1
Scope it to one real screen
Say it like this
"I'll critique GrainForge's first screen specifically, a blank chat box a furniture maker sees the moment they open the app, before they've typed anything."
Why this works
Grounds an abstract "critique this pattern" prompt in one concrete artifact.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who built it and why. Uncover what it was tested against. Demand to see a real first session. Isolate what's missing. Test it myself."
Why this works
Signals a structured audit, not a list of vague design opinions.
Stage 3
Name the silent claim
Say it like this
"A blank box doesn't just fail to teach scope, it actively implies there is no scope. Every question looks equally welcome, and every answer looks equally confident."
Why this works
This is the sharpest move in an AUDIT critique: naming what the artifact implies, not just what it lacks.
Stage 4
Isolate what's missing
Say it like this
"No example request. No scope note. No confidence signal on the output. No way to flag a wrong answer before it turns into cut wood."
Why this works
Turns "this feels risky" into four specific, checkable gaps.
Stage 5
Prove it with a real test
Say it like this
"Dax's apprentice typed a request for a curved stair railing, well outside GrainForge's scope. It came back with a full, confident cut list anyway. The shop cut 600 dollars of hardwood before anyone caught it didn't fit."
Why this works
A real cost, in real dollars, turns the critique from a taste opinion into a measured failure.
Stage 6
Close on the fix
Say it like this
"Replace the blank box with one example and one scope sentence. That's the whole fix, and it costs a single screen, not a redesign."
Why this works
Ends on something buildable, not a vague call for "better onboarding."

Let's learn

Here is what a blank chat box promises silently, before a single word gets typed into it: that anything you type is a fair question, and whatever comes back is a fair answer.

GrainForge is an AI tool that reads a written description of a furniture piece and generates a cut list, joint types, board lengths, angles, ready to hand to a table saw.

Before GrainForge, Dax drew every cut list by hand, about forty minutes per piece for anything with real joinery. GrainForge does it in under a minute. That part works.

First-session outcome, within 5 minutes of opening the app
100% 50% 0 34% 9% Blank chat box Guided starter screen
A third of first-time users never sent a single message to a blank box. That's not patience, it's a screen that gave them nothing to hold onto.
Knowledge spark: what does "out of scope" mean here? A request GrainForge was never trained to handle well, like a curved architectural piece instead of flat furniture joinery. The tool doesn't know it's out of its depth, so it answers just as confidently as it would for a request it actually handles well.

At its worst: Dax's newest apprentice, alone at the screen for the first time, typed a request for a curved stair railing. GrainForge doesn't do architectural curves, only flat furniture joinery, but nothing on the blank box said so. It returned a full, clean-looking cut list anyway. The shop cut it. Six hundred dollars of hardwood, and the pieces didn't fit together at all.

The blank box never lied. It just never said there was anything it couldn't do, which is its own kind of lie.
The decision I would take back GrainForge's first screen fails silently on an out-of-scope request: it never says "this is outside what I'm built for," it just answers anyway, as confidently as it would for a request it handles well. That decision was never made on purpose, it was inherited from copying a general chatbot's blank-box convention because it looked clean in an investor demo. It stopped being harmless the day a real, unscoped request reached real wood.

What I would leave alone: for a request squarely inside GrainForge's scope, a simple cabinet joint, a table leg, the blank box works fine, the model is genuinely strong there. The fix isn't scaffolding every request, it's naming the edge before someone falls off it.

The lesson: a clean, minimal first screen isn't automatically a good one. Minimal and honest are two different design goals, and a blank box optimizes only for the first.

Now here is the same thing as a story

The short version above is the critique, argued plainly. This one is how Dax actually found the gap.

For the first three weeks, GrainForge's blank chat box felt like freedom. Dax could describe a piece in his own words, no menus, no dropdowns, and get back a cut list faster than he could draw one by hand. By the fourth week, it felt like a dare.

Hand sketched metaphor scene titled A blank page, or a workbench set up for you. Left, a question mark box icon labeled BLANK BOX, caption a cursor, nothing else. Right, a box icon labeled SET UP FOR YOU, caption tools laid out, scope named.
A real workbench tells you, just by looking at it, what it's set up to build. A blank cursor tells you nothing at all.

Dax had used GrainForge himself for weeks before letting anyone else near it, and he already knew, mostly by trial, the kinds of requests that worked well. His newest apprentice didn't have that trial behind her. She sat down at the screen for her first real task and saw exactly what Dax had seen on day one: a blinking cursor, and nothing else.

Hand sketched labeled parts diagram titled What's missing from the blank box. Center question mark box icon labeled Empty First Screen. Four callouts: no example request, no scope note, no confidence signal, no fallback path.
Four things missing from one screen. Any single one of them would have caught what actually happened next.

She typed a request for a curved stair railing, a piece she'd built by hand before at a previous job. GrainForge doesn't handle architectural curves, only flat furniture joinery, but she had no way of knowing that. Nothing on the screen had ever mentioned a boundary existed.

Hand sketched quadrant titled Sorting a first message by fit. Axes how far outside GrainForge's scope, and how confident the output looks. Flat cabinet joinery and simple table legs sit top left, in scope and confident. Curved stair railing and architectural arch sit top right, out of scope but just as confident.
The danger zone isn't low confidence. It's the top right corner, where an out-of-scope request looks exactly as sure of itself as an in-scope one.

GrainForge came back with a full cut list, board lengths, angles, joint callouts, formatted identically to every request that actually worked. Nothing about it looked uncertain.

Hand sketched timeline titled Dax's own test of the blank box. Four milestones: apprentice's first try, a curved railing typed cold. Confident cut list, looked completely valid highlighted. Wood cut, 600 dollars wrong fit. Scope note added, to the first screen.
The second milestone is the one that mattered. A confident answer is exactly as dangerous as a wrong one when nothing distinguishes the two.

The crew cut the boards. Two days later, at assembly, the pieces didn't meet at a single joint. Six hundred dollars of hardwood, and two days of shop time, gone on a request the tool was never built to handle.

We did not lose 600 dollars of wood to a bad model. We lost it to a screen that never once suggested there was anything it couldn't do.

Hand sketched flow diagram titled Auditing a blank chat box, five steps. Five steps: who built it why, what it was tested on, what a first session shows, what's missing highlighted, test it yourself.
Step four, isolating what's missing, is where this critique actually lands. Everything before it is just getting there.

Dax added one line above the box the next week: an example request, and one sentence naming what GrainForge doesn't attempt. He didn't touch the model at all.

Hand sketched icon list titled What a good first screen shows instead. Four items: a document icon labeled one real example request, a scale icon labeled what it won't attempt, a gauge icon labeled how sure the output is, a person icon labeled a way to flag it's wrong.
None of these four touch the model underneath. All four touch what a first-time user is told before they ever type a word.

I copied the blank-box pattern because it looked clean, modern, the way every popular chatbot looks. It took watching real material get cut wrong to notice that "looks like every other chatbot" and "is honest about its own edges" were never the same design goal.

AUDIT, applied to one screenNot a style critique. AUDIT is what forces the question of what a design silently claims, not just what it looks like.

A
Ask who built it, and why.
A small team, optimizing the first screen to look sleek in an investor demo, not tested against a real first-time furniture maker.
Names the incentive behind the design, not just the design itself.
U
Uncover what it was tested against.
Only the founding team's own requests, people who already knew the tool's scope from building it.
A design never tested on a real novice has no evidence behind its simplicity.
D
Demand to see a real first session.
An apprentice, cold, alone, nothing on screen but a cursor and no memory of what worked before.
The single most revealing thing you can ask for in any onboarding audit.
I
Isolate what's missing.
No example, no scope note, no confidence signal, no flag-it path. Four specific gaps, not a vague sense of risk.
The hardest step, and the one that turns "this feels risky" into something you can fix.
T
Test it yourself.
Dax's own apprentice's real request exposed the gap, on real material, at a real cost.
A critique that stops at reading the screen is a critique. One that tests it is evidence.
Material waste cost per month, before and after the scope note
$700 $350 0 $640, month 3 scope note added Month 1 Month 6
One spike, one fix. Waste settled near a tenth of the incident's cost within two months of adding a single sentence.

The recap, one line per letter: ask is a demo-optimized first screen, uncover is that it was only tested by people who already knew its scope, demand is a real apprentice's cold first session, isolate is the four missing pieces, and test is the 600-dollar proof the gap was real.

And if you want to be sure it really works, try it somewhere elseSame five letters, a home energy-audit app instead of a furniture shop. This time the missing scope note is about safety, not wood.

Toma Radulescu is a homeowner using WattScope, an AI app that answers questions about home energy efficiency from a blank chat box of its own.

Mapped onto AUDIT: ask is who built WattScope's first screen, a small team that copied the same blank-box convention for the same reason, a clean demo. Uncover is that it was tested only by energy-efficiency staff, never a homeowner asking something genuinely unsafe, like rewiring a breaker panel themselves. Demand is watching what happens the first time someone types exactly that. Isolate is the same four gaps: no example, no scope note, no confidence signal, no path to a licensed professional. Test is trying it yourself with a real unsafe request and seeing what comes back with the same unearned confidence as a safe one.

Hand sketched decision tree titled WattScope, sorting a homeowner's first message. Root, first message received, branching to four outcomes: clearly in scope leads to answer directly, vague missing details leads to ask one question, outside scope leads to say so plainly, safety-related leads to point to a licensed pro.
The fourth branch is the one a blank box has no way to reach. Nothing routes a genuinely unsafe question anywhere but a confident answer.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "one example, one scope note, before the box," and stop.
Cost: there's no budget for a redesign this quarter. Say so honestly, and start with a single sentence above the existing box, since even one honest line beats a silent one.
The model gets better, for real: if GrainForge's accuracy on architectural curves eventually improves enough to genuinely support them, the scope note still needs updating, not removing, the same problem returns the moment scope shifts again.

Where people run it wrong.
They treat a blank box as a neutral, minimal choice, when it's actually a specific claim that nothing is out of bounds.
They test onboarding only with people who already know the tool's scope, which hides exactly the failure a real novice would hit.
They fix a discovered gap with a help page or FAQ link instead of a change to the one screen people actually see first.

How to use it live. When someone hands you an onboarding screen to critique, ask yourself one question first: what does this screen silently claim about what's in scope, just by what it does and doesn't show. Name that claim before you say anything about how clean it looks.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "critique an onboarding that starts with a blank chat box"?
Tap to flip
ANSWER
AUDIT: ask, uncover, demand, isolate, test. Isolate is the hardest step, turning a vague sense of risk into four specific gaps.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dax Okafor, who runs a small furniture workshop and uses GrainForge to generate cut lists from written descriptions.
3 · THE SILENT CLAIM
What does a blank chat box implicitly claim?
Tap to flip
ANSWER
That anything typed into it is a fair question, and every answer will look equally confident, whether the request is in scope or not.
4 · WHAT'S MISSING
Name the four things missing from GrainForge's blank box.
Tap to flip
ANSWER
An example request, a scope note, a confidence signal on the output, and a one-tap way to flag a wrong answer.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting the tool fail silently on an out-of-scope request instead of naming the edge, a decision inherited from copying a generic chatbot pattern, not made on purpose.
6 · THE NUMBER
Fill in the blank: the curved-railing incident cost the shop about ___ dollars in wasted hardwood.
Tap to flip
ANSWER
600 dollars, plus two days of shop time, on a request GrainForge was never built to handle.
7 · THE FIX IN ACTION
After the scope note was added, what changed?
Tap to flip
ANSWER
Material waste cost dropped from a 640-dollar spike to roughly a tenth of that within two months, without any change to the underlying model.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the higher-stakes version of the gap?
Tap to flip
ANSWER
WattScope, a home energy-audit app. There, the same missing scope note means a genuinely unsafe request, like rewiring a panel, gets the same confident answer as a safe one.

Check yourself Score: 0 / 0

Multiple choice
1. According to this answer, what's the real problem with GrainForge's blank chat box?
  • A. Chat interfaces are inherently a bad design choice for any AI product.
  • B. It never signals that some requests are out of scope, so an out-of-scope answer looks exactly as confident as an in-scope one.
  • C. GrainForge's underlying model is too inaccurate to be usable.
  • D. The chat box loads too slowly for new users.
Show hint
Look at the direct answer and the quadrant diagram.
Show answer
B. The curved-railing request was well outside scope but came back looking exactly as polished as any in-scope answer.
True or false
2. True or false: this answer's fix required retraining GrainForge's underlying model.
  • True
  • False
Show hint
Look at "the choice I would take back" and the icon list of what a good first screen shows.
Show answer
False. The fix was one example request and one scope-note sentence added to the first screen, no model changes at all.
Fill in the blank
3. Fill in the blank: with the blank chat box, ___ percent of first-time users sent no message at all within five minutes.
Show hint
Look at the bar chart of first-session outcomes.
Show answer
34 percent. That dropped to 9 percent once a guided starter screen with one example and a scope note replaced the blank box.
Short answer, where it wouldn't matter
4. Name a request type where GrainForge's blank box works fine as-is, and why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A request squarely inside scope, like a simple cabinet joint or table leg. GrainForge handles those well, so the fix is naming the edge, not scaffolding every request equally.
Short answer, apply it yourself
5. Think of an AI chat product you've used that opened with a blank box. What's one thing you asked it early on that turned out to be outside what it could really do well?
Show hint
Think of a time a chatbot answered confidently but the answer turned out to be wrong or made up.
Show answer
Model answer: Most people can recall asking a chatbot something oddly specific, a real citation, a local business detail, and getting a confident, plausible, wrong answer, exactly the gap this critique is about.
Before you close the answer
Why this works
Tests whether you can look past how clean a screen looks and name what it silently implies about scope, and whether you can propose a fix that lives in the design rather than in a document nobody reads first.
Follow-up traps
"Isn't a scope note just going to make users feel limited and use the product less?" Response: the data says the opposite here, first-message rate rose once people had something concrete to react to instead of a blank cursor with no boundary in sight.

"Couldn't you solve this by training the model to say 'I'm not sure' more often instead of changing the screen?" Response: that helps, but it's a model-level fix on a slower timeline. A scope note ships today and catches the exact failure that already cost real material.
If pressed
GrainForge's real fix also tags every generated cut list with the specific joint types it used, so a woodworker can spot-check whether the output matches a technique they actually recognize, even before the scope note existed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more