Artifact critiqueIntermediateResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #16

Critique an internal tool built with no user research because the users sit nearby.

GUARD the scenario: Loom & Lattice Architecture, a mid-size firm whose internal structural-detail-checking plugin was tested only with the architects down the hall

Interviewer's question: "Critique an internal tool built with no user research because the users sit nearby." Renata Szabo built a structural-detail-checking AI plugin for Loom & Lattice Architecture's CAD software, testing it only with the senior architects who work down the hall from her at the main studio.

The direct answer
"The users sit nearby" is not a research pass, it's a sampling bias wearing the clothes of convenience. The people down the hall were never a stand-in for everyone who'd actually use this tool, they were just the easiest people to ask. The real cost lands on whoever's workflow genuinely differs and never got a say, and it shows up quietly, as a usage rate that craters in exactly the group nobody thought to include.
Do this, in order
  1. Name who the tool was actually tested with, and who it wasn't.Why: "the users" tested were really one office, one seniority band, and treating that as representative is the whole mistake.
  2. Find where the untested group's real workflow differs, not just assume it doesn't.Why: proximity bias fails unevenly, on exactly the workflows the nearby group never had a reason to think about.
  3. Run one real, recorded fifteen-minute session with someone from every office and seniority band.Why: a hallway chat with whoever's around isn't research, it's convenience mistaken for coverage.
  4. Track adoption by office and seniority, not as one company-wide number.Why: a healthy blended adoption rate can completely hide one office's usage quietly dying.
  5. Give the untested group an actual way to flag "this doesn't fit," not just a feedback box.Why: without a real lever, their only option is to quietly stop using the tool, which teaches you nothing until it's too late.

How to answer this, stage by stage

Nobody is grading whether you can say "do more user research." They're grading whether you can name exactly whose workflow got left out, and how you'd know before it costs you.

Stage 1
Scope it to one real tool and one real firm
Say it like this
"I'll critique Renata Szabo's structural-detail plugin at Loom & Lattice Architecture, tested only with the senior architects at the main studio."
Why this works
Grounds an abstract bias question in one concrete, checkable tool.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, unequal impact, ability to contest, reduce, detect."
Why this works
Signals a method built for exactly this shape of question, not a general "be more inclusive" answer.
Stage 3
Name both groups plainly
Say it like this
"The tool was tested with senior architects down the hall from Renata. It was never tested with junior drafters, or with the satellite office in a different city."
Why this works
The G step. Naming both sides plainly is what separates a real critique from vague concern.
Stage 4
Say where the assumption fails unevenly
Say it like this
"The satellite office does historic-preservation detailing with a completely different annotation style. The main studio never encounters that workflow, so the tool was never built to recognize it."
Why this works
Shows the harm isn't generic, it lands specifically on the workflow the nearby group never had a reason to consider.
Stage 5
Name who can't push back, and why
Say it like this
"A satellite-office architect has no lever at all. The tool shipped 'done' based on approval from whoever happened to be down the hall, and nobody two floors away, in a different city, was ever asked."
Why this works
This is GUARD's strongest move, and the heart of the critique.
Stage 6
Give the concrete design reduction
Say it like this
"A real, recorded fifteen-minute workflow-shadowing session with at least one person from every office and every seniority band, before research ever gets called done."
Why this works
This is the direct answer, a concrete product decision, not a policy or a training session.
Stage 7
Say how you'd catch it in production
Say it like this
"Track adoption by office and seniority. If the satellite office's usage rate quietly craters while the main studio's stays healthy, that's the exact segment nobody asked, showing up before anyone files a complaint."
Why this works
The D step. Detection has to be specific enough to find the actual gap, not just "monitor usage."

Let's learn

Picture a tool that passed every test its builder ever ran on it, and still failed the first week it left the building.

Loom & Lattice Architecture's structural-detail-checking plugin sits inside their CAD software and flags likely errors in a drafted detail, a missing dimension, an unsupported span, before it goes to review. Renata Szabo built it and tested it entirely with the senior architects who work down the hall from her at the main studio.

Knowledge spark: what is proximity bias, in plain terms? Building something around whoever's easiest to reach, and mistaking their approval for everyone's approval. It isn't malicious. It's just what happens when "ask the users" quietly becomes "ask the users who are already in the room."

By every measure Renata had, the tool was a success. Main-studio adoption was strong, feedback was glowing, and nobody down the hall had a single complaint.

Hand sketched comparison diagram titled Two people, one lever. Left panel, a person icon labeled Main studio architect, caption asked, tested with, has the lever. Right panel, a question mark icon labeled Satellite office architect, caption never asked, no lever at all.
Same tool, same company. Only one of these two people ever got a say in how it works.

The turn: "no complaints" was never proof the tool worked everywhere. It was proof the tool worked for the one group who'd been consulted, and everyone else had simply never been given a way to say otherwise.

Who can't push back, and why A drafter at Loom & Lattice's satellite office, doing historic-preservation detailing with a completely different regulatory annotation style, has no lever at all. The tool shipped "done" the moment the people down the hall approved it. Nobody two floors away, in a different city, was ever in the room to say it didn't fit.

At its worst: the satellite office quietly stops using the plugin altogether, reverts to the old manual detail-check process, and nobody at headquarters notices for months, because a tool nobody uses generates no complaints either.

What I would leave alone: the plugin's core detail-checking logic for standard, non-preservation work doesn't need to change at all. The gap isn't in what the tool does well, it's in who it was ever tested against.

The lesson: "the users sit nearby" describes a sample of convenience, not a research pass, and the two look identical right up until someone two floors away actually tries to use the thing.

Now here is the same thing as a story

The short version above is what you'd say in a design review. Read this one for how the gap actually surfaced.

For three months, Renata Szabo's plugin was the best-reviewed internal tool the firm had ever shipped. Then someone two floors away, in a different city, actually tried to use it.

Hand sketched timeline titled How the gap was found. Four milestones: plugin ships month one, main studio adopts month two, new hire's question month three highlighted, satellite usage checked found the gap.
Three quiet months of glowing reviews, and it took one new hire's honest question to end them.

A newly hired architect at the satellite office, still new enough to ask questions nobody else thought to ask anymore, raised her hand in a training session: "Why doesn't it recognize any of our preservation annotations? It just flags all of them as errors." Renata didn't have an answer, because she'd never once watched anyone at that office actually draft a detail.

The tool wasn't broken. It had simply never met the people it was about to fail.

Renata pulled the adoption numbers by office that same afternoon. Main-studio usage sat at 91 percent, exactly as glowing as every internal report had said. Satellite-office usage sat at 12 percent, a number nobody had been tracking separately, buried inside a company-wide average that looked perfectly fine.

Plugin adoption rate, by office
100% 50% 0% 91% Main studio 12% Satellite office
One blended company-wide number would have shown a healthy 70 percent and hidden this completely.
Satellite office usage rate, week by week after launch
100% 50% 0% Main studio: flat Week 1 Week 12
Twelve straight weeks of quiet decline, and nobody was watching this line separately until a new hire asked why.

Renata's fix wasn't a bigger model or a longer feature list. She sat with two drafters at the satellite office for fifteen real minutes each, recorded, watching them work through an actual preservation detail from start to finish, and added a second annotation profile the plugin could recognize. Satellite usage climbed back above 70 percent within a month.

GUARD, in one screenNot a diversity checklist. GUARD is what makes an invisible research gap show up before production does.

G
Groups. Name both, plainly.
Senior architects at the main studio, tested and consulted. Junior drafters and the satellite office, never asked at all.
Naming both sides is what turns a vague worry into a real critique.
U
Unequal. Where the harm actually lands.
The satellite office's historic-preservation annotation style, a workflow the main studio never encounters at all.
The failure isn't random, it lands exactly on the workflow the nearby group had no reason to think about.
A
Ability to contest. The strongest move.
A satellite-office drafter had no lever. The tool shipped "done" based on approval from whoever was down the hall.
The hardest step, and the heart of the critique: someone's workflow was decided against without ever being asked.
R
Reduce. The concrete design change.
A real, recorded fifteen-minute workflow-shadowing session with someone from every office and seniority band.
A product decision, not a policy document or a training session.
D
Detect. How you'd catch it early.
Adoption tracked by office and seniority. The satellite office's usage cratering while the main studio stayed flat, visible weeks before anyone complained.
Shows the gap in production before someone external has to point it out.
Hand sketched icon list titled What a real research pass includes. Four items: one session per office location, one session per seniority band, a real recorded workflow not a chat, fifteen real minutes minimum.
Four requirements, and none of them are a hallway conversation with whoever's around.

The recap, one line per letter: groups is the main studio against the satellite office and junior drafters, unequal is the preservation-annotation workflow the harm actually lands on, ability to contest is a satellite architect with no lever at all, reduce is a real fifteen-minute shadowing session per office and band, and detect is adoption tracked separately by office.

And if you want to be sure it really works, try it somewhere elseSame five letters, a regional grocery chain instead of an architecture firm. Nothing else about the two jobs is alike.

Pinemont Grocers built an internal AI tool for shelf-stocking schedules, researched only with staff at the flagship downtown store, never with staff at the chain's rural stores.

Hand sketched quadrant titled Who was actually asked. Axes distance from the builder and seniority. Senior main studio sits close and senior. Junior main studio sits close and junior. Senior satellite office sits far and senior. Junior satellite office sits far and junior.
The two closest to the builder got asked. The two farthest away never did, regardless of seniority.

Mapped onto GUARD: groups is flagship downtown staff, tested and consulted, against rural-store staff, never asked at all. Unequal is that rural stores restock on a completely different weekly delivery schedule the flagship store never deals with. Ability to contest is a rural stocker with no lever, since the tool shipped "done" the moment downtown staff approved it. Reduce is a real recorded shadowing session at one rural store before calling research finished. Detect is tracking schedule-override rate by store type, since rural stores overriding the tool constantly is the exact tell nobody was watching for.

Hand sketched labeled parts diagram titled What the tool's research was actually built from. Center box icon labeled The plugin, with four callouts: main studio habits, senior architects only, no satellite input, no junior drafter input.
Four inputs, and every single one came from the same hallway.

Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "the users down the hall were a sample of convenience, not a research pass, and the real gap shows up as adoption cratering in the group nobody asked," and stop.
Cost: if a full research pass genuinely isn't affordable before launch, at minimum ship with adoption tracked separately by group from day one, so the gap surfaces in weeks instead of months.
The model gets better, for real: if the plugin's overall accuracy improves, the gap doesn't close on its own, since a better model trained on the same narrow group just gets more confidently wrong about the workflow it never saw.

Hand sketched flow diagram titled Where the appeal should be, and is not. Four steps: tool flags detail, drafter disagrees, no path back highlighted, tool wins by default.
Four steps, and the third one is the empty space where a real appeal should have lived.

Where people run it wrong.
They treat "nobody down the hall complained" as proof the tool works everywhere, instead of proof it works for the one group who got a say.
They watch one company-wide adoption number instead of breaking it out by the exact groups who might differ.
They build a feedback box instead of a real lever, so the group with the different workflow can only vote by quietly disappearing.

How to use it live. When someone tells you research wasn't needed because "the users sit nearby," ask yourself first: nearby to whom, and who does that leave out. Name that group specifically, and the critique writes itself.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits critiquing a tool built with no user research because the users sit nearby?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Built for risk and fairness questions, not AUDIT, which is only for judging a document's or report's trustworthiness.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Renata Szabo, who built Loom & Lattice Architecture's structural-detail plugin, testing it only with the senior architects down the hall.
3 · THE GROUPS
Who was tested with, and who was never asked?
Tap to flip
ANSWER
Senior architects at the main studio were tested and consulted. Junior drafters and the satellite office, doing historic-preservation work, were never asked at all.
4 · UNEQUAL IMPACT
Where does the assumption fail unevenly here?
Tap to flip
ANSWER
The satellite office's historic-preservation annotation style, a workflow the main studio never encounters and the tool was never built to recognize.
5 · THE DESIGN REDUCTION
What's the concrete fix, and what would you take back?
Tap to flip
ANSWER
Treating a hallway chat with whoever's nearby as a finished research pass. The fix: a real, recorded fifteen-minute shadowing session per office and seniority band.
6 · THE NUMBER
Fill in the blank: main-studio adoption sat at 91 percent, while satellite-office adoption sat at only ___ percent.
Tap to flip
ANSWER
12 percent. A company-wide blended average would have shown a healthy 70 percent and hidden the gap completely.
7 · THE REPLAY
Same gap, real fix. What changed?
Tap to flip
ANSWER
Renata added a second annotation profile after fifteen real minutes watching satellite drafters work, and their usage climbed back above 70 percent within a month.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different company. Which one, and who was left out there?
Tap to flip
ANSWER
Pinemont Grocers' shelf-stocking tool. Rural-store staff, on a different delivery schedule from the flagship downtown store, were never consulted.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: it took a ___ question during a training session for Renata to first learn the satellite office's workflow was never considered.
Show hint
Look at the trigger in the story section.
Show answer
New hire's. Someone new enough to ask a question nobody else thought to ask anymore.
Multiple choice
2. Why was "nobody down the hall complained" not real proof the tool worked?
  • A. Because complaints are always exaggerated.
  • B. Because it was only proof the tool worked for the group who had a lever to complain in the first place.
  • C. Because the main studio never actually used the tool.
  • D. Because complaints were disabled in the software.
Show hint
Look at "the turn" in Section 1.
Show answer
B. A group with no lever can't complain even when something genuinely doesn't work for them.
True or false
3. True or false: the plugin's core detail-checking logic for standard, non-preservation work was the actual problem here.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The core logic worked fine. The gap was entirely in who the tool was ever tested against, not in what it does well.
Short answer, apply it yourself
4. Think of a tool at your own company that was likely built around whoever was easiest to ask. Who might have been left out, and how would you find out?
Show hint
Think about a different office, shift, or seniority band that does the same job differently.
Show answer
Model answer: Something like a night-shift or remote team whose workflow differs from the daytime, in-office team the tool was actually tested with, findable by checking usage rates by shift or location.
Short answer, the number question
5. If satellite-office usage had stayed flat at 58 percent instead of dropping to 12 percent, would this still count as a real research gap? Why or why not?
Show hint
Think about what the decline itself, not just the raw usage number, was evidence of.
Show answer
Model answer: Possibly a smaller one. A flat 58 percent would still be well below the main studio's 91 percent, suggesting real friction, even without the dramatic decline that made the gap undeniable.
Short answer, name the reversal
6. What decision would you take back here, and why did it make sense at the time?
Show hint
Look at "the decision I would take back" language across the answer.
Show answer
Model answer: Treating approval from the people down the hall as proof of readiness. It made sense because they were the easiest, fastest group to test with, not because they were representative of everyone who'd use it.
Before you close the answer
Why this works
Tests whether you can name a specific excluded group and a specific harm, rather than giving a generic "do more user research" answer with nobody actually in it.
Follow-up traps
"Isn't a fifteen-minute session too short to catch everything?" Response: it's a floor, not a ceiling, and fifteen real, recorded minutes watching an actual workflow catches far more than zero minutes ever will.

"What if the satellite office is too small to justify separate research?" Response: size doesn't change whether their workflow differs, only how loud the resulting silence is when nobody asks.
If pressed
Loom & Lattice's actual usage-tracking dashboard now flags any office whose adoption rate falls more than 15 points below the company average for two consecutive weeks, routing it straight to Renata's team instead of waiting for a quarterly review to notice.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more