Artifact critiqueIntermediateQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #10

Design the daily dashboard for an AI feature team.

Design the daily dashboard to show who the aggregate number is hiding, not just how the aggregate number looks.

The direct answer
Give the daily dashboard a segment-sliced pre-qualify rate next to the aggregate, split by income documentation type, not instead of it. Put a fairness ratio, the weakest segment's rate over the strongest, on that same screen as its own daily number, with a floor that forces a same-day review. An aggregate that only describes the typical applicant is not a health metric, it is a way of not looking.
Do this, in order
  1. Put a segment-sliced pre-qualify rate on the dashboard, next to the aggregate, every single day.Why: an average can sit healthy for months while one real group's number collapses underneath it.
  2. Add a fairness ratio, the worst segment's rate over the best, as its own tile with a floor that forces same-day review.Why: a ratio nobody is watching is a number that only exists for a report nobody reads until it is too late.
  3. Track how many flagged applicants quietly leave the chat without finishing a full application.Why: FirstStep's soft message never becomes a formal denial, so a quiet exit is the only trace most of this harm leaves.
  4. Gate every new instant-decision model version on the same segment slice before it ships, not after.Why: the gap that hit Consuela's segment came from a retrained model that looked better on the aggregate and worse on one group nobody sliced for.
  5. Limit the daily segment tiles to groups where a real gap has actually been found.Why: fifty tiles nobody reads bury the one gap that is real, which is worse than showing none at all.
  6. Route a same-day alert to a person who can pause the model, not just to a chart in the corner.Why: a number that only moves a chart is decoration. A number that reaches someone who can act is a guardrail.

How to answer this, stage by stage

Nobody is grading whether you can name a fairness risk. They are grading whether you can turn it into a tile someone actually looks at tomorrow morning. Eight moves get you there.

1
Scope it to one real product, one real owner
Say it like this
"Let's ground this. Cazarrow Financial runs FirstStep, a chatbot that asks a few questions and gives an instant pre-qualify estimate for a personal loan. Emmett Vasilenko leads the small analytics team that owns FirstStep's daily dashboard."
Why this works
Grounds the design in a real product and a real owner before a single tile gets named.
2
Say your structure out loud
Say it like this
"I'd run this through GUARD. Name the groups affected, find where the harm lands unevenly, ask who actually gets to push back, then design the one change and the one alert that would catch it."
Why this works
Gives the interviewer a map in two seconds, before you name a single tile.
3
Name both groups, not just the one in the room
Say it like this
"Two groups touch this chatbot's output. Emmett's team reads a dashboard. The applicant reads one line of chat. I want to design for both, not just the one paying my salary."
Why this works
This is GUARD's first move, and skipping it is the most common way this kind of question turns generic.
4
Say plainly where the harm hides
Say it like this
"An aggregate pre-qualify rate can sit at fifty eight percent for nine weeks straight while one segment, self-employed and gig-income applicants, falls from forty one percent to nineteen. The average never moves enough to look wrong, because that segment is a small share of total volume."
Why this works
This is the reframe the whole answer turns on. A candidate who skips it just says "watch for bias" and moves on.
5
Name who can't push back
Say it like this
"FirstStep's message isn't a denial, it's a soft nudge: approval isn't likely based on what you told us today. Nobody has to send an adverse-action notice for a nudge. So the applicant who closes the tab right there never generates a complaint, a ticket, or a record anyone at Cazarrow will ever see."
Why this works
This is GUARD's hardest step, and the one that actually answers a risk question instead of describing one.
6
Give the one design decision
Say it like this
"So here's what I'd build. A segment-sliced pre-qualify rate on the daily dashboard, right next to the aggregate, plus a fairness ratio with a floor that forces same-day review, not a quarterly audit."
Why this works
Matches the direct answer. A named tile beats "we'd build in more fairness checks."
7
Prove it with the failure that actually happened
Say it like this
"Here's what happens without it. Cazarrow retrained FirstStep's model on eighteen months of new data to catch up with more gig-economy applicants. Aggregate accuracy went up. Nobody sliced the eval by documentation type, so nobody saw the self-employed pre-qualify rate quietly fall by more than half. A new hire doing an unrelated report noticed it by accident, nine weeks later."
Why this works
A four-sentence failure beats a paragraph about the importance of fairness.
8
Close on the option you ruled out and what it costs
Say it like this
"We looked at just running a fairness review every quarter, the way compliance already did for the full loan book, and ruled that out. It catches the harm eventually, and eventually means three more months of Consuela's segment quietly getting turned away. The real cost of the fix is that every new model version needs a few extra days to clear the segment check before it ships. That's the trade I'd take over a quiet count of people we'll never hear from again."
Why this works
Naming a rejected option and a real cost turns "watch for bias" into a defensible decision.
If you remember one thing A number that stays healthy on average while one group's real number collapses underneath it is not measuring health. It is measuring whoever is loud enough to move the average.

Let's learn

What happens when a health dashboard tells the truth about most people and stays quiet about the rest?

FirstStep is a chat window on Cazarrow Financial's loan page. Type in your income, your rent, and how you get paid, and it tells you in under a minute whether a full loan application is likely to go anywhere.

Hand sketched comparison titled One dashboard, two very different views. Left, Emmett at a gauge labeled dashboard, caption sees one clean number, can slice it, escalate it, fix it. Right, a person icon labeled Consuela on her phone, caption sees one soft line, no score shown, no appeal button.
Same product, same morning. One side gets a lever. The other gets one line of chat.

Before FirstStep, every applicant filled out a full application and waited two to four days for a loan officer to look at it, even the ones who never had a real shot. FirstStep cuts that to about forty seconds and tells three out of five people, before they've spent an evening gathering documents, whether it's worth the effort. Cazarrow's team watches one number every morning to know FirstStep is healthy: the aggregate share of chats that end in a positive pre-qualify estimate. It usually sits between fifty six and fifty nine percent, and it hadn't really moved in nine weeks.

Instant pre-qualify rate, week nine, by who is asking
100% 0% 58% 63% 19% Aggregate Salaried Self-employed
The one tile Emmett's team watches every morning sat at fifty eight percent. It never showed that self-employed applicants were pre-qualifying at less than a third of the salaried rate.

Here is the turn. Say it plainly: the aggregate number was not broken. It was doing exactly what an average does, describing the middle of a crowd. Self-employed and gig-income applicants, people paid from a business account instead of a steady paycheck, had gone from a forty one percent pre-qualify rate to nineteen percent over those same nine weeks. The aggregate barely moved, because that group was only about one in six of everyone who opened the chat.

The average wasn't lying. It just stopped being about anyone in particular.
Knowledge spark: what is an adverse action notice? A letter a lender has to send when it turns down a finished loan application, saying why and giving the applicant a chance to respond. FirstStep's soft "approval isn't likely" line happens before a real application exists, so it never triggers one. No finished application, no notice, no formal record anywhere.

At its worst, this doesn't look like a scandal. It looks like nothing. Consuela Marrufo runs a tailoring shop and opens FirstStep on a Tuesday afternoon in April. She gets the same line five other self-employed applicants got that week: "Based on what you've told us, approval isn't likely today." She could still finish the full application. She doesn't. She closes the tab and starts looking for a different lender. Nobody at Cazarrow logs that as a denial, because she never technically applied. Nobody at Cazarrow logs it at all.

Hand sketched decision tree titled What happens after the soft no. Root box reads FirstStep, approval is not likely today. One branch, finishes the application anyway, leads to a person reviews it, a denial comes with a reason and a right to reply. Other branch, closes the tab like most self employed applicants do, leads to no application, no notice, no appeal, nobody at Cazarrow sees it.
Only one of these two paths has a person, a reason, or a way to push back. FirstStep can't tell which path most self-employed applicants take, because it never asks.
The decision that mattered Building FirstStep's health dashboard around one aggregate tile, with no segment slice and no fairness check. It was the right call at launch, when almost every applicant was salaried. Nobody ever came back to ask if that was still true.

What I would leave alone: FirstStep's dashboard doesn't need a daily tile for every trait anyone could imagine. Age, for one, has shown no real gap in this data. Adding a tile for every group anyone could think of just buries the one gap that's real under twenty that aren't.

The lesson: an average is a promise that everyone inside it is doing about the same. The day one group stops being close to that average is the day the average starts lying, on purpose or not, and nobody built anything to notice.

Now here is the same thing as a story

The short version sits above. Read on for the Tuesday standup where a new hire's aside was the only reason anyone found out.

Emmett Vasilenko has run analytics for Cazarrow Financial's lending products for three years. Before FirstStep had a dashboard at all, he was the one who built it, back when Cazarrow's whole loan book was salaried employees with a pay stub and a landlord letter.

For most of two years, the one tile on his morning screen barely moved. Fifty five percent one week, fifty nine the next, back to fifty seven. A boring number was the whole point. It meant FirstStep was doing today what it did yesterday.

The fading happened in three beats, and none of them looked like a mistake at the time. First, Emmett stopped opening the old weekly breakdown report every Monday, the one that split the pre-qualify rate out by a dozen applicant traits, because it kept agreeing with the aggregate anyway. Second, when Cazarrow's product team made growing self-employed and gig-income applicants a real priority, nobody rebuilt that weekly report after a reorg quietly dropped it. Third, by the time FirstStep's model got retrained on newer data, the aggregate tile on Emmett's dashboard was the only number left in the room.

The trigger was a Tuesday standup. Nia Faridah, three weeks into the analytics team, was building an unrelated churn report and mentioned it almost as a throwaway line: "is it normal that almost none of the self-employed folks in my sample got a positive pre-qualify? Looked like one in five to me."

Emmett pulled the real number that afternoon. Self-employed applicants had gone from a forty one percent pre-qualify rate nine weeks earlier to nineteen percent now. The aggregate had barely moved, fifty six to fifty nine percent the whole time, because that segment was only about one in six of everyone who opened the chat.

Here is the part that actually cost something. Nobody at Cazarrow could say how many people like Consuela Marrufo had seen that soft line and quietly closed the tab over those nine weeks, because none of them had finished a real application. No application means no adverse-action notice, no fair-lending flag, no record in any system built to catch this. When Emmett's VP asked how long this had been happening, the honest answer was nine weeks, and nobody had known for eight of them.

It was never the aggregate lying. It just stopped being about anyone in particular.

Three years earlier, setting up that very first dashboard, the decision took about twenty minutes in a room with two other people. One aggregate tile, the daily pre-qualify rate. It was the right call. The applicant pool really was one group back then, and slicing it six ways would have shown six nearly identical charts saying the same thing. Nobody ever put a date on that decision. Nobody ever came back to ask whether the applicant pool was still one group.

Run the same Tuesday through the fixed design. The fairness ratio, self-employed rate over salaried rate, crosses below its floor of zero point eight in week three, not week nine, and it crosses because of a same-week trend, not one noisy day. The dashboard flags itself. A same-day review finds the real cause, a documentation-weighting change in the retrained model, and engineering ships a fix that Friday. Six weeks of exposure become gone. Consuela Marrufo, and everyone who opens FirstStep's chat after her that week, never sees the message that sent her looking for a different lender.

One design watched a number that used to describe everyone. The other watches whether it still does.

What I would tell myself, back in that twenty-minute meeting: the day you build a dashboard around the word everyone, write down what everyone means today. Someday it will not mean that anymore, and nothing you built will tell you when.

GUARD, and the five checks the old dashboard skipped

This is a design question, but the honest brief for a daily dashboard is really a fairness check that runs every morning, so GUARD does the work here, not a design-only framework.

G
Groups. Who is affected.
Emmett's team, who read a dashboard every morning. Applicants like Consuela Marrufo, who read one line of chat and never learn a model decided anything about them.
In this answer: not applicants in general. Consuela, specifically, and the roughly one in six FirstStep users who share her situation.
U
Unequal. Where the harm lands.
Self-employed and gig-income applicants pre-qualify at nineteen percent. Salaried applicants pre-qualify at sixty three percent. The aggregate, fifty eight percent, hides that gap because the harmed group is a minority of total volume.
This is the whole reason one aggregate tile fails here. A number built from everyone can look fine while it is actively wrong for someone.
A
Ability to contest. Who never gets to push back.
FirstStep's message is a soft nudge, not a denial, so it never triggers a formal adverse-action notice. An applicant who closes the tab leaves no complaint, no ticket, and no record anywhere in Cazarrow's compliance system.
Consuela never gets a lever. Nobody decided that on purpose. It fell out of treating a pre-qualify estimate as a suggestion instead of a real decision.
R
Reduce. The specific design change.
A segment-sliced pre-qualify rate on the daily dashboard, next to the aggregate, plus a fairness ratio with a floor that forces same-day review.
Not a policy document. A tile Emmett's team looks at every single morning, the same way they already look at the aggregate.
D
Detect. How you'd know before someone external tells you.
The ratio would have crossed its floor by week three instead of surfacing by accident in week nine, from a new hire's aside in a standup.
Three weeks of exposure instead of nine, and a fix that happens before a regulator, a journalist, or a lawyer ever has to point it out.

Three things worth naming directly here, since this is where the real judgment lives. The easy alternative on offer was to keep the existing quarterly fair-lending review and just ask compliance to look harder next time. That got ruled out on purpose: eventually is not a plan, it is three more months of Consuela's segment quietly getting turned away before anyone with the power to fix it even looks. The failure worth naming by name is distribution shift: FirstStep's retrained model was never taught a rule that says reject self-employed applicants, it simply saw eighteen months of new applications with a different mix of income documentation than it had seen before, and nobody evaluated it on that split before it shipped. The guardrail is a segment-sliced golden set, a batch of past applications with known good outcomes across documentation types, gating every new model version before release, not after. And the floor itself is not a hard rule that fires on any single bad reading. One odd day can be one small segment having a rough morning. What actually forces a same-day review is a same-week trend below the floor, checked against that segment's own rolling average, because a number that fires on noise gets ignored within a week. That gate costs something real too: every new instant-decision model version now needs a few extra days to clear the segment check before it reaches anyone, and the segment tables themselves are a daily compute cost Cazarrow didn't used to pay. That's the trade. Slower releases and a bigger reporting bill, in exchange for catching a real harm in days instead of months.

And if you want to be sure it really works, try it somewhere else

Same five letters, a housing voucher waitlist instead of a loan, and the gap runs along documentation type again, proof the method isn't a fluke of this one industry.

Grantham County Housing Authority runs TriagePath, a chatbot that pre-screens applicants for a subsidized housing voucher waitlist by asking about income, household size, and paperwork, before a caseworker is ever assigned. Arlen Njoku runs intake operations there.

G, groups. Grantham's intake team, who watch a daily dashboard. Applicants on the voucher waitlist, who talk to TriagePath before any human is in the loop.
U, unequal. Applicants with steady paycheck income pass TriagePath's documentation check at seventy nine percent. Applicants paid through gig platforms or freelance invoices pass at thirty three percent, because TriagePath's automated reader wants a regular monthly deposit pattern to match against, and irregular income never produces one. The aggregate pass rate, seventy one percent, looks fine because steady-income applicants are most of the waitlist.
A, ability to contest. An applicant flagged as not meeting the documentation requirement never gets assigned a caseworker, so there is nobody to appeal to yet. Their spot on the waitlist simply stops moving, and nothing tells them why.
R, reduce. The same fix. A documentation-type slice on the daily intake dashboard, next to the aggregate pass rate, with a fairness-ratio floor that pulls a caseworker in before an applicant's file goes cold.
D, detect. A daily automated check instead of Grantham's existing annual program audit, the only thing that used to look at this at all.

Voucher pre-screen pass rate, by income documentation type
100% 0% 71% 79% 33% Aggregate Steady income Gig income
Grantham's old dashboard only showed the aggregate. Seventy one percent looked fine right up until someone finally split it by income type.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever aggregate number your dashboard already shows, put the same number split by the group most likely to be hurt quietly, right next to it, every day.
Cost: engineering says a real segment pipeline is two months out. Don't read the aggregate alone as a stopgap in the meantime, pull the segment split by hand from raw logs once a week until the real pipeline ships.
The model got better, for real: say FirstStep's overall accuracy genuinely improves next quarter. That still isn't the same claim as every segment doing fine. A better model can even make a gap easier to miss, because the aggregate number gets healthier at the same time one segment gets worse.

Where people run it wrong.
They build the segment slice, then leave it as a chart nobody's on the hook for reading, instead of a number with a floor and an owner.
They wait for the quarterly compliance review to catch a gap that's been running for weeks, instead of a daily number that would have caught it in days.
They fix the dashboard's mistake with a training deck on fairness, instead of a second number the dashboard checks automatically.

How to use it live. Say the reframe before naming a single tile: "an aggregate number can look perfectly healthy while it's actively failing one group, because the group it's failing is too small to move it." That buys you room to give the real design, instead of reciting "add a bias check" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a risk-and-fairness design question like this one?
Tap to flip
ANSWER
GUARD: name the groups, find where harm lands unevenly, ask who can't push back, design the reduce, build the daily detect.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Emmett Vasilenko, who leads analytics for Cazarrow Financial's FirstStep chatbot, and Consuela Marrufo, a self-employed applicant FirstStep quietly steered away.
3 · WHAT THE AGGREGATE HID
What did the single aggregate tile hide?
Tap to flip
ANSWER
The self-employed pre-qualify rate falling from 41 percent to 19 percent over nine weeks, while the aggregate barely moved because that segment was only about one in six applicants.
4 · THE TWO GROUPS
Name the two groups this answer is built around.
Tap to flip
ANSWER
Emmett's team, who can see the dashboard and change the model. Applicants like Consuela, who get one soft line of chat with no score shown and no appeal button.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building the health dashboard around one aggregate tile, with no segment slice and no fairness check. It was right when almost every applicant was salaried; nobody revisited it once the mix changed.
6 · THE NUMBER
Fill in the blank: over nine weeks, the self-employed pre-qualify rate fell from 41 percent to ___ percent, while the aggregate stayed near ___ percent.
Tap to flip
ANSWER
19 percent; 58 percent. The gap was real and the average never showed it.
7 · THE REPLAY
Same story, new dashboard, what changes?
Tap to flip
ANSWER
The fairness ratio crosses its 0.8 floor by week three instead of week nine. A same-day review catches the retrained model's documentation bug and fixes it that Friday.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
TriagePath, a housing voucher pre-screen chatbot at Grantham County Housing Authority. Same GUARD steps, a documentation-type gap instead of a self-employment gap.

Check yourself Score: 0 / 0

True or false
1. True or false: FirstStep's aggregate pre-qualify rate moved noticeably during the nine weeks the self-employed segment's rate was collapsing.
  • True
  • False
Show hint
Check what share of total applicants the self-employed segment actually was.
Show answer
False. The aggregate stayed in a fifty six to fifty nine percent band the whole time, because self-employed applicants were only about one in six of total volume.
Multiple choice
2. What is the real problem with a health dashboard that only shows one aggregate pre-qualify rate for a lending chatbot?
  • A. It costs too much compute to run every day.
  • B. It can hide a real gap for one segment behind a healthy-looking average.
  • C. It only works for chatbots with low traffic.
  • D. It ignores how fast the chatbot responds.
Show hint
Think about what an average is built to describe, and what it isn't.
Show answer
B. An average describes the middle of a crowd. A small, badly-treated group can sit entirely outside that middle without moving it.
Fill in the blank
3. Over nine weeks, the self-employed applicants' pre-qualify rate fell from 41 percent to ___ percent.
Show hint
Check the first chart in "Let's learn."
Show answer
19 percent. That's less than a third of the salaried rate of sixty three percent, and the aggregate never showed it.
Multiple choice
4. Why couldn't Emmett's team just keep an informal eye on fairness instead of building a mandatory segment tile and a ratio floor?
  • A. FirstStep's soft message never becomes a formal denial, so it never trips the usual complaint or audit alarms on its own.
  • B. Informal review is against company policy at every lender.
  • C. The model can only be checked once a quarter for technical reasons.
  • D. Self-employed applicants file more support tickets than other applicants.
Show hint
Think about what actually happens when an applicant sees a soft "unlikely to qualify" message.
Show answer
A. No finished application means no adverse-action notice and no formal record, so nothing in the usual compliance system ever flags it on its own.
Short answer, apply it yourself
5. Pick a product you use that gives everyone a slightly different experience based on some detail about them. What's one group that product's own dashboard might be quietly failing, with no way for that group to say so?
Show hint
Look for a place where the product makes a judgment call about a person before a human ever gets involved.
Show answer
Model answer: A ride-share app that estimates wait times. In a neighborhood with fewer drivers signed up, the app might just show a longer wait, quietly, with no explanation, while its city-wide average wait time looks fine. The people affected have no way to flag it, because a long wait doesn't look like a bug. It looks like the truth.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the "decision that mattered" box in "Let's learn."
Show answer
Model answer: Building FirstStep's health dashboard around one aggregate tile, with no segment slice. It made sense three years ago, when almost every applicant was salaried and the aggregate really did describe everyone using the product. Nobody revisited that call once the applicant mix changed.
Before you close the answer
Why this works
Tests whether you can design detection into a tool your team already opens every morning, not just promise a review later. Most candidates answer a GUARD question by describing a policy. Few describe an actual tile.
Follow-up traps
"Isn't a 0.8 floor just a different threshold that could also be wrong?" Response: that's why it's a floor that forces review, not an automatic shutoff. A ratio under 0.8 doesn't turn FirstStep off, it puts a person in the loop the same day, the way a low reading sends a nurse to a bedside instead of writing a prescription by itself.

"What if the gap is real risk, not bias, self-employed income really is harder to verify?" Response: that's exactly why the fix isn't "approve more self-employed applicants," it's "stop letting a documentation problem look like nobody checked." The segment tile doesn't decide the answer, it decides that somebody has to look.
If pressed
The retrained model's drop didn't come from an explicit rule. It came from a shift in the training data's documentation mix: eighteen months of new applications had a higher share of bank-statement income proof than pay-stub income proof, and the model had never been evaluated on that split before it shipped. The guardrail is a segment-sliced golden set, past applications with known good outcomes for both documentation types, that any new model version has to clear before release, not just the aggregate accuracy number.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more