Design the daily dashboard for an AI feature team.
Design the daily dashboard to show who the aggregate number is hiding, not just how the aggregate number looks.
- Put a segment-sliced pre-qualify rate on the dashboard, next to the aggregate, every single day.Why: an average can sit healthy for months while one real group's number collapses underneath it.
- Add a fairness ratio, the worst segment's rate over the best, as its own tile with a floor that forces same-day review.Why: a ratio nobody is watching is a number that only exists for a report nobody reads until it is too late.
- Track how many flagged applicants quietly leave the chat without finishing a full application.Why: FirstStep's soft message never becomes a formal denial, so a quiet exit is the only trace most of this harm leaves.
- Gate every new instant-decision model version on the same segment slice before it ships, not after.Why: the gap that hit Consuela's segment came from a retrained model that looked better on the aggregate and worse on one group nobody sliced for.
- Limit the daily segment tiles to groups where a real gap has actually been found.Why: fifty tiles nobody reads bury the one gap that is real, which is worse than showing none at all.
- Route a same-day alert to a person who can pause the model, not just to a chart in the corner.Why: a number that only moves a chart is decoration. A number that reaches someone who can act is a guardrail.
How to answer this, stage by stage
Nobody is grading whether you can name a fairness risk. They are grading whether you can turn it into a tile someone actually looks at tomorrow morning. Eight moves get you there.
Let's learn
What happens when a health dashboard tells the truth about most people and stays quiet about the rest?
FirstStep is a chat window on Cazarrow Financial's loan page. Type in your income, your rent, and how you get paid, and it tells you in under a minute whether a full loan application is likely to go anywhere.
Before FirstStep, every applicant filled out a full application and waited two to four days for a loan officer to look at it, even the ones who never had a real shot. FirstStep cuts that to about forty seconds and tells three out of five people, before they've spent an evening gathering documents, whether it's worth the effort. Cazarrow's team watches one number every morning to know FirstStep is healthy: the aggregate share of chats that end in a positive pre-qualify estimate. It usually sits between fifty six and fifty nine percent, and it hadn't really moved in nine weeks.
Here is the turn. Say it plainly: the aggregate number was not broken. It was doing exactly what an average does, describing the middle of a crowd. Self-employed and gig-income applicants, people paid from a business account instead of a steady paycheck, had gone from a forty one percent pre-qualify rate to nineteen percent over those same nine weeks. The aggregate barely moved, because that group was only about one in six of everyone who opened the chat.
At its worst, this doesn't look like a scandal. It looks like nothing. Consuela Marrufo runs a tailoring shop and opens FirstStep on a Tuesday afternoon in April. She gets the same line five other self-employed applicants got that week: "Based on what you've told us, approval isn't likely today." She could still finish the full application. She doesn't. She closes the tab and starts looking for a different lender. Nobody at Cazarrow logs that as a denial, because she never technically applied. Nobody at Cazarrow logs it at all.
What I would leave alone: FirstStep's dashboard doesn't need a daily tile for every trait anyone could imagine. Age, for one, has shown no real gap in this data. Adding a tile for every group anyone could think of just buries the one gap that's real under twenty that aren't.
The lesson: an average is a promise that everyone inside it is doing about the same. The day one group stops being close to that average is the day the average starts lying, on purpose or not, and nobody built anything to notice.
Now here is the same thing as a story
The short version sits above. Read on for the Tuesday standup where a new hire's aside was the only reason anyone found out.
Emmett Vasilenko has run analytics for Cazarrow Financial's lending products for three years. Before FirstStep had a dashboard at all, he was the one who built it, back when Cazarrow's whole loan book was salaried employees with a pay stub and a landlord letter.
For most of two years, the one tile on his morning screen barely moved. Fifty five percent one week, fifty nine the next, back to fifty seven. A boring number was the whole point. It meant FirstStep was doing today what it did yesterday.
The fading happened in three beats, and none of them looked like a mistake at the time. First, Emmett stopped opening the old weekly breakdown report every Monday, the one that split the pre-qualify rate out by a dozen applicant traits, because it kept agreeing with the aggregate anyway. Second, when Cazarrow's product team made growing self-employed and gig-income applicants a real priority, nobody rebuilt that weekly report after a reorg quietly dropped it. Third, by the time FirstStep's model got retrained on newer data, the aggregate tile on Emmett's dashboard was the only number left in the room.
The trigger was a Tuesday standup. Nia Faridah, three weeks into the analytics team, was building an unrelated churn report and mentioned it almost as a throwaway line: "is it normal that almost none of the self-employed folks in my sample got a positive pre-qualify? Looked like one in five to me."
Emmett pulled the real number that afternoon. Self-employed applicants had gone from a forty one percent pre-qualify rate nine weeks earlier to nineteen percent now. The aggregate had barely moved, fifty six to fifty nine percent the whole time, because that segment was only about one in six of everyone who opened the chat.
Here is the part that actually cost something. Nobody at Cazarrow could say how many people like Consuela Marrufo had seen that soft line and quietly closed the tab over those nine weeks, because none of them had finished a real application. No application means no adverse-action notice, no fair-lending flag, no record in any system built to catch this. When Emmett's VP asked how long this had been happening, the honest answer was nine weeks, and nobody had known for eight of them.
Three years earlier, setting up that very first dashboard, the decision took about twenty minutes in a room with two other people. One aggregate tile, the daily pre-qualify rate. It was the right call. The applicant pool really was one group back then, and slicing it six ways would have shown six nearly identical charts saying the same thing. Nobody ever put a date on that decision. Nobody ever came back to ask whether the applicant pool was still one group.
Run the same Tuesday through the fixed design. The fairness ratio, self-employed rate over salaried rate, crosses below its floor of zero point eight in week three, not week nine, and it crosses because of a same-week trend, not one noisy day. The dashboard flags itself. A same-day review finds the real cause, a documentation-weighting change in the retrained model, and engineering ships a fix that Friday. Six weeks of exposure become gone. Consuela Marrufo, and everyone who opens FirstStep's chat after her that week, never sees the message that sent her looking for a different lender.
One design watched a number that used to describe everyone. The other watches whether it still does.
What I would tell myself, back in that twenty-minute meeting: the day you build a dashboard around the word everyone, write down what everyone means today. Someday it will not mean that anymore, and nothing you built will tell you when.
GUARD, and the five checks the old dashboard skipped
This is a design question, but the honest brief for a daily dashboard is really a fairness check that runs every morning, so GUARD does the work here, not a design-only framework.
Three things worth naming directly here, since this is where the real judgment lives. The easy alternative on offer was to keep the existing quarterly fair-lending review and just ask compliance to look harder next time. That got ruled out on purpose: eventually is not a plan, it is three more months of Consuela's segment quietly getting turned away before anyone with the power to fix it even looks. The failure worth naming by name is distribution shift: FirstStep's retrained model was never taught a rule that says reject self-employed applicants, it simply saw eighteen months of new applications with a different mix of income documentation than it had seen before, and nobody evaluated it on that split before it shipped. The guardrail is a segment-sliced golden set, a batch of past applications with known good outcomes across documentation types, gating every new model version before release, not after. And the floor itself is not a hard rule that fires on any single bad reading. One odd day can be one small segment having a rough morning. What actually forces a same-day review is a same-week trend below the floor, checked against that segment's own rolling average, because a number that fires on noise gets ignored within a week. That gate costs something real too: every new instant-decision model version now needs a few extra days to clear the segment check before it reaches anyone, and the segment tables themselves are a daily compute cost Cazarrow didn't used to pay. That's the trade. Slower releases and a bigger reporting bill, in exchange for catching a real harm in days instead of months.
And if you want to be sure it really works, try it somewhere else
Same five letters, a housing voucher waitlist instead of a loan, and the gap runs along documentation type again, proof the method isn't a fluke of this one industry.
Grantham County Housing Authority runs TriagePath, a chatbot that pre-screens applicants for a subsidized housing voucher waitlist by asking about income, household size, and paperwork, before a caseworker is ever assigned. Arlen Njoku runs intake operations there.
G, groups. Grantham's intake team, who watch a daily dashboard. Applicants on the voucher waitlist, who talk to TriagePath before any human is in the loop.
U, unequal. Applicants with steady paycheck income pass TriagePath's documentation check at seventy nine percent. Applicants paid through gig platforms or freelance invoices pass at thirty three percent, because TriagePath's automated reader wants a regular monthly deposit pattern to match against, and irregular income never produces one. The aggregate pass rate, seventy one percent, looks fine because steady-income applicants are most of the waitlist.
A, ability to contest. An applicant flagged as not meeting the documentation requirement never gets assigned a caseworker, so there is nobody to appeal to yet. Their spot on the waitlist simply stops moving, and nothing tells them why.
R, reduce. The same fix. A documentation-type slice on the daily intake dashboard, next to the aggregate pass rate, with a fairness-ratio floor that pulls a caseworker in before an applicant's file goes cold.
D, detect. A daily automated check instead of Grantham's existing annual program audit, the only thing that used to look at this at all.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever aggregate number your dashboard already shows, put the same number split by the group most likely to be hurt quietly, right next to it, every day.
Cost: engineering says a real segment pipeline is two months out. Don't read the aggregate alone as a stopgap in the meantime, pull the segment split by hand from raw logs once a week until the real pipeline ships.
The model got better, for real: say FirstStep's overall accuracy genuinely improves next quarter. That still isn't the same claim as every segment doing fine. A better model can even make a gap easier to miss, because the aggregate number gets healthier at the same time one segment gets worse.
Where people run it wrong.
They build the segment slice, then leave it as a chart nobody's on the hook for reading, instead of a number with a floor and an owner.
They wait for the quarterly compliance review to catch a gap that's been running for weeks, instead of a daily number that would have caught it in days.
They fix the dashboard's mistake with a training deck on fairness, instead of a second number the dashboard checks automatically.
How to use it live. Say the reframe before naming a single tile: "an aggregate number can look perfectly healthy while it's actively failing one group, because the group it's failing is too small to move it." That buys you room to give the real design, instead of reciting "add a bias check" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the gap is real risk, not bias, self-employed income really is harder to verify?" Response: that's exactly why the fix isn't "approve more self-employed applicants," it's "stop letting a documentation problem look like nobody checked." The segment tile doesn't decide the answer, it decides that somebody has to look.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?