Artifact critiqueAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #15
Build the roadmap review ritual you would run monthly for an AI team.
ORDERthe meeting that decided the roadmap by rank instead of by whoever spoke last
What would a monthly roadmap review for an AI team actually need to contain, on purpose, instead of whatever felt urgent that week? Wayfinder Housing Trust runs Haven Match, a tool that matches a person experiencing homelessness to an open shelter bed based on stated needs. Minh Tran, the nonprofit's Head of Product, built the ritual after two years of roadmap decisions made in hallway conversations.
The direct answer
Run one fixed hour a month with five parts in this order: read the single outcome number out loud, list what actually changed since last month, backtest any proposed change against real history before debating it, force a ranked call of advance, pause, or kill on every item, and name what stays untouched. The ranking is the whole point. A meeting that reviews ideas without ranking them by what breaks first if skipped isn't a ritual, it's a status update.
Do this, in order
Force a ranked call, advance, pause, or kill, on every item, every month.Why: without a forced rank, the loudest idea in the room wins by default, which is exactly how the roadmap got here.
Agree on the one outcome number before ranking anything.Why: a ranking built without a shared outcome is just five people's opinions taking turns.
Backtest every proposed change against real history before debating it.Why: an idea that sounds obviously good in a meeting can lose to a five-minute check against last quarter's actual data.
Rank by what's hardest to undo, not by what's loudest.Why: a reversible experiment and an irreversible trust loss don't belong on the same list without this distinction.
Name what stays untouched, out loud, every time.Why: a ritual that only ever announces changes teaches the team that stability is never a real decision.
How to answer this, stage by stage
Nobody is scoring whether you can list agenda items. They're scoring whether the agenda forces an actual rank, not a discussion.
Stage 1
Scope it to one real team
Say it like this
"I'll build this for Haven Match at Wayfinder Housing Trust, a shelter-matching tool where the roadmap used to get rewritten by whichever caseworker or funder brought up something loudest that week."
Why this works
Keeps the ritual from turning into a generic meeting-agenda template.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as ORDER. Outcome, what we're actually trying to move. Reversibility, what's hardest to undo. Dependency, what unblocks what. Evidence, what we can learn cheaply first. Rank, the actual order, defended."
Why this works
Tells the interviewer this is a repeatable method, not a one-off list of good ideas.
Stage 3
Reframe: a ritual without a forced rank is just a status update
Say it like this
"The question isn't 'what should we talk about once a month.' It's 'what forces this room to actually rank things instead of taking turns being persuasive.' Without the rank, the meeting has no teeth."
Why this works
This is where a strong answer separates from "we'll have a monthly sync."
Stage 4
Give the one decision
Say it like this
"Five parts, one hour, fixed order: the outcome number, what changed, the backtest, the forced rank, and what stays untouched. Nothing gets ranked without a backtest first, and nothing gets discussed that isn't already backtested."
Why this works
This is the direct answer, stated as a buildable agenda instead of a value about discipline.
Stage 5
Prove it with the compressed failure
Say it like this
"Before the ritual, Wayfinder's roadmap changed direction about six times a quarter, mostly from hallway conversations. Nobody could say which change actually helped, because nothing was ever tested against real placement data first."
Why this works
Compresses the whole failure into the missing step the ritual was built to add.
Stage 6
Name the AI-specific trap, then close
Say it like this
"The trap is treating a backtest number as certain. A model's match quality on last quarter's data can still look different on next month's clients if the mix of needs shifts, so the ritual has to recheck the outcome number the month after any change ships, not assume the backtest was the last word."
Why this works
Shows the guardrail is about model evaluation drift, not a generic caution about change.
Let's learn
What does a monthly roadmap review for an AI team actually need in it, so it decides things instead of just discussing them?
Haven Match reads a client's stated needs, family size, medical requirements, whether they have a pet, and ranks open shelter beds by best fit. Before the ritual, Wayfinder had no fixed process for deciding what to build next. Feature ideas came from caseworker complaints, funder suggestions, and whatever a board member had read about a new model that week. A change could ship because three people mentioned it in one week, or die because nobody happened to bring it up.
Five fixed steps, same order every month. The rank in step four is the part that actually does the deciding.
Here's the turn: the problem wasn't too many good ideas. It was that no idea ever competed against another idea on the same terms. A caseworker's complaint about pet-friendly filtering and a board member's excitement about a new model both got the same hallway-length hearing, and whichever one got said most recently usually won.
Ad hoc, un-ranked roadmap changes per month, trailing six months
Fewer changes isn't the goal by itself. It's what happens once every change has to earn a spot on a ranked list instead of just being the newest thing said out loud.
At its worst, this doesn't just waste meeting time. Wayfinder shipped a "smarter" ranking change in one quarter that nobody had backtested, and successful placements actually dropped for six weeks before anyone noticed, because the only measure anyone was tracking was whether the change had shipped, not whether it worked.
Successful placements, clients still housed past 72 hours, before and after the ritual
The line chart above shows how many changes shipped, un-ranked. This one shows what those changes actually did to the outcome the whole roadmap was supposed to move.
The choice I would take back
Early on, Wayfinder decided that any team member could raise a roadmap idea directly to Minh at any time, since the org was small and formal process felt like overhead. That made sense with three people and one product. It stopped making sense once ideas outnumbered the hours to properly test any of them.
What I would leave alone: caseworkers can still flag urgent bugs any time, outside the monthly ritual. A ritual for prioritizing new work is not the same thing as an incident process, and forcing bug reports to wait a month would be its own kind of harm.
The lesson: a roadmap doesn't get chaotic because people care too much. It gets chaotic because nothing ever has to compete against anything else on the same terms. Build the ranking step, and caring stops being the same thing as winning.
Now here is the same thing as a story
The short version above is what you'd put in a one-page process doc. Read this one for how a shared tablet full of sticky-note ideas became the thing that finally forced a rank.
A shared tablet sat on the front desk at Wayfinder's main office, the one three staff members used to log client intakes. Somewhere along the way, it also became where people typed roadmap ideas whenever they thought of one: "add a pet-friendly filter," "what if we used the new model everyone's talking about," "a caseworker asked why it never explains its ranking."
Every idea on the tablet looked the same length. They were never the same size at all.
Minh Tran had run product at Wayfinder for three years, since before Haven Match existed. He was good at turning a caseworker's half-formed complaint into a clear feature. For a long while, that skill was enough. The tablet stayed short. Ideas got built, roughly in the order they arrived.
Then the tablet started filling up faster than anyone could act on it. Not because of one bad month. It built up slowly, over about a year, with no single moment anyone could point to. A funder mentioned a new model in a check-in call. A board member forwarded an article. A caseworker asked, again, why the tool didn't explain its ranking. Each idea got a hallway conversation, and whichever one had been said most recently felt the most urgent.
Knowledge spark: why can't you just build the best idea first?
"Best" isn't obvious until you check it against real outcomes. A ranking change that sounds smart in a meeting might do nothing for actual placements, or even hurt them, and the only way to know before shipping it is to run it against real history first, not to judge how convincing it sounded out loud.
One quarter, a new ranking approach got built because it sounded genuinely smart in a meeting, and everyone wanted to try the new model behind it. Nobody backtested it against real placement history first, because there was no fixed step in the process that required it. Successful placements, clients still housed past 72 hours, dropped from 61 percent to 54 percent over six weeks before a caseworker flagged that something felt off.
The tablet never sorted ideas this way. Everything sat in one flat list, waiting its turn to be argued about.
The tablet wasn't full of bad ideas. It was full of ideas that had never been asked to compete against each other on the same terms.
So here is the decision I would take back: letting any idea reach Minh directly, at any time, without ever passing through a shared outcome number or a backtest. That was fine with three people and a handful of ideas a month. It stopped being fine once the tablet held more ideas than any one person could hold in their head at once.
The step that was missing sits second from the left, exactly where a hallway conversation used to stand in for it.
With the ritual in place, the replay runs differently. The same funder mention, the same board article, the same caseworker question all land on the tablet the same way. But now they wait for the first Tuesday of the month, where Minh reads Haven Match's one outcome number out loud, backtests each idea against real history, and forces a rank: advance, pause, or kill. The ranking model idea that once shipped untested now gets backtested first, and it either earns its place above the pet-friendly filter or it doesn't. And the thing I'd tell myself, if I could go back: I never needed to say no to more ideas. I needed a room where ideas had to prove themselves against each other before any one of them got to be the loudest.
ORDER, in one screen
O
Outcome. What is every candidate competing to move?
Successful placements: clients still housed with the matched shelter past 72 hours. Not "features shipped," not "model accuracy" alone.
Without a shared outcome, ranking is just five people's opinions taking turns.
R
Reversibility. Which decision is hardest to undo?
A ranking-weight tweak reverts in a day. Removing the caseworker override doesn't come back once staff stop trusting the tool enough to lean on it.
Two ideas that look equally exciting in a meeting can carry completely different costs if they're wrong.
D
Dependency. What unblocks what?
You can't fairly judge a new ranking model until the backtest pipeline against real outcomes exists, so the pipeline has to come before any model-change debate, not after.
Some order is forced by reality, and pretending otherwise wastes a meeting arguing about the wrong thing first.
E
Evidence. What could you learn cheaply before committing the month?
Run any proposed ranking change against last quarter's real placement history before it ever reaches a discussion, let alone a rank.
This is the hardest step, and the one the whole ritual turns on: no backtest, no debate.
R
Rank. State the order, defend the top pick.
The backtest pipeline ranks first, since nothing else can be fairly judged without it. The pet-friendly filter, small and reversible, ranks lowest of the live candidates.
A rank you can defend in one sentence is a real rank. One that needs a paragraph of caveats isn't.
The recap, one line per letter: outcome is successful placements past 72 hours, reversibility is telling a ranking tweak apart from removing the override, dependency is building the backtest pipeline before ranking anything against it, evidence is backtesting every idea before it's even discussed, and rank is stating the order and defending the top pick in one breath.
And if you want to be sure it really works, try it somewhere else
Thermwell Field Services runs DispatchLine, a tool that assigns HVAC technicians to service calls based on skill match, drive time, and part availability. Colette Marchetti, the dispatch lead, used to field roadmap ideas the same way Wayfinder once did, whoever radioed in an idea during the week got it raised at the next all-hands, in whatever order people remembered to bring it up. Mapped onto ORDER: outcome is on-time arrival rate, not dispatch-model accuracy alone. Reversibility separates a routing-weight adjustment, reversible in a day, from removing the dispatcher's manual override, which technicians would stop trusting the system enough to use again if pulled. Dependency means building a backtest against last month's actual routes before debating any new dispatch model. Evidence is running that backtest before the meeting, not during it. Rank forces a call on every item: advance, pause, or kill, defended in one sentence. The story underneath is different too, a delegation flip rather than the un-ranked hallway chaos at Wayfinder: once dispatch decisions started routing through a monthly ritual, the regional manager who used to personally override any technician's route he disagreed with had to let the ranked call stand instead of reclaiming it himself.
The same four weeks repeat every month, whether the roadmap is shelter beds or service trucks.
Same five parts as Haven Match's ritual, just measuring trucks instead of shelter beds.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "outcome number, what changed, backtest first, forced rank, name what stays," and stop.
Cost: no budget for a formal backtest pipeline yet. Say so honestly, and start with a manual spreadsheet check against last quarter's data, since a rough backtest beats no backtest at all.
The model gets better, for real: if a proposed change backtests as a clear win, that's the ritual working as intended, and the honest move is to rank it first and ship it fast, not slow it down out of habit.
Where people run it wrong.
They hold the meeting but skip the backtest, so the same hallway-persuasion dynamic just moves onto a calendar invite.
They rank by how exciting an idea sounds instead of by what breaks first if it's skipped or how hard it is to undo.
They never say out loud what's staying the same, so the team only ever hears about the roadmap when something's changing.
How to use it live. The moment someone asks you to build a review ritual, ask yourself: what in this agenda actually forces a rank, instead of just giving everyone a turn to talk? Build that part first, and the rest of the agenda follows.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one-line job?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Its job is ranking candidates by what breaks first if skipped, not by who argued loudest.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Minh Tran, Head of Product at Wayfinder Housing Trust, who had run Haven Match's roadmap for three years off hallway conversations before building the ritual.
3 · THE HABIT
What did the team stop doing as ideas piled up?
Tap to flip
ANSWER
They stopped checking any idea against real placement history before building it, and started judging ideas by how recently and persuasively they'd been raised.
4 · THE MISSING STEP
What single step, if added, would have caught the bad ranking change before it shipped?
Tap to flip
ANSWER
A backtest against last quarter's real placement history, required before any idea reaches a debate, let alone a ranked decision.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting any roadmap idea reach Minh directly at any time, with no shared outcome number or backtest step in between. Fine at three people, not fine once ideas outnumbered the hours to test them.
6 · THE NUMBER
Fill in the blank: successful placements dropped from 61 percent to ___ percent over six weeks after the untested ranking change shipped.
Tap to flip
ANSWER
54 percent, and nobody caught it sooner because nothing tracked placements as an ongoing number, only whether the change had shipped.
7 · THE REPLAY
Same tablet full of ideas, ritual now in place, what changes?
Tap to flip
ANSWER
Every idea waits for the first Tuesday, gets backtested against real history, and only then gets ranked: advance, pause, or kill. Un-ranked changes per month dropped from 7 to 1 within three months.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what changed in the story?
Tap to flip
ANSWER
Thermwell Field Services' DispatchLine. The story shifts to a delegation angle: a regional manager who used to personally override routes had to let the ranked call stand instead.
Check yourself Score: 0 / 0
True or false
1. True or false: this answer's ritual lets caseworkers raise urgent bugs only once a month, during the review.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. Urgent bugs still get flagged any time; the monthly ritual is for prioritizing new work, not for incidents.
Multiple choice
2. Why does the ritual require a backtest before any idea gets discussed, not just before it ships?
A. Because backtests take too long to run during the meeting itself.
B. Because an idea that sounds convincing in conversation can still be a loss against real outcomes, and the meeting shouldn't spend time debating ideas that fail that check.
C. Because Wayfinder's funders require a backtest report for every roadmap item.
D. Because the model can only be evaluated once a month.
Show hint
Look at the knowledge spark and the "evidence" step.
Show answer
B. The backtest exists to stop a persuasive-sounding idea from winning a debate it would lose against real data.
Fill in the blank
3. Fill in the blank: ad hoc, un-ranked roadmap changes dropped from 7 a month to ___ a month within three months of the ritual starting.
Show hint
Look at the line chart in Section 1.
Show answer
1 a month. Fewer changes wasn't the goal itself, it was the side effect of everything having to earn a ranked spot.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at "the choice I would take back."
Show answer
Model answer: Letting any idea reach Minh directly at any time. That made sense with three people and few ideas, and stopped making sense once ideas outnumbered the hours to test them.
Short answer, apply it yourself
5. Think of a team or group you're part of that decides what to do next informally. What's one thing that would change if every idea had to be ranked against every other idea by the same rule, instead of by who mentioned it most recently?
Show hint
Think of a family, a club, or a small team where whoever spoke last tends to set the agenda.
Show answer
Model answer: A family deciding weekend plans by whoever asked most recently, instead of by what actually matters most to the most people that week.
Short answer, where it wouldn't matter
6. Name a case at Wayfinder where the monthly ritual genuinely shouldn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An urgent bug a caseworker needs fixed right away. That's an incident, not a roadmap idea, and it shouldn't wait for the first Tuesday of the month.
Before you close the answer
Why this works
Tests whether you'll design a ritual that actually forces a decision, or one that just gives everyone a scheduled turn to be persuasive.
Follow-up traps
"What if the backtest data doesn't exist yet for a brand-new idea?" Response: then building the smallest version of that data, a quick historical proxy or a short pilot, becomes the dependency that has to rank first, ahead of the idea itself.
"Doesn't a rigid five-step agenda kill genuine urgency?" Response: no, because true incidents route around the ritual entirely; the ritual only governs new roadmap work competing for the same limited hours.
If pressed
The backtest step specifically replays each candidate change against the last full quarter of real intake records, not a synthetic test set, because Haven Match's actual need-mix, family size, medical flags, pets, shifts enough month to month that a synthetic set quietly stops matching reality.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.