Artifact critiqueAdvancedAI Opportunity & Model Strategy / Opportunity identification for AI / #19

Build a scoring rubric for ranking ten candidate AI features.

ORDER · the six-line scorecard Dagmara Grimshaw built after a near miss on rate night

Haventhorne Hotels runs 34 boutique properties and ships one new AI feature every quarter, guest-facing or back-of-house. One spring, ten teams pitched ten different candidates at once: a concierge chat, a pricing engine, a voice assistant, a maintenance sensor, and six more, all competing for the same engineering slot. VP of Product Dagmara Grimshaw had to rank all ten before the room handed the slot to whoever pitched hardest, which is exactly what it was about to do.

The direct answer
Score every candidate on the same six weighted lines before ranking anything: guest and revenue lift, reversibility, whether real signal already exists to build and check it, how cheap it is to test the core assumption first, how fast a wrong output gets caught, and how far a wrong output can spread before someone catches it. Run the real numbers, and when two totals tie, break it on guest and revenue lift, the heaviest line, decided in advance, never on whoever argued longest in the room.
Do this, in order
  1. Score every candidate on the same six weighted lines before anyone ranks anything out loud.Why: a rubric only protects the ranking if it exists before the room starts arguing for a favorite.
  2. Weight dependency and detection cost like they actually matter.Why: for a model-based feature, they decide whether the score means anything at all.
  3. Run the actual scoring before committing the quarter's engineering slot, not after a pitch has already won the room.Why: a real number beats a confident deck every time it gets said out loud.
  4. Set the tie-break rule in advance, on the single heaviest-weighted line, never argued case by case.Why: this is what stops the rubric from quietly becoming a new way to justify the same favorite.
  5. Hold any candidate that scores low on dependency or detection cost in a silent test, not a launch.Why: a wrong output nobody can catch quickly is exactly the failure the rubric exists to price in.
  6. Revisit the weights between quarters, with real scores on file, not a fresh gut feeling each time.Why: the rubric is only repeatable if the criteria don't reshuffle to fit whichever feature is popular that quarter.

How to answer this, stage by stage

Nobody's grading whether Dagmara can make dynamic pricing sound exciting for five minutes. They're grading whether the rubric she names would actually have caught six rooms sitting under floor before they went live everywhere.

1
Scope it to one company, one list, one real fork
Say it like this
"Let's make this concrete. I'm VP of Product at Haventhorne Hotels, thirty-four boutique properties. Ten teams pitched ten AI features for one engineering slot, and I had to rank all ten before the room picked a favorite by feel."
Why this works
Naming the real company and the real list stops the answer from staying a vague debate about "good prioritization."
2
Name your method out loud before you use it
Say it like this
"I'd run this as ORDER. Outcome, what the rubric actually has to protect. Reversibility, how hard each pick is to walk back. Dependency, what has to already be true for it to work. Evidence, what's cheap to check first. Rank, the real scored table, defended."
Why this works
Two seconds of structure tells the interviewer they're about to get a method, not an opinion.
3
Reframe what the question is really testing
Say it like this
"This isn't really asking me to rank ten features. It's asking whether I'll build something that ranks the same way twice, no matter who's loudest in the room that quarter."
Why this works
Separates a real rubric from a fancy-looking excuse for whichever feature already had momentum.
4
Name the outcome the rubric has to protect
Say it like this
"Here's the outcome. The rubric isn't there to find the most exciting feature. It's there to protect Haventhorne from betting the whole quarter's trust on a pick nobody actually checked."
Why this works
Naming the outcome before scoring anything is what stops the ranking from becoming opinion wearing a spreadsheet.
5
Build the three checks a normal feature list skips
Say it like this
"Reversibility asks how public and hard to undo a miss is, once it's live. Dependency asks whether real, usable signal already exists to build and check this, or whether we'd be starting from nothing. Evidence asks what's cheap to test before we commit the whole quarter to it. Those three are the ones a plain feature list never asks."
Why this works
This is the part of the rubric that's actually specific to a model-based feature, not a generic priority list with AI features plugged in.
6
Show the actual scored table, not a description of one
Say it like this
"Six lines, weighted: guest and revenue lift twenty-five percent, dependency twenty percent, reversibility fifteen, evidence cost fifteen, detection cost fifteen, blast radius ten. Score each candidate one to five on every line. The personalized upsell offer comes out on top at four point three out of five. The in-room voice assistant comes out lowest, two point two five, because almost nothing about it is ready to be checked yet."
Why this works
This is the direct answer, said with real numbers, not a claim about how rigorous the process sounds.
7
Prove it with the near miss, compressed
Say it like this
"Here's what happens if you skip the dependency and reversibility checks. The dynamic pricing engine had the loudest champion and the best revenue story. Three hours before it was due to go live on every channel at all thirty-four properties, six room types were already priced under the floor written into our OTA contracts, and some of those contracts needed twenty-one days' notice before we could even correct the rate back up."
Why this works
Shows the real cost of skipping Reversibility and Dependency in four sentences, not a slide deck.
8
Close on the tie-break rule and what ships next
Say it like this
"When two totals tie, I break it on guest and revenue lift, the heaviest line, decided before anyone saw a single score. The upsell offer ships this quarter. Dynamic pricing goes back to the team with the floor built in as a hard constraint, not a suggestion, and it gets scored again next quarter with real evidence behind it."
Why this works
Ends on something concrete the interviewer can hold the candidate to later, not just a confident closing line.

Let's learn

Haventhorne Hotels runs thirty-four boutique properties. Every quarter, the product team ships one new AI feature, something a guest touches or something that runs behind the front desk.

Before Dagmara Grimshaw built a rubric, ten AI ideas got pitched in one two-hour meeting every spring. Whoever had the best deck and the loudest sponsor in the room walked out with the quarter's one engineering slot. Nobody wrote a real number down.

Hand sketched two panel comparison titled How Haventhorne used to rank a feature list. Left panel, a person icon labeled Old way, caption: whoever pitches loudest, biggest sponsor wins. Right panel, a document icon labeled New way, caption: same six weighted checks, every candidate.
Two ways to fill one engineering slot. Only one of them can be checked afterward.

Now, all ten candidates get the same six-line scorecard first, about ninety minutes of real scoring using numbers already sitting in Haventhorne's own booking and reservation systems, before anyone argues for a favorite out loud.

Hand sketched numbered icon list titled The ten candidates on Haventhorne's list. Ten rows, each a small icon and one line of text: Concierge chat for guest requests. Dynamic rate pricing engine. Personalized welcome and itinerary notes. Predictive housekeeping order. AI-drafted review responses. No-show and late-cancel scoring. Front-desk multilingual translation. In-room voice assistant. Predictive maintenance on room gear. Personalized upsell offers at check-in.
Guest-facing and back-of-house, hardware and pricing logic, all competing for one slot.
What does "dependency" actually check? Whether real, usable examples already exist to build and test a feature, or whether the team would be starting from nothing. A concierge chat has years of front-desk messages to learn from. A voice assistant has almost no real recordings from an actual hotel room yet.
Hand sketched left to right flow diagram titled The order the rubric actually gets built in. Five boxes connected by arrows, the third box highlighted: Name the outcome. Score reversibility. Score dependency. Run a cheap evidence check. Rank all ten.
The rubric gets built in this order, every quarter, before a single candidate gets scored.

Here is the turn. Ranking these ten wrong isn't really about picking a slightly worse feature first. It's about which mistake you can quietly fix later and which one you can't take back once it's live on every guest's screen, or every booking channel, at once.

We were not ranking ten features. We were ranking whichever exec argued the longest.

At its worst, this costs weeks, not hours. A pricing bug that goes live on every OTA channel at every property doesn't get quietly fixed the next morning. Some of Haventhorne's contracts need twenty-one days' notice before a live rate can be corrected back up, so an underpriced room stays bookable long after anyone's found the mistake.

The choice I would take back Haventhorne let the loudest pitch and the biggest sponsor decide the quarter's AI feature, with no shared, written criteria. Nobody could compare a chatbot idea against a pricing engine on the same terms. I would score every candidate on the same six lines before the room hears a single pitch.

What I would leave alone: a small wording change to a guest confirmation email, or a minor color tweak to the booking screen, doesn't need this rubric. Run the full six-line score only on candidates that are genuinely new AI capability, competing for the same engineering quarter.

The lesson: a rubric's real job isn't to make a ranking look scientific. It's to stop a room from mistaking confidence for evidence.

Every candidate, scored one to five and weighted, ranked top to bottom
0 5.0 Upsell offers 4.30 Welcome notes 4.05 Concierge chat 3.55 Housekeeping order 3.45 Review drafts 3.30 Front-desk translation 3.05 Predictive maintenance 3.05 Dynamic pricing 2.55 No-show scoring 2.55 Voice assistant 2.25
Top-ranked candidateAll other candidates
Two genuine ties on the list, front-desk translation with predictive maintenance, and dynamic pricing with no-show scoring, both break the same way: on guest and revenue lift, the heaviest line.

Now here is the same thing as a story

What you'd actually say sits above. Read this one for the ninety minutes it took to find out the loudest pitch in the room was also the most dangerous one.

Dagmara Grimshaw has run product at Haventhorne Hotels for five years. Ask her which guest complaint costs the most staff time and she'll tell you before you finish the question: a room that isn't ready at 3pm, every time.

Every spring, ten teams pitch their AI idea for the year's one open engineering slot. That spring, dynamic rate pricing had the loudest room. Bohumil Ravenshaw, the CFO, walked in with a slide that promised a real revenue lift across all thirty-four properties, and he'd been talking about it in the hallway for a month before the meeting even started.

Nobody else in the room had numbers that confident. The concierge chat idea had a modest, believable case. The voice assistant idea barely had a case yet at all. By the time the meeting ended, dynamic pricing had the room's vote, the way the loudest idea always had the room's vote.

Engineering built it over six weeks. It tested clean in every simulation the team ran on old data. So it went into a live night-audit test run across all thirty-four properties, scheduled to go fully live to every OTA channel at 6:14am.

Hand sketched horizontal timeline titled The night nobody wanted to be first to notice. Four milestones left to right, the third one highlighted: Test starts, 12:00am, all 34 properties. Rate dips under floor, 1:40am, six room types. Analyst flags it, 3:14am. Launch paused, 6:00am, hours before go-live.
Ninety-four minutes between the first wrong price and the trace that caught it.

At 12:00am, the test run started fine. By 1:40am, six room types, mostly the garden suites nobody watches closely, had quietly dropped under the floor rate written into Haventhorne's own OTA contracts, the minimum price the hotel had promised each channel it would never charge less than. The model had never been shown that floor. It only knew how to fill rooms, and empty garden suites were the cheapest rooms left to discount.

Nobody caught it at 1:40am. Dagmara caught it at 3:14am, running the same pre-launch trace she ran on every feature before it touched a real guest, out of habit more than suspicion.

Six rooms, under floor, on a system about to go live everywhere in three hours.

We did not almost mistarget one price. We almost mispriced every channel at every property, for weeks nobody could quietly undo.

She paused the launch at 3:20am. Nobody outside the product team ever saw those six wrong prices. But if the trace had run an hour later, or if she hadn't run it at all, those prices would have gone live on every OTA at 6:14am, and some of Haventhorne's contracts required twenty-one days' written notice before a live rate could be raised back up. A model that had never been shown its own contract's floor would have stayed live, discounting rooms nobody meant to discount, for three weeks minimum.

Bohumil wasn't careless. His revenue case was real; dynamic pricing genuinely could lift revenue once it worked. He'd simply never been asked the question the rubric now asks every candidate: does real, usable signal already exist to build and check this, and what's cheap to test before betting the whole quarter on it. Nobody had asked, because nobody had a rubric yet. The room just voted for the best pitch.

So here is the decision Dagmara took back. Haventhorne used to let the loudest pitch and the biggest sponsor decide the quarter's feature, with nothing written down to compare one candidate against another on the same terms. She built the six-line scorecard the next week: guest and revenue lift, reversibility, dependency, evidence cost, detection cost, blast radius, each weighted, each scored the same way for every candidate before anyone pitched a word.

Run against that scorecard, dynamic pricing scores low, not because it's a bad idea, but because it hadn't earned the right to run unattended yet. It goes back to engineering with the floor built in as a hard constraint, and it gets scored again once there's a real, checkable test behind it.

And the thing I'd tell myself, standing in that room the day the loudest pitch won: confidence is not evidence, and a rubric's whole job is to make you check the difference before the room forgets to.

ORDER, turned into the six columns Haventhorne scores every candidate on

This isn't a straight two-way tradeoff. Haventhorne had ten genuinely different candidates competing for one engineering quarter, guest-facing and back-of-house, hardware and pricing logic alike. That's what ORDER ranks. PICK only picks a side.

Hand sketched labeled parts diagram titled The six lines on Haventhorne's scorecard. A center document icon labeled The Rubric, with six callouts around it: Guest and revenue lift, Reversibility, Dependency, Evidence cost, Detection cost, Blast radius.
Six lines, one weight each, scored the same way for every candidate.
OOutcome. What the rubric actually has to protect.
Not the flashiest feature, and not whichever exec's revenue story sounds biggest. The rubric's real job is protecting Haventhorne's ability to compare ten genuinely different candidates on the same terms, and to catch a mistake before it's live everywhere at once. Everything scored below exists to serve that one outcome.
Name the outcome before scoring a single candidate. Skip this and "best pitch" quietly becomes the real ranking rule again.
RReversibility. How hard each pick is to walk back once it's live.
A personalized welcome note that underperforms is one guest's inbox, quietly reworked, nobody outside the team notices. Dynamic pricing gone live wrong is every OTA channel at every property, and some of those contracts needed twenty-one days' notice before the rate could even be corrected back up. Same size of mistake on paper. Wildly different cost once it's actually live.
This is why order matters more than preference. One miss is a quiet fix. The other is contractually stuck for weeks.
Hand sketched two panel comparison titled Which one can Dagmara still take back. Left panel, a box icon labeled Welcome notes, caption: one guest's inbox, quiet to turn off. Right panel, a box icon labeled Dynamic pricing, caption: live on every channel, hard to unsay.
Reversibility isn't a reason to avoid the bold idea. It's a reason to check the evidence before shipping it.
DDependency. What has to already be true for each candidate to succeed.
A candidate needs real, usable signal to be built and checked against, not just a good idea. The upsell offer had years of loyalty and booking history already logged, a strong four out of five. The in-room voice assistant had almost no real recordings from an actual guest room yet, a two out of five, because that signal doesn't exist until someone builds it first.
This is the line a normal feature list skips entirely. A new screen doesn't need a labeled dataset behind it before anyone can build it. A model's judgment call does.
What does "blast radius" mean for a wrong output? How far one bad answer travels before someone catches it. A wrong upsell offer stops at one guest's check-in screen. A wrong price on a shared pricing engine can hit every OTA channel at every property in the same minute, before any person has looked at a single number.
EEvidence. What's cheap to check before committing the whole quarter.
The pre-launch trace Dagmara ran at 3:14am is exactly this check: cheap, a few minutes, using numbers the system already had. It's what caught six wrong prices three hours before they'd have gone live everywhere, for a cost of nearly nothing.
Cheap to check, and it settled the whole question on its own, no guess needed about how ready dynamic pricing "felt."
The Garden King suite's quoted rate, night-audit test run, against the $189 contract floor
$210 $150 Contract floor, $189 Flagged, 3:14am 12:00a 1:40a 3:14a
Quoted rateUnder floor, then flagged
The rate crossed under floor by 1:40am. The cheap check that caught it did not run until 3:14am, ninety-four minutes later, and it still beat the 6:14am go-live by three hours.
RRank. The actual scored table, defended.
Weighted one to five across all six lines: personalized upsell offers on top at 4.30, the in-room voice assistant lowest at 2.25. Two pairs tie on raw total, front-desk translation with predictive maintenance, and dynamic pricing with no-show scoring. Both ties break the same way: on guest and revenue lift, the heaviest line, decided before anyone saw a score.
If this rank would look the same no matter what Outcome you'd named in step one, it was picked by gut and the outcome got written afterward to match it.
When totals tie Break it on guest and revenue lift, the single heaviest-weighted line, fixed before anyone sees a score. Front-desk translation (guest lift: 4) beats predictive maintenance (guest lift: 3). Dynamic pricing (guest lift: 2) beats no-show scoring (guest lift: 1). Nobody argues a tie in the room, ever.

One alternative is worth naming and rejecting directly: scoring every candidate purely on projected revenue impact, and ranking by that single number. It lost, because two of Haventhorne's highest-revenue candidates, dynamic pricing and no-show scoring, also carried the least usable signal and the hardest-to-catch failures; a revenue-only rubric would have ranked exactly what nearly went live under floor as the top pick. The AI-specific failure worth naming here is a missing constraint mistaken for a pricing problem: the model wasn't wrong about demand, it had simply never been given the contract floor as a hard rule, so it optimized cleanly toward a number nobody had told it mattered. The guardrail is the dependency and evidence lines themselves: no candidate ships past a low dependency score without a cheap, real check run against it first, the same trace that caught the six rooms at 3:14am. And the trade-off is accepted on purpose: a slower path to the most exciting-sounding revenue feature, in exchange for catching a three-week, contractually-stuck mistake before it ever reached a single OTA channel.

And if you want to be sure it really works, try it somewhere else

Same six lines, a farm equipment rental cooperative instead of a hotel chain, and the honest ranking doesn't change shape.

Cottersgate Equipment Cooperative rents tractors, balers, and irrigation gear to about 400 member farms. Ops lead Domenica Trevanion faced the same fork: three AI candidates for one open engineering slot. A bilingual chat assistant for the rental desk, since half the co-op's seasonal renters speak a first language the desk staff don't. A predictive maintenance system for the tractor fleet, using sensors that don't exist yet. And a late-return risk score, flagging which renters are likely to bring equipment back late, so the desk can double-book shared inventory with more confidence.

Hand sketched quadrant chart titled Sorting Cottersgate's three candidates. X axis, how ready is the evidence, from barely any to well tested. Y axis, hard to walk back if wrong, from easy to adjust to hard to unsay. Bilingual rental chat sits mid-right, evidence fairly ready and easy to adjust. Predictive maintenance sits far left and low, barely any evidence yet and easy to adjust. Late-return risk score sits mid-right and high, evidence fairly ready but hard to unsay.
The item in the top-right corner is the one that needs a silent test before it tells anyone anything.

Same steps, mapped onto Cottersgate. Outcome: protect whether a member actually trusts what the co-op tells them, not whether Cottersgate looks advanced at the annual meeting. Reversibility: a late-return score that wrongly flags a longtime member as high-risk is hard to walk back inside a co-op where everyone knows everyone; a chat assistant that mistranslates once is a quiet, private fix. Dependency: the chat assistant can lean on years of rental-desk scripts and FAQs already written down; predictive maintenance needs sensors Cottersgate hasn't installed yet, so its dependency score stays low until that hardware exists. Evidence: the chat assistant is cheap to check, real bilingual staff can review a sample of its answers in an afternoon; the late-return score is not cheap to check quietly, since testing it for real means actually flagging a member. Rank: the bilingual chat assistant ships first. Predictive maintenance gets queued behind the sensor rollout it depends on, not killed. The late-return score waits for a silent test, scored against real outcomes with no member ever seeing a flag, before it tells the desk anything.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to Rank, name the pick and why, everything else is support.
Cost: there's no budget for a full sensor rollout before the meeting. Say so plainly, and score predictive maintenance honestly low on dependency rather than pretending the sensors already exist.
The model got better, for real: a newer translation model needs far less review before it's trustworthy. Say that too, and move the chat assistant's evidence cost down, which raises its total.

Where people run it wrong.
They skip Dependency because "the model is smart enough to figure it out," when the real gap is that no usable signal exists yet, not that the model is weak.
They score every candidate on the same weights every quarter without ever checking whether guest impact still deserves the heaviest line.
They let the rubric's own tie-break rule get argued case by case instead of fixed in advance, which quietly turns it back into "whoever's louder wins."

How to use it live. Before scoring anything, ask out loud: "does real, usable signal already exist for this, or would we be starting from nothing?" If the honest answer is "starting from nothing," that's the whole Dependency check, and it's reason enough to queue the feature instead of rank it as ready.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits building a rubric that ranks ten AI feature candidates?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Built to score real candidates by what's hardest to undo and whether they're even ready to compete.
2 · THE CAST
Who holds each role in this story, and where do they work?
Tap to flip
ANSWER
Dagmara Grimshaw is VP of Product at Haventhorne Hotels, thirty-four properties. Bohumil Ravenshaw is the CFO who championed dynamic pricing, the candidate that nearly went live under the contracted rate floor.
3 · THE OUTCOME
What does the rubric's Outcome step actually protect?
Tap to flip
ANSWER
Haventhorne's ability to compare ten genuinely different candidates on the same terms, and to catch a mistake before it's live everywhere at once, not whichever pitch sounds biggest.
4 · REVERSIBILITY
Why did dynamic pricing score low on reversibility even though its revenue case was real?
Tap to flip
ANSWER
Once live, it hits every OTA channel at every property at once, and some contracts required 21 days' notice before a wrong rate could even be corrected back up.
5 · THE OLD DECISION
What decision would Dagmara take back?
Tap to flip
ANSWER
Letting the loudest pitch and the biggest sponsor decide the quarter's AI feature, with no shared, written criteria to compare candidates on the same terms.
6 · THE NUMBER
Fill in the blank: six room types sat under the contracted rate floor for about ___ minutes before Dagmara's routine trace caught it, and the flag came ___ before the scheduled go-live.
Tap to flip
ANSWER
About 94 minutes (1:40am to 3:14am); the flag came 3 hours before the 6:14am go-live.
7 · THE RANK
State the actual top and bottom picks, in one line.
Tap to flip
ANSWER
Personalized upsell offers ranks top at 4.30 out of 5. The in-room voice assistant ranks lowest at 2.25, because almost nothing about it has usable signal yet.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs ORDER again on a different organization. Which one, and who runs it?
Tap to flip
ANSWER
Cottersgate Equipment Cooperative, a farm equipment rental co-op. Ops lead Domenica Trevanion ranks three AI candidates the same way.

Check yourself Score: 0 / 0

Multiple choice
1. Why did the in-room voice assistant score only 2 out of 5 on dependency?
  • A. It's too expensive to build.
  • B. No real dataset of guest voice commands from an actual hotel room exists yet, so the signal would need to be built from scratch.
  • C. Guests don't want voice assistants in hotel rooms.
  • D. The hardware hasn't been ordered yet.
Show hint
Check the Dependency letter and the spark box in Let's learn.
Show answer
B. Dependency scores whether real, usable signal already exists, not cost or demand.
True or false
2. True or false: the dynamic pricing near miss happened because the model got less accurate than usual that week.
  • True
  • False
Show hint
Check the story section and the AI-specific failure named in the framework recap.
Show answer
False. The model was never given the contracted rate floor as a constraint. That's a missing rule, not an accuracy dip, which is why more training data alone wouldn't have fixed it.
Fill in the blank
3. Guest and revenue lift is weighted ___ percent, the heaviest single line, and the top-ranked candidate, ___, scored ___ out of 5.
Show hint
Check stage 6 of the walkthrough.
Show answer
25 percent; personalized upsell offers; 4.30 out of 5.
Short answer, name the failure
4. Why couldn't Haventhorne just keep letting the room vote by a show of hands, the way it used to?
Show hint
Check the Outcome letter and the opening of Let's learn.
Show answer
Model answer: A vote rewards whoever pitches most confidently that day, not whichever candidate is actually most ready. It can't tell a well-evidenced idea from a well-rehearsed one, and it has no way to compare a chatbot against a pricing engine on the same terms.
Short answer, apply it yourself
5. Think of a ranked list you've seen decided at work or school, a class project, a roadmap, a budget. What would a real six-line rubric like this one have caught that a group vote missed?
Show hint
Check the Reversibility and Evidence letters, what's hard to undo and what's cheap to check first.
Show answer
Model answer: A rubric forces someone to name what's cheap to check before committing, and how hard a wrong pick is to undo, both of which a group vote skips in favor of whoever argues best in the room.
Short answer, work the number
6. If Haventhorne raised Dependency's weight from 20 percent to 30 percent, taking the extra 10 points from guest and revenue lift, would the top-ranked candidate change? Work it out.
Show hint
Check the Rank step's scores for upsell offers and welcome notes on guest lift and dependency.
Show answer
Model answer: no. Upsell offers scores 4 on both guest and revenue lift and dependency, so moving 10 points from one to the other cancels out, it stays at 4.30. Welcome notes ticks up from 4.05 to 4.15 because its dependency score (4) beats its guest-lift score (3), but it still doesn't catch the leader.
Before you close the answer
Why this works
Tests whether a candidate can build something that ranks the same way twice, not just talk confidently about "prioritization frameworks." It specifically checks whether the criteria are actually about a model, real signal, detection cost, blast radius, rather than a generic feature checklist wearing an AI company's name tag.
Follow-up traps
"Doesn't this just replace one bias, the loudest pitch, with a different bias, whoever picks the weights?" Response: the weights get set before anyone sees a single candidate's score, and only get revisited between quarters, never case by case, so nobody can nudge a weight after the fact to rescue a favorite.

"What happens when two totals genuinely tie?" Response: break it on the single heaviest-weighted line, guest and revenue lift, decided in advance. Dynamic pricing and no-show scoring tied at 2.55; dynamic pricing wins the tie because its guest and revenue lift score, 2, beats no-show scoring's, 1.
If pressed
A dependency score can't go above a 2 out of 5 until the candidate has at least 150 real, labeled examples covering every case it claims to handle, not just the common one. That number is what keeps "the model will probably figure it out" from quietly passing as evidence.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more