InterviewAdvancedAI Opportunity & Model Strategy / Opportunity identification for AI / #23

Pitch me an AI feature for a product I use daily, and tell me why it has not been built yet.

TRACE · a live feature pitch, tested against itself, on Daymark, the calendar app already open on Bettany Ozeri's second monitor and what she found when she checked her own guess before she said it out loud

Daymark is a calendar app. It blocks a person's day into meetings, holds, and travel time, and tries to keep that day honest when something changes underneath it. Bettany Ozeri has used it every workday for three years. Corran Grieveson, running her final interview for a senior AI PM role, pointed at the tab open on her screen and asked her to pitch it one AI feature, then explain why that feature does not exist yet.

The direct answer
Pitch Reflow: an AI feature that rewrites the rest of your day the moment a meeting runs long or falls through, instead of just flagging the clash. It has not shipped because the real block was never the idea. Doing this well means knowing which meetings matter to you, and to whom, not just what time they are at, and no calendar app has ever earned the access or the trust to know that. Daymark's own limited version already proves it: a single-tap suggested move that stops exactly where that judgment would have to start.
Do this, in order
  1. Pitch Reflow, one concrete feature, not "smarter AI scheduling."Why: a vague pitch gives the interviewer nothing they can actually test.
  2. Reject "nobody thought of it" before naming any real cause.Why: it is the nicest answer on offer and almost never the true one.
  3. Name the real block as a trust and data problem, not a hard-to-code problem.Why: rewriting a day well means knowing who matters to you, and no calendar app has ever been given that.
  4. Check whether a smaller, lower-trust version already ships before you claim a total gap.Why: if it does, the company already tried, and stopped somewhere real, not somewhere random.
  5. Run one test: does the small version stop exactly where the trust line would sit?Why: that is the move that separates a real diagnosis from a guess dressed up as one.
  6. Build it in stages of earned trust, not one leap to full autonomy.Why: a full leap on day one is the version most likely to move the one meeting that breaks someone's whole day.

How to answer this, stage by stage

Nobody is grading whether Reflow is a good feature. They are grading whether Bettany can turn a fun thirty-second pitch into a real diagnosis the moment the follow-up lands, instead of just saying "and that would be great."

1
Grab the product the moment it is named, and pick one concrete feature inside it
Say it like this
"Let's use Daymark, since we've both got it open right now. I'm not going to pitch 'better AI scheduling,' that's not a feature, it's a mood. I'll pitch one specific thing it should do, then tell you honestly why I don't think it exists yet."
Why this works
Naming the product instantly and refusing a vague category tells the interviewer this candidate will not hide behind a buzzword.
2
Say the two-part structure out loud before doing either half
Say it like this
"I'll do this in two moves. First the pitch, thirty seconds, one feature. Then the harder half: why it's probably not built. I'll actually reason through that instead of guessing, with a quick timeline, the honest list of real reasons, and one test to pick between them."
Why this works
Naming the two-part shape up front stops the answer collapsing into "and that would be cool," which is where most feature pitches quietly die.
3
Pitch the feature itself, mechanically, not just the outcome
Say it like this
"Call it Reflow. Right now, when my 9 o'clock runs long, Daymark draws a red line where it overlaps my 9:30 and hands me the mess. Reflow would look at the rest of my day, decide which meetings can slide fifteen minutes, which ones are untouchable, and which ones it genuinely isn't sure about, and only ask me about that last group."
Why this works
A mechanic an interviewer can picture, not just a feeling, is what makes a pitch read as a real product idea instead of a wish.
4
Ground it in real, personal evidence before diagnosing anything
Say it like this
"This isn't hypothetical for me. Two Tuesdays ago my 9 o'clock ran to 9:40, and it cost me twelve minutes of dragging four meetings around by hand and texting two people to ask if twenty past worked. That's not rare, it happens on maybe one day in five. And it's not just me: Daymark's own feedback board shows requests for exactly this roughly tripling over the last two years."
Why this works
A number with a timeframe, and a personal stake, turns "I think this would help" into evidence someone can actually check.
5
Say the flattering explanation out loud, then reject it, by name
Say it like this
"The easy answer here is 'nobody thought of it.' I don't believe that for a second. This is one of the most requested features in every calendar app's feedback board, for years. Somebody at Daymark has absolutely proposed this exact thing in a roadmap review already."
Why this works
Naming and refusing the most flattering explanation is the single move that separates a real diagnosis from a compliment dressed up as one.
6
Recut the real candidates, then commit to the one that actually fits
Say it like this
"There's really four honest reasons. One, it's harder to build than it looks from outside. Two, some version already exists and I just haven't found it. Three, it's real and known, but genuinely lower priority than whatever else is on the roadmap. Four, it needs a level of access to my priorities and my relationships that's genuinely hard to ship responsibly. My money's on four. Spotting a fifteen-minute gap is a solved problem. Knowing which fifteen minutes are safe to take from me is not a scheduling problem, it's a trust problem."
Why this works
Listing every honest candidate, then committing to one with a stated reason, is what makes this a diagnosis instead of a shrug.
7
Turn your best guess into a live, checkable test
Say it like this
"Here's a way to actually check that instead of just asserting it. If I'm right, Daymark should already have a smaller, lower-trust version that stops right at the point real judgment would start. And it does. Daymark shipped 'Smart Move' about a year ago: it suggests one replacement slot and makes you tap to accept it, every single time, even for a ten-minute nudge. It never once acts without asking. That's exactly where you'd expect the wall to be if the real problem is trust, not detection."
Why this works
This is the strongest move in the whole answer, because it is the one line that could have proven the guess wrong instead of just supporting it.
8
Close on the one build decision, and the trade-off you are accepting
Say it like this
"So if I owned this: no full autonomy on day one. I'd let Reflow act on its own only for moves under fifteen minutes with a very high confidence score, and ask for everything else, the same way Smart Move already does. Slower to feel magic. But it means the first time it gets something wrong, it's a small thing, not the one meeting I couldn't afford to lose."
Why this works
Ending on a real decision, with the trade-off named plainly, is what an interviewer remembers after the talking stops.

Let's learn

Daymark is a calendar app. It holds a person's meetings, blocks their focus time, and tries to keep the day honest when something changes underneath it.

For most of the time Bettany has used it, when a meeting ran long, Daymark did exactly one thing well: it drew a red line the second two events overlapped. That part took under a second, and it was never wrong. Everything after the red line, deciding what to actually do about the rest of the day, was hers. On a normal overrun, that meant opening three or four events by hand, guessing which ones could slide, and texting someone to ask if twenty past worked. About twelve minutes, on a meeting that only ran forty minutes long.

Hand sketched comparison diagram titled 9:00, then 9:40. Left panel, a document icon labeled The plan, caption 9:00 client call, 9:30 sync, 10:00 interview, 11:00 focus block, all lined up clean. Right panel, a gauge icon labeled 9:40, caption the call runs long, four events collide, twelve minutes of dragging things by hand.
Twelve minutes of manual triage, every time a forty-minute overrun hits a packed morning.

That is not rare. On roughly one weekday in five, something on her calendar runs over or falls through, and the same twelve minutes of dragging things around follows it.

It is not just her. Daymark's own public feedback board carries a version of this request every quarter, and it has not been getting quieter.

Requests for "fix my day automatically," Daymark's feedback board, by quarter
200 100 0 190, still rising Q1 Q2 Q3 Q4 Q5 Q6 Q7 Q8
Requests per quarterMost recent quarter
Sixty-two requests in the first quarter, one hundred ninety in the eighth. That rules out "nobody thought of it" before the diagnosis even starts.
Daymark could see the clash instantly. It had no idea which fifteen minutes were safe to take from her.
Knowledge spark: what is a reversible action? One a person can easily undo without real cost. Moving a meeting by ten minutes is reversible. Cancelling a client's only slot of the week is not. The safer an AI's mistake is to undo, the more freedom it can safely be given to act without asking first.
The choice I would take back Smart Move, Daymark's one real step toward this, was built to ask for a tap on every single change, even a ten-minute nudge nobody would ever object to. That was the right call the day it shipped, when nobody had reason to trust it yet. It is also why Daymark still does not know, a year later, which of those taps people actually wanted skipped, because it never once collected that signal.

What it costs at its worst: skip straight to the flashy version, an AI that silently rewrites your day with no confirmation at all, and the first time it is wrong, it moves the one meeting that mattered, the client call, the school pickup block, the interview itself, and it does it without asking. A tool like that does not get a second chance. It also demos beautifully in a room full of executives exactly once, then gets turned off by the first real user it embarrasses.

What I would leave alone: the red-line conflict flag itself. Spotting two things booked at the same time needs zero judgment about anyone's priorities. Wiring an AI onto that would be solving a problem that is already solved, for free, by a clock.

The lesson: a feature that has to guess what matters to you has to earn the right to know what matters to you first. You cannot build your way past that with a smarter model. You have to be given the access, one small, provably safe decision at a time.

Now here is the same thing as a story

Say the short version above out loud if an interviewer cuts you off early. Read this one for the two weeks that actually built the answer, before Corran ever asked the question.

Bettany Ozeri has run her own calendar carefully for longer than Daymark has existed. Colour-coded blocks, a hard stop for lunch, a standing rule that nothing gets booked before her daughter's drop-off. When she started using Daymark three years ago, the conflict flag felt like magic for about a month. It caught things she would have missed herself on a busy morning, and it caught them in under a second.

Then the magic went ordinary, the way useful things do. First she noticed the flag, fixed the clash, and moved on without thinking about it. Then fixing the clash became a small ritual, the same order every time: check the next meeting, check the one after, text whoever needed twenty minutes back. Then, somewhere in the second year, she stopped noticing she was doing it at all. It was just what a busy Tuesday cost.

Two Tuesdays before the interview, it happened again. Her 9 o'clock ran to 9:40. She was mid-drag on the third meeting, phone in one hand, laptop in the other, when a colleague leaned past her desk and said, half-joking, "Doesn't that thing do this for you yet?"

She laughed it off. Then she didn't stop thinking about it.

Hand sketched timeline titled Two years of Daymark, one gap that never closed. Five milestones left to right: conflict flag ships, catches the clash instantly. Requests for fix my day triple, over 2 years. Smart Move ships, one tap suggested move. No confirmation beta, quiet, 8 months ago small cohort, this milestone emphasized in a different color. Corran asks the question, today.
Five points on the same two years, and the gap that never closed sits right in the middle of them.

Over the next two weeks, in the gaps between real work, Bettany did what she would later do out loud in Corran's office. She checked Daymark's own feedback board and found the trend line climbing, not flat. She checked Daymark's own changelog and found Smart Move, the one-tap suggested reschedule, shipped about a year earlier. And buried three posts down in a product blog nobody reads twice, she found a mention of a ninety-day pilot: a small cohort had tried a version of Smart Move with the confirmation tap removed entirely, silent and automatic.

The extra minutes were never really the problem. Nobody, not Bettany and not Daymark, had ever decided what her day was actually for in a way a machine could act on.
Hand sketched icon list titled Four honest reasons Reflow might not exist. Four numbered rows: one, much harder to build than it looks from outside. Two, a smaller version exists, and I just had not found it. Three, real and known, but genuinely lower priority. Four, needs access to priorities and relationships to ship safely, this row in a different color.
Four honest guesses, written out before any of them got to be the answer.

She laid out all four honestly, the way she'd have to defend them if pushed. Too hard technically didn't sit right, because Daymark had already shipped harder things, multi-calendar merges, travel-time blocking, timezone handling across forty countries. Already exists somewhere she hadn't found turned out to be half true, and led her straight to Smart Move. Real but low priority was possible, but it didn't explain why the pilot that did exist had been quietly shelved rather than expanded. That left the fourth guess standing alone: to move someone's day well, without asking, a model needs to know which meetings are precious and to whom, and Daymark's data model was never built to hold that.

Hand sketched labeled parts diagram titled What Reflow would actually need to know. Center icon a gauge labeled Reflow's call. Four labeled callouts around it: sees meeting times. Sees who is invited. Missing, what matters to you. Missing, who you cannot disappoint.
Everything on the left, Daymark already has. Everything on the right, nobody has ever asked her for.

She could guess at the meeting where Smart Move's always-ask rule got decided, even without sitting in it: a small team, a brand-new feature nobody had earned trust for yet, and a reasonable, cautious call that every move needed a tap. Nobody in that room was wrong. But a year of every tap being logged as "accepted" or "rejected," with no record of which accepts were reluctant and which were instant, meant the one signal that could have justified loosening the rule was never collected in the first place.

Opt-in rate on Daymark's own AI scheduling features, by how much access each one needs
100% 50% 0 94% Conflict flag 68% Smart Move, one tap 11% No-confirmation beta
Calendar structure onlyAdds attendee availabilityNeeds a priority signal Daymark does not have
Ninety-four percent down to eleven, in three steps. The drop tracks how much judgment each feature asks the model to make, not how useful it is.

Before she trusted the fourth guess, she wrote the check she would need to survive Corran asking "how do you know." One alternative would have been to just call it "an execution gap" and move on, since a competitor calendar app had, a few months earlier, demoed a flashy autonomous version at a launch event to visible applause. She turned that explanation down. A demo works once, in a room full of people who will forgive one wrong guess. A calendar has to be right every single ordinary Tuesday, for a person who will not forgive it twice.

Hand sketched decision tree titled Does a smaller, lower-trust version already exist? Root box, check Daymark's own feature list. Three branches: yes and it stops exactly at judgment, leads to confirms trust and data not raw difficulty, this branch in red. Yes but never asks for approval, leads to would point to low priority instead. No smaller version exists anywhere, leads to would point to genuinely too hard or missed.
One question, three possible answers, and only one of them matched what Bettany actually found.

The evidence test was the one thing that could have proven her wrong instead of just agreeing with her: does a smaller, lower-trust version already exist, and does it stop exactly where real judgment would have to start? It did. Smart Move suggests, and asks, every time, for a year, without exception, even for the ten-minute nudges nobody would ever contest. That is not what "too hard" looks like. That is not what "low priority" looks like either, since a low-priority feature does not usually get a careful, cautious confirmation rule built into it on purpose. That is exactly what a company being careful with something it does not yet have permission to do on its own looks like.

By the time Corran asked the question, forty minutes into the interview, Bettany did not need to invent an answer. She had one, tested against her own best attempt to prove it wrong, and she said it in under four minutes, ending on the one build decision she would actually make: high-confidence moves under fifteen minutes go automatic, everything else still asks, the same shape Smart Move already uses, just with a real threshold behind it instead of "always ask, forever."

Hand sketched flow diagram titled Four rungs of trust, and where Daymark stopped. Four connected boxes left to right: flags the clash, suggests one move, you tap to accept, moves your day for you, this last box outlined in a different color to show it has never shipped.
Daymark climbed three rungs over two years. The fourth is the one nobody has been given permission to build.

What I would tell myself, the first time that red line ever saved me four minutes: do not mistake the part that got easy for the part that was actually hard. The hard part had not even been attempted yet.

TRACE, so a pitch survives its own follow-up question

Not a script for sounding clever about scheduling. TRACE is what stops "and that would be great" from passing as an answer, the moment an interviewer asks the one question that actually matters: why doesn't it exist yet?

TTimeline. Lay out what shipped, including things that looked like the whole answer.
Daymark's conflict flag has existed for years and catches every clash instantly. Requests for a fuller fix roughly tripled over the last two years on Daymark's own feedback board. Smart Move, a one-tap suggested reschedule, shipped about a year ago. A quiet, no-confirmation pilot ran about eight months ago to a small cohort, and nothing public followed it. Then, today, the interview.
The gap that matters isn't between "wanting this" and today. It's the eight months between that quiet pilot and right now, with no follow-up anyone outside Daymark ever heard about.
RRecut. Slice apart every honest reason it might not exist.
Too hard to build. Already exists in a smaller form. Real and known, but genuinely lower priority. Needs a level of access to priorities and relationships that's hard to ship responsibly. Four different slices, and only one of them explains why the existing version stops exactly where it does.
A pitch that skips straight to one cause without naming the others isn't a diagnosis, it's a guess with confidence attached.
AAssume nothing. The flattering story does not get to stand in for a real answer.
"Nobody thought of it" is the version that makes every product team look best, and it is also the version a rising two-year request trend rules out almost immediately. Bettany named it and rejected it before reaching for anything else.
The most comfortable explanation is cheap to reach for and almost never survives five minutes of checking.
CCause candidates. Reason through the real one, out loud.
Detecting a fifteen-minute gap is a solved problem, Daymark solves harder scheduling math elsewhere already. Deciding which fifteen minutes are safe to take from a specific person, without asking, needs to know what matters to them and to whom. That's not a scheduling problem. That's a trust and data-access problem, the kind that only makes sense for a feature a model has to judge, not just compute.
Only this candidate explains the specific shape of the gap, not just the fact that a gap exists.
EEvidence test. The one check that could have proven the guess wrong.
Does a smaller, lower-trust version already exist, and does it stop exactly at the point real judgment would start? Smart Move does exactly that: suggests, always asks, never once acts alone, for a full year. That's the signature of a trust wall, not a technical one.
This is the strongest move in the whole framework. It's the one line that risks being wrong, instead of just repeating the guess with more confidence.

Three things worth saying plainly, since interviewers push here. The rejected alternative: treating a competitor's flashy, no-confirmation demo as proof the gap is "just execution," which Bettany turned down because a demo only has to survive one room once, and a calendar has to survive every ordinary Tuesday for a person who will not forgive a wrong silent move twice. The AI-specific risk, named by name: an autonomous move made with total confidence and no explanation, moving the one meeting that actually mattered, is a silent failure that looks identical whether it's right or wrong until the damage is already done, so any real build attaches a plain "why I moved this" note to every automatic change, and ships only once a move clears a set bar on a labeled eval set of real move-or-don't-move judgments, not a promise to always get it right. And the trade-off, accepted on purpose: acting automatically only on high-confidence, easily reversible moves is slower to feel magical than full autonomy, and worth it, because the alternative's one wrong guess costs the whole relationship, not just one meeting.

And if you want to be sure it really works, try it somewhere else

Same five letters, a library cataloguing desk instead of a calendar app. This time the trust wall isn't about a person's priorities. It's about what happens when a mistake can't be undone at all.

Cartouche is cataloguing software a public library uses to log, classify, and shelve incoming books. Ashling Marrash catalogues donations there five days a week. Pitched the same question, cold, in her own interview, she names AutoShelf: an AI feature that classifies and shelves an incoming donation end to end, no review, the moment it's scanned.

Mapped onto TRACE: timeline is two years of requests for "just let it file itself" on Cartouche's own forum, and a "suggested subject headings, one click to accept" feature that shipped fourteen months ago and has never been extended past a suggestion. Recut is the same four honest slices: too hard, already exists smaller, known but low priority, or too sensitive to hand over fully. Assume nothing rules out "nobody thought of it," the same way a rising request count did for Reflow. Cause candidates land somewhere adjacent but not identical: classifying an ordinary paperback donation is a solved problem, Cartouche already does it well. Deciding, alone, to shelve a fragile local-history pamphlet or a first edition under the wrong heading is a mistake that might never get found and corrected, because almost nobody browses that shelf by accident. Evidence test: the suggested-headings feature stops exactly at rare and local-history items, where a human still makes the final call every time, and sails right through ordinary donations without a second look.

Hand sketched quadrant diagram titled Sorting Cartouche's backlog by what is actually hard about it. X axis how easy to undo, from easy to fix to hard to fix. Y axis how much judgment it takes, from rule based to needs a librarian's eye. Four points plotted: typo in a title field, low on both axes. Suggested subject heading, medium judgment, easy to undo. Ordinary donation full auto file, medium on both. Rare or local history item, high on both axes, top right corner.
The item furthest from safe to automate is also the one a library can least afford to get wrong.

Swap the trigger and it still runs.
Speed: an interviewer caps the whole answer at ninety seconds. Skip straight to the pitch and the evidence test; the timeline and the four-way recut are what you'd add back if asked to go deeper.
Cost: there's no feedback board to check live. Say what you'd check instead: "I'd pull the last year of support tickets that mention this by hand, and see if the shape of the ask has changed."
The model got better, for real: say Daymark's scheduling engine gets meaningfully smarter at guessing free time. The evidence test still runs exactly the same way, a better engine changes how good the guesses would be, not whether the model has been given permission to act on them alone.

Where people run it wrong.
They stop at "nobody thought of it," because it's flattering to imagine you found a gap everyone else missed.
They find one real reason and stop looking, instead of naming all four honestly before picking.
They treat a competitor's demo as proof the whole problem is solved, without checking whether it holds up under a real, repeat, no-audience Tuesday.

How to use it live. The moment an interviewer asks "why hasn't this been built," buy two seconds by saying the two-part structure out loud first: "let me pitch it, then actually diagnose that, rather than guess." Saying the plan is often the calmest four seconds in the whole answer, and it's free.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits pitching a feature live, and then diagnosing why it isn't built?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Used here to diagnose why a plausible feature hasn't shipped, not why a metric dropped.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Bettany Ozeri, the candidate pitching Reflow. Corran Grieveson, the interviewer who pointed at Daymark on her screen and asked the question.
3 · THE TIMELINE
When did the real evidence for Reflow actually start piling up?
Tap to flip
ANSWER
Feature requests roughly tripled over two years. Smart Move shipped about a year ago as a partial answer. A quiet no-confirmation beta ran about eight months ago to a small cohort, with no public follow-up since.
4 · THE RECUT
What four honest reasons could explain why Reflow isn't built?
Tap to flip
ANSWER
Too hard technically. Already exists in a smaller form. Known but genuinely low priority. Too sensitive to ship responsibly with today's access. The diagnosis lands on the fourth.
5 · THE OLD DECISION
What decision would Bettany take back?
Tap to flip
ANSWER
Smart Move was built to ask for a tap on every single change, even a harmless ten-minute nudge. Right call with zero trust built up. It also meant Daymark never collected the signal on which moves people actually wanted skipped.
6 · THE NUMBER
Fill in the blank: about ___ percent of users kept Smart Move's one-tap suggestions on. Only about ___ percent opted into the no-confirmation autonomous beta.
Tap to flip
ANSWER
68 percent. 11 percent. That gap, not a guess, is the evidence for a trust problem, not a technical one.
7 · THE EVIDENCE TEST
What's the one test that actually separates "too hard" from "not trusted yet"?
Tap to flip
ANSWER
Check whether a smaller, lower-trust version already exists and stops exactly at the point real judgment would start. Smart Move does. That confirms the block is trust, not detection.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what's the parallel finding?
Tap to flip
ANSWER
Cartouche, a library cataloguing tool. Ashling Marrash finds the same shape of gap: autonomous shelving stalls specifically on rare and local-history items, where a wrong call is hard to undo, not on ordinary donations.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these is the strongest evidence that Reflow's real blocker is trust, not raw technical difficulty?
  • A. Reflow sounds complicated to build.
  • B. Daymark already ships Smart Move, a smaller version that stops exactly at the point autonomous judgment would start.
  • C. Feature requests for it exist on the feedback board.
  • D. A competitor has announced something similar.
Show hint
Check the evidence test step in the TRACE recap.
Show answer
B. A and C are context, not proof, and D is the explanation Bettany explicitly rejected. Only B is a checkable test that could have come back the other way.
True or false
2. True or false: because feature requests for this tripled over two years, that alone proves Daymark simply never thought of building it.
  • True
  • False
Show hint
Look at what Smart Move and the quiet beta actually prove about whether Daymark had considered this.
Show answer
False. Rising demand is evidence against "nobody thought of it," not proof of it. Smart Move and the eight-month-old beta both show Daymark had clearly thought about it and tried a version.
Fill in the blank
3. Fill in the blank: Smart Move keeps about ___ percent of users opted in. The no-confirmation autonomous beta was kept on by only about ___ percent.
Show hint
Check the bar chart in the story section, and flashcard 6.
Show answer
68. 11. The drop tracks how much judgment each feature asks the model to make, not how useful either one is.
Short answer, where it wouldn't matter
4. Name one part of Daymark that should NOT get an AI judgment layer, based on this diagnosis. Why not?
Show hint
Look at "What I would leave alone" in Let's learn.
Show answer
Model answer: The red-line conflict flag itself. Spotting two overlapping events needs no judgment about anyone's priorities, it's a solved, deterministic problem already handled for free by a clock.
Short answer, apply it yourself
5. Think of a product you use daily that already ships a smaller, lower-trust version of some bigger, more autonomous feature you wish it had. What would the evidence test say about why the bigger version isn't built?
Show hint
Check whether the smaller version still asks for your approval on even the tiniest, most obviously safe action.
Show answer
Model answer: If the smaller version exists and still requires approval every single time, even for actions nobody would ever object to, that's a sign the real block is earned trust in the model's judgment, not the underlying technical capability.
Short answer, the number question
6. If the no-confirmation beta's opt-in rate had come back at 55 percent instead of 11, would Bettany's diagnosis still hold? Why or why not?
Show hint
Think about what the low number is actually standing in for.
Show answer
Model answer: Probably not, or at least much weaker. A 55 percent opt-in would suggest most people trusted it fine, pointing more toward "still a bit rough" or "low priority to expand" than a deep trust wall. The diagnosis leans hard on that number being genuinely low.
Before you close the answer
Why this works
Tests whether a candidate can turn a fun, easy pitch into a real diagnosis the moment a real follow-up lands, and whether they'll reach for "nobody thought of it," the story that flatters everyone in the room, instead of checking it.
Follow-up traps
"Couldn't this just be a hard machine learning problem, nothing to do with trust?" Response: Smart Move already solves the easier detection problem well, so if it were purely a hard modeling problem, the simpler shipped version wouldn't stop precisely at the point autonomy would need judgment. It stops exactly there.

"If a competitor ships the full autonomous version first, doesn't that prove it's just an execution gap, not a trust one?" Response: A flashy demo only has to survive one room, once, with an audience ready to forgive a wrong guess. A calendar has to survive every ordinary Tuesday for a person who won't forgive a wrong silent move twice, and one bad move ends the feature for that user for good.
If pressed
A real build wouldn't ship on a single global confidence number. It would need its own eval set of real move-or-don't-move judgments, labeled by actual users after the fact, not synthetic scenarios, covering recurring versus one-off meetings, internal versus external attendees, and how often an invite gets declined versus silently ignored, before any auto-accept threshold went anywhere near a real calendar.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more