CaseAdvancedModel Fluency & the AI PM Role / Working with ML engineers and researchers / #5

How do you keep a research team connected to user problems without constraining their exploration?

ORDER · which connection mechanism Nikolina Grewal's research pod gets first at Emberfield Systems

Emberfield Systems builds Wattloom, an app that quietly shifts a home's dryer, dishwasher, and EV charging into cheaper electricity hours and nudges the thermostat to match, without the homeowner noticing a difference in comfort. Nikolina Grewal leads the applied research pod chasing a harder version of the same idea: guessing which appliance is running from the smart meter alone, so Wattloom can work in the millions of homes that only have one meter. A product manager wants to embed one of Nikolina's researchers into Wattloom's two-week sprint, the same fix a sibling pod tried six months ago. Hjalmar Askeland, VP of Product, has to decide what Nikolina's pod actually gets first.

The direct answer
Give the research pod low-friction, self-serve access to real household usage data and real complaint text, not a sanitized summary. Pair it with an optional monthly drop-in session where anyone can bring a real problem to the room. Hold sprint-embedding in reserve, since it is the one option nobody can cleanly undo, and reject a curated quarterly deck outright, since a filtered summary is the exact failure this whole decision exists to avoid.
Do this, in order
  1. Give the research pod self-serve access to real usage data and real complaint text, not a filtered summary.Why: every other option on the table quietly depends on this being true first.
  2. Add an optional monthly drop-in session alongside it.Why: it gives the human layer without asking anyone to give up sprint capacity.
  3. Hold sprint-embedding until the lighter options prove they aren't enough.Why: once other teams start planning around a researcher's time, pulling back reads as the pod refusing to help.
  4. Reject the quarterly curated deck outright, don't build it at all.Why: a sibling pod already lived on exactly this kind of sanitized summary, and it broke on account types nobody's roadmap ever mentioned.
  5. Watch voluntary use, not attendance, as the real signal.Why: it's the cheap check that shows whether a mechanism is actually working before anything heavier gets built.
  6. Leave infrastructure-only research, like on-device inference speed, out of this decision entirely.Why: it has no real household-behavior question in it, so a user-connection mechanism protects nothing there.

How to answer this, stage by stage

Nobody is grading whether you can name four ways to keep a team aligned. They're grading whether you can tell a real connection from one that only performs connection, and defend the order you'd build them in.

1
Scope it to one pod and one real ask on the table
Say it like this
"Let's put this on one team. Emberfield Systems builds Wattloom, a home energy app. Nikolina Grewal runs the applied research pod chasing a new load-disaggregation model. A product manager just asked to embed one of her researchers into Wattloom's sprint, because a sibling pod did that and got praised for it."
Why this works
A named team and a real ask stop the answer from floating in the abstract.
2
Answer the literal question before ranking anything
Say it like this
"You keep a research pod connected by making the real, messy household data cheap and voluntary to reach, not by pulling researchers into someone else's roadmap. Exploration doesn't get constrained by contact with real problems. It gets constrained by contact with someone else's calendar."
Why this works
This is the reframe the rest of the answer defends. Skip it and the ranking looks like a random preference.
3
Say your structure out loud
Say it like this
"I'll run this as ORDER. Outcome, what the whole ranking is protecting. Reversibility, which option is hardest to walk back. Dependency, what actually has to be true first. Evidence, what's cheap to check before committing. Rank, the real call, defended."
Why this works
Two seconds of structure before four options land tells the interviewer you have a method, not just an opinion.
4
Name the outcome, so the rank isn't a guess
Say it like this
"Every option here is fighting for the same thing: does the pod's next model actually work on the messy homes, the solar-and-battery ones, the window-AC ones, without the pod turning into a second engineering team that only ships whatever's already on the roadmap."
Why this works
Naming the outcome up front stops "the PM asked for it" from quietly becoming the tiebreaker.
5
Give the ranked call, committed
Say it like this
"So here's the call. Ship self-serve access to the real, raw usage data and complaint text first, no approval needed to look at it. Pair it with a monthly drop-in session, optional, nobody has to attend. Hold the sprint embedding. Don't build the quarterly curated deck at all."
Why this works
This is the direct answer, said plainly, before any number shows up to defend it.
6
Prove it with a real precedent, and name what got rejected
Say it like this
"Six months ago a sibling pod got embedded straight into the billing team's sprint. For two quarters that looked like alignment, every ticket landed, every demo shipped. But they never touched a real bill dispute, only the billing team's own dashboard, and the model they shipped broke immediately on account types nobody on that roadmap had ever mentioned. I also looked at just handing Nikolina's pod a quarterly summary deck instead, and rejected it. Same failure, just administered slower."
Why this works
A real precedent inside the same company beats an abstract worry, and naming the rejected option before anyone asks shows real judgment.
7
Close on the one line anyone could go check
Say it like this
"You'll know it's working when researchers pull the raw ticket archive without being told to, not when everyone shows up to a meeting they can't skip. That's the number I'd actually watch: voluntary use, week over week, not attendance."
Why this works
Ends on something an interviewer could actually go verify later, not just a confident-sounding pick.

Let's learn

The quarterly deck lived in a shared folder called Research_Alignment_FINAL_v3, four pages, the same four category tags every time: appliance issue, schedule mismatch, comfort complaint, billing question. It was laminated once, that's how permanent it felt.

Wattloom is Emberfield's app for keeping a home's electricity use cheap without anyone having to think about it. It quietly delays a dryer cycle by forty minutes, moves a dishwasher run past the evening price spike, and shifts an EV charge to the middle of the night, all while holding the thermostat steady. It works because it can already tell, from separate submeters, exactly which appliance is drawing power. Most homes don't have those submeters. They have one meter, for the whole house.

So Nikolina's pod is chasing a harder model: guess which appliance is running from the single smart-meter stream alone, no submeters needed. That only works if the model has actually seen the messy real cases: a window air conditioner cycling on its own schedule, a solar panel and battery quietly cancelling out half the signal, a space heater running at 2am in a house with an elderly resident who keeps odd hours. None of that shows up in a category tag that just says "appliance issue."

Hand sketched labeled parts diagram titled What used to reach the research pod. A center document icon labeled Quarterly Deck, with four labels radiating outward: Category tags only. No raw complaint text. Curated by product marketing. Reviewed once a quarter.
This was the only door into real households the pod had, for a year, and nobody had built it to be a door.

Here's the turn. A product manager wants to fix this by embedding one of Nikolina's researchers directly into Wattloom's two-week sprint, full time, same as a sibling pod did six months ago for a dynamic-pricing model. It sounds like the obvious fix: get research physically closer to the product. But closeness to a calendar is not the same thing as closeness to a real household. Ruaridh Tindale's pod proved that already, the expensive way.

Where the disaggregation model missed the appliance, by household type
80% 40% 0% 58% 21% Window AC 71% 26% Solar and battery 44% 17% Overlap homes
Trained on the sanitized category summaryTrained on raw usage and complaint access
The model doesn't need to be perfect to ship. It needs to clear an agreed miss rate on a held-out set of real, messy homes, not the tidy ones the summary kept describing.

Ship it blind and the cost isn't a slightly worse model. Nikolina's pod could spend a whole quarter refining a beautifully engineered version of a model that only works on the plain, single-meter, no-solar homes, exactly the homes where the old, simpler model already worked fine. Every home that actually needed the new model, the ones with solar, with window units, with two appliances running at once, keeps guessing wrong, the same way Ruaridh's billing model broke on the accounts nobody on his roadmap had ever mentioned.

We didn't cut the research pod off from real households. We just built the only door in through product marketing's desk.
The choice I would take back A year ago, every support ticket got filtered down to one of four category tags before it left the support tool. That was the right call at the time: it protected customer privacy and kept the support tooling simple, and nobody outside support needed more than a category count. Nobody rebuilt that pipe once a research pod actually needed the real words in a complaint, not just its label.

What I would leave alone: a separate research thread inside the same pod is testing whether a smaller version of the model can run directly on the thermostat's own chip, to cut how long a recommendation takes to appear. That's a pure speed question. No household behavior in it at all, so none of this connection question applies there.

The lesson: access is not the same claim as calendar. A pod can sit in every meeting in the building and still never touch what's real, if what reaches the room has already been rounded off before anyone in it opens their mouth.

Now here is the same thing as a story

The short version above is what you'd actually say out loud. Read this one for the two quarters a whole research pod spent believing it was finally aligned.

Emberfield's applied research floor is quiet by six most nights, except for the two weeks a quarter when demo prep runs late and someone's ordering in food at eight. Nikolina Grewal has run the load-disaggregation pod for a year and a half, four researchers, one open question: can a single smart-meter stream tell you what's running in a house, without a single extra sensor.

Down the hall, Ruaridh Tindale ran a different pod, chasing a model that could predict when a household was about to get a shockingly high bill, before it happened. Six months ago, his pod got what everyone called a win: full seats on the billing team's two-week sprint. Standups, tickets, demos, all of it. For a while it genuinely felt like alignment. Ruaridh told people it was the best thing that had happened to his research in a year.

Hand sketched numbered icon list titled What sprint-embedding cost Ruaridh's pod. Four rows, each an icon and a line of text: one, a person icon, Own research questions, gone. Two, a gauge icon, Two quarters, no new finding. Three, a document icon, Only roadmap tickets got worked. Four, a scale icon, Model broke on real bill data.
Nobody told Ruaridh's pod to stop exploring. Their calendar just never had room left for it.

What actually happened was smaller and slower. Every hour of every sprint already had a ticket number attached to it. Nobody on the pod had time left over to go read a real, messy bill dispute, the kind where a customer calls furious about a charge nobody can quite explain. Instead, the pod worked from the billing team's own success dashboard: clean numbers, clean cohorts, a pilot group of suburban single-family homes on simple flat-rate plans.

Ruaridh's pod didn't stop exploring because anyone told them to. They stopped because every hour already had a ticket number on it.

The bill-shock model shipped. It worked beautifully on the pilot cohort. Then it hit real accounts: apartment buildings on shared meters, households that switched pricing plans mid-month, customers on a low-income assistance program with its own separate billing rules. Nobody on the pod had touched a single account like that in two quarters. Not because anyone hid those accounts. They just never showed up on a billing-team roadmap, so nobody in the room ever had a reason to mention them.

Now, in the next building over, a Wattloom product manager is proposing the exact same fix for Nikolina's pod: embed a researcher full time, same as Ruaridh's team. Hjalmar Askeland, watching what happened down the hall, isn't ready to say yes just because it looks like the fastest way to show commitment.

Hand sketched two panel comparison titled One sentence became four words. Left panel, a green tinted document icon labeled The raw complaint, caption dryer keeps kicking off mid-cycle. Right panel, a red orange tinted document icon labeled What research received, caption appliance issue, other.
Valdis Pellegrino's actual ticket, and the four words that were all Nikolina's pod ever saw of it.

He pulls one real example to make the point concrete. Valdis Pellegrino, a Wattloom user, filed a ticket that read: "my dryer keeps kicking off mid-cycle, right when it's supposed to be running on the cheap window." Somewhere in the support tool that became "Appliance issue, other." That's the version that would have reached even an embedded researcher, unless the actual pipe carrying real language into the room got fixed first. Sitting in the sprint room doesn't fix a filter sitting upstream of the sprint room.

So here's what Hjalmar decides not to do: build the quarterly curated deck that a product marketing colleague half-offered as a compromise. Same failure as Ruaridh's dashboard, just delivered slower and with nicer formatting. And here's what he does instead. Nikolina's pod gets self-serve, low-friction access to the real ticket text and the real usage traces, anonymized by account, not by content. Alongside it, an optional monthly drop-in slot on Wattloom's calendar, no one required to attend, no notes assigned, just an open half hour where a real problem or a real research question can get put on the table. Sprint embedding stays on the shelf, not thrown out, just not reached for first, because it's the one move nobody can cleanly take back once it's running.

What I'd tell myself, watching Ruaridh's pod get praised for finally being aligned: a seat in the room was never the thing that mattered. What reached the seat was.

ORDER, for deciding what actually earns a research pod's trust first

PICK would fit if this were only two options and an even trade. But there are four candidates here, and the real question is which one earns the quarter first, and which one you hold back. That's ORDER's job.

OOutcome. What the rank actually has to protect.
All four options are competing for the same thing: whether the pod's next model actually works on the messy real homes, the solar-and-battery ones, the window-AC ones, without the pod quietly turning into a second engineering team that only ships whatever's already committed to the roadmap. Not "look aligned." A model that clears its miss-rate bar on real, ugly households, and a pod that still gets to chase its own next question.
Name the outcome before ranking anything. Skip this and the rank is just whichever option the PM asked for loudest.
Hand sketched flow diagram titled What has to unblock what. Four boxes connected by arrows, left to right, the first box highlighted in green: Raw data access. Real directions found. Sessions earn trust. Embed, if needed.
Every box after the first one only works if the first one actually happened.
RReversibility. Which choice is hardest to undo.
A monthly drop-in session is easy to cut. Skip a month, nobody's plan breaks, no promise gets broken. Sprint embedding is different. Once a researcher owns tickets inside a two-week cadence, other teams start planning their own capacity around that researcher being there. Pulling them back out, even after one bad quarter, reads as the research pod refusing to help, and that damages trust in a way that outlasts the org chart that caused it.
This is why order matters, not preference. One mistake costs a month. The other one is load bearing for how the whole team plans, the moment it ships.
Hand sketched two panel comparison titled One is easy to cut, one is not. Left panel, a green tinted document icon labeled Drop-in problem session, caption optional, skip a month, nothing breaks. Right panel, a red orange tinted person icon labeled Embedded in the sprint, caption pull them out and trust breaks with it.
Reversibility isn't a reason to avoid the harder one forever. It's a reason to be sure of it first.
DDependency. What has to already be true.
For any of these mechanisms to actually work, research needs the real, raw language in a complaint and the real usage traces behind it, not a category tag standing in for both. This is the trap Ruaridh's pod fell into with a seat in every single sprint: proximity to the room was never the same thing as access to the real signal. A researcher sitting three feet from a product manager, reading the same sanitized dashboard, learns exactly as little as a researcher reading a quarterly deck from across the building.
This is why the order isn't a guess about which mechanism looks the most committed. Embedding without fixing the data pipe underneath it doesn't fix anything, it just moves the same gap into a more expensive room.
EEvidence. What's cheap to check first.
Before committing to anything heavier, Hjalmar's team shipped the self-serve raw-data catalog and the drop-in session, both cheap, both reversible, and watched what happened. Ruaridh's old mandatory monthly sync sat at 100 percent attendance every single month, because it was mandatory. The self-serve catalog had no such rule behind it, and its use climbed anyway, from a fifth of the pod querying it in week one to nearly all of them by week six. Attendance told Hjalmar nothing. Voluntary use told him everything.
Cheap, and it's the number that actually settled it, not whichever option looked the most like commitment on a slide.
Hand sketched timeline titled The research pod's one quarter, spent. Four milestones along a horizontal line, the first one highlighted in green: week one, data access ships, forced, blocks the rest. Week three, drop-in sessions start, chosen, not forced. Week six, usage checked, who actually queried it. Week ten, embedding decision, only if still needed.
Only the first milestone was ever forced. Everything after it was a choice, checked against real use before the next one got built.
Share of the research pod using each connection point, week by week
100% 50% 0% mandatory sync, always full crosses 80%, week 6 Week 1 Week 5 Week 8
Voluntary raw-data catalog queriesMandatory sync attendance
The flat line looks perfect and says nothing. The climbing line is the one that told Hjalmar the pod actually wanted real contact, not a box to check.
RRank. The actual call, defended.
Self-serve raw-data and complaint access ships first, this month, no approval gate to look at it. The optional monthly drop-in session ships alongside it. Sprint embedding stays held, not cancelled, revisited only if the lighter options genuinely fail to close the gap. The quarterly curated deck never gets built at all.
If this rank would be identical with a different outcome in the O step, say "make research look aligned on a slide," it was picked by pressure, not judgment. Swap the outcome to "grow trust between research and product as fast as possible," and the rank still holds, because a mechanism that quietly breaks on real accounts doesn't build trust, it just delays the moment it gets spent.

Three things worth stating directly, since the real judgment sits here. The alternative worth naming and rejecting is the quarterly curated deck, built by product marketing as a middle-ground compromise. It loses because Ruaridh's pod already ran that exact experiment with a live sprint seat standing in for the deck, and a sanitized summary is a sanitized summary no matter how it's delivered or how often it's refreshed. The AI-specific failure worth naming is silent relevance drift: a research pod's whole sense of what's worth exploring quietly stops matching real households, because its own evals got built from the same filtered categories feeding everything else. The guardrail is a standing check, a monthly sample of raw tickets pulled against whatever the pod currently believes its top problems are, so a mismatch shows up as a number, not as a broken model three months later. And the trade-off is accepted on purpose: raw access costs real engineering, an anonymization pipeline, a privacy review, ongoing upkeep, in exchange for a pod that isn't quietly guessing at problems from a summary, which is the more expensive failure and the harder one to see coming.

And if you want to be sure it really works, try it somewhere else

Same five letters, a farm instead of a home, and this time the sanitized summary wears a different name: the quarterly agronomy report.

Loambrook Agritech builds Cropscope, a tool that predicts when a field needs water from soil-moisture sensors and satellite imagery, so a farm can irrigate the right acres on the right day instead of running every line on the same fixed schedule. Svetlana Kirkland, Head of R&D, is deciding what her irrigation-timing research team gets, after a regional agronomist asked to have a data scientist join the farm-ops team's weekly rotation full time.

Hand sketched quadrant chart titled Sorting Cropscope's four options. X axis, cost to undo, from cheap to hard. Y axis, how close to real fields, from filtered to raw. Drop-in field notes sits cheap and moderately raw. Self-serve sensor and ticket access sits cheap and very raw. Embedded in irrigation sprints sits hard to undo and moderately raw. Quarterly agronomy report sits cheap to undo but heavily filtered.
The report is cheap to cancel and useless anyway. The two options worth keeping sit in the top half of this chart, not the bottom.

Same steps, mapped onto Loambrook. Outcome: protect whether an irrigation recommendation actually holds up on real, mixed-crop, mixed-soil farms, not just the one tidy pilot farm the model was tuned against. Reversibility: a drop-in field-notes session, dropped in on a farm-ops call once a month, is trivial to cancel; a data scientist rotated into weekly farm-ops planning is not, since crews start building their week around that person showing up. Dependency: the model only earns trust if it sees raw sensor drift and the agronomist's actual field notes, not the quarterly agronomy report's rounded-off "irrigation needs improved in the northeast block." Evidence: Svetlana's team shipped self-serve sensor and note access first and watched who actually opened it without being asked, exactly the same voluntary-use test Hjalmar ran. Rank: self-serve raw sensor and field-note access ships first, the drop-in session rides alongside it, full rotation into farm-ops stays held, and the quarterly agronomy report gets cancelled rather than kept as a fallback.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the rank: ship the raw-access point first, it's cheap and it's the one thing every other option secretly needs anyway.
Cost: there's no budget to properly anonymize the raw complaint text this quarter. Ship a visibly reduced version, real language kept, only the account identifiers stripped, and say so plainly, rather than quietly shipping the old flattened category tags again under a new name.
The model got better, for real: say the lab accuracy on the disaggregation model already looks great. Keep the same access gate anyway. Lab accuracy on curated data was never the same claim as accuracy on a real, messy home.

Where people run it wrong.
They pick the option that photographs well on a roadmap slide, an embedded headcount, over the one that actually fixes the data gap underneath it.
They treat a meeting's attendance as proof of connection, instead of checking who used it when nobody made them.
They let "protect privacy" quietly turn into "strip out everything specific," when the real fix is anonymizing the identifiers and keeping the actual words intact.

How to use it live. Before ranking anything, ask out loud: "which of these can I undo next quarter without anyone getting hurt, and which one, once it's running, is nobody going to want to be the one who cancels it?" Whichever answer feels uncomfortable to say is usually the one to rank last, not first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
Which framework fits ranking which connection mechanism a research pod gets first?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Built for ranking real candidates by what's hardest to undo and what's actually ready, exactly what "which mechanism first" needs.
2 · THE CAST
Who holds each role in this story, and where do they work?
Tap to flip
ANSWER
Hjalmar Askeland is VP of Product at Emberfield Systems, which builds Wattloom. Nikolina Grewal leads the applied research pod. Ruaridh Tindale led the sibling pod whose sprint embedding broke on real billing accounts.
3 · THE OUTCOME
What does the rank actually have to protect?
Tap to flip
ANSWER
Whether the pod's next model works on the messy real homes, without the pod becoming a second engineering team that only ships roadmap tickets. Not "look aligned."
4 · THE DEPENDENCY
What has to already be true before any connection mechanism actually works?
Tap to flip
ANSWER
Research needs the real, raw usage data and the real words in a complaint, not a category tag. Ruaridh's pod had a seat in every sprint and still never touched a real dispute ticket.
5 · THE OLD DECISION
What decision would Hjalmar take back, and why did it make sense at the time?
Tap to flip
ANSWER
Filtering every support ticket down to one of four category tags before it left the support tool, a year ago, to protect privacy and keep the tooling simple. Right then; nobody rebuilt the pipe once research needed the real words.
6 · THE NUMBER
Fill in the blank: on window-AC homes, raw data and complaint access cut the appliance-miss rate from ___% to ___%.
Tap to flip
ANSWER
From 58% down to 21%, the biggest single gap of the three household types tested.
7 · THE RANK
State the final call, defended in one line.
Tap to flip
ANSWER
Self-serve raw-data and complaint access ships first, with an optional drop-in session alongside it. Sprint embedding stays held, and the quarterly curated deck never gets built, because a filtered summary already failed once, down the hall.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what plays the role of the rejected curated deck there?
Tap to flip
ANSWER
Cropscope, Loambrook Agritech's irrigation-timing tool. The quarterly agronomy report plays the failing role, the same rounded-off, filtered stand-in for real field data.

Check yourself Score: 0 / 0

Fill in the blank
1. Raw data and complaint access cut the window-AC appliance-miss rate from ___% down to ___%.
Show hint
Check the grouped bar chart under "Let's learn."
Show answer
58% down to 21%. Same underlying architecture, the only thing that changed was whether the training data included the messy real homes or just the sanitized categories.
Multiple choice
2. Why doesn't a quarterly curated deck fix the same gap that sprint embedding is supposed to fix?
  • A. It costs too much for product marketing to produce every quarter.
  • B. Its own filtering step, done to protect privacy and keep it short, is the actual gap, so what reaches research still isn't the real signal.
  • C. Research pods are contractually barred from reading marketing-authored documents.
  • D. It only ever covers the most recent month of tickets.
Show hint
Check the Dependency (D) step and what happened to Ruaridh's pod.
Show answer
B. A filtered summary is a filtered summary no matter how it's delivered. Ruaridh's pod proved this already, with a live sprint seat standing in for the deck and the exact same gap underneath it.
True or false
3. True or false: once the self-serve data catalog and drop-in sessions are working well, sprint embedding is permanently off the table.
  • True
  • False
Show hint
Check the Rank (R) step of the ORDER recap.
Show answer
False. It's held, not cancelled. It's the option evidence could still justify later. It's just not the one you reach for first, because it's the hardest of the four to walk back.
Short answer, name the old decision
4. What old decision would Hjalmar take back, and why did it make sense when it was made?
Show hint
Look at the "choice I would take back" key point in "Let's learn."
Show answer
Model answer: Filtering every support ticket down to one of four category tags before it left the support tool, a year ago, to protect customer privacy and keep the tooling simple. It made sense before any team outside support needed more than a count. Nobody rebuilt the pipe once research needed the real language.
Short answer, apply it yourself
5. Pick a team you know that's meant to explore, not just execute. What's one lightweight, voluntary way you could hand it real raw inputs instead of a sanitized summary?
Show hint
Look for the thing that's cheap to build, cheap to cancel, and doesn't need anyone's sign-off to use.
Show answer
Model answer: A design research team that only ever sees usability-test highlight reels could instead get a rolling, self-serve library of full raw session recordings, tagged but not trimmed, with no requirement to watch any of it. Whether they open it unprompted tells you if the connection is real.
Short answer, work the number
6. If the self-serve catalog's voluntary query rate had stalled at 20% instead of climbing to 90% by week eight, would the same rank still hold?
Show hint
Check the Evidence (E) step, and what the flat line versus the climbing line was actually testing.
Show answer
Model answer: no, not automatically. A stalled rate is exactly the signal the Evidence step exists to catch. It wouldn't prove embedding is the fix, but it would mean the lighter mechanism isn't landing, and that's worth understanding, low friction alone versus a pod that's genuinely stopped caring, before building anything heavier.
Before you close the answer
Why this works
Tests whether you can tell voluntary, genuine access from a mechanism that only performs alignment, and whether you'll commit to the reversible first step instead of reaching for the heaviest option because it looks like the most commitment. Most candidates stop at "put them in the room."
Follow-up traps
"Why not just embed the researcher, it's the fastest way to guarantee alignment?" Response: fastest isn't free. Once a researcher owns sprint tickets, other teams plan around that capacity, and Ruaridh's pod already proved embedding alone doesn't even fix the real gap, since the researcher still only saw the billing team's own dashboard.

"Isn't handing researchers real complaint text a privacy problem?" Response: yes, and that's the accepted trade-off. Anonymize the account identifiers, keep the actual sentence intact. A flattened category tag is exactly what erased the signal in the first place.
If pressed
The self-serve catalog logs every query it gets. That log becomes a leading indicator research management actually watches, the same way a support team watches ticket volume, so a flat or falling query rate gets treated as a signal months before anyone would otherwise notice a whole quarter went quiet.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more