ConceptFoundationalModel Fluency & the AI PM Role / AI PM role variants: platform, applied, infra, research / #1

Describe the difference between an applied AI PM and a platform AI PM in terms of who their customer is.

PICK · a broker's screen and an internal rate model both claim the same customer, at a truckload freight-quoting company

Drossel Freight runs two things on top of the same idea. Lanecast predicts what a truckload should cost to move from one place to another. Boundfare is the screen where a broker types in a lane and sees a bindable price in about three seconds. Torvin Bramlett owns Lanecast. Isaure Callary owns Boundfare. For a year, neither of them had much reason to think hard about who the other one was actually building for.

The direct answer
An applied AI PM's customer is the person on the other end of a single decision the model just made for them: here, the freight broker staring at a bindable price on Boundfare, who needs one trustworthy number in about three seconds. A platform AI PM's customer is every internal team building something of their own on the same rate-prediction model: here, Boundfare's own team, plus a network-planning team, plus whoever calls Lanecast next. Confuse the two and you either build a broker's screen for an engineer's eyes, or a model that only ever serves the first team that asked.
Do this, in order
  1. State the split itself, first: the applied PM's customer is the broker about to bind a load; the platform PM's customer is every internal team building on the same rate model.Why: this is the one line the rest of the answer hangs off. Get it backwards and everything after is arguing about the wrong thing.
  2. Design the broker's screen for the broker, not for the engineers who can already read the model's raw output.Why: a confidence band and a model-version tag mean nothing to someone who just needs to know if the price is safe to click.
  3. Design the model's API for every team that will call it, not only the first one that asked.Why: a model wired to one team's exact shape becomes something a second team has to rebuild from scratch instead of reuse.
  4. Spend your worry on the mistake that hides, not the one that shouts.Why: the broker-screen fix showed up in nine days and cost about $96,000. The narrow-model fix showed up five months later and cost about $310,000.
  5. Count the real number of teams calling the shared model before deciding the roles should stay split.Why: one team for two straight quarters is a sign the platform role isn't earning its keep yet.
  6. Keep the applied PM in the room when the platform PM sets retrain cadence and thresholds.Why: those decisions decide whether the broker's price is still accurate by the time anyone looks at it.

How to answer this, stage by stage

Nobody is grading whether you know two job titles. They're grading whether you can name who each one actually answers to, with real numbers behind it, in a room where someone might ask you to prove it.

1
Scope it to the two products and the two real people
Say it like this
"Let me ground this in something real. Drossel Freight runs two things: Lanecast, the model that predicts a fair rate for a truckload lane, and Boundfare, the screen where a broker actually books one. Torvin owns Lanecast. Isaure owns Boundfare. That's the whole question, really: who is each of them actually building for."
Why this works
Naming two real products and two real owners turns a definitions question into something with an actual, checkable answer.
2
Say the structure out loud
Say it like this
"I'll run this as PICK. Position: my actual answer to who the customer is, up front. Impact: who gets hurt when either PM mixes that up. Cost asymmetry: which mistake is cheap and loud, which one is hidden and expensive. Kill criteria: what would tell me the split at a specific company is wrong."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about two job titles.
3
Give the position, committed
Say it like this
"Here's the actual answer. Isaure's customer is the broker staring at a bindable price on Boundfare's screen, someone who needs one number he can trust in about three seconds, with no idea what's happening underneath. Torvin's customer is every internal team pulling from Lanecast to build something of their own: Boundfare, the network-planning team, and whoever calls it next. Neither of them is building for 'the model.' They're building for two different people who happen to need the same number."
Why this works
This is the literal answer to the question, stated before any story, with no hedge.
4
Prove the applied side with the real failure
Say it like this
"Here's what happens when that gets confused. Early on, Isaure's team shipped Boundfare showing the raw Lanecast output: predicted rate per mile, an 80 percent confidence band, a model-version tag. That's useful to Torvin's engineers. It's noise to a broker booking a load in three seconds. Book-through rate went from 54 percent to 29 percent in nine days. About $96,000 in lost margin, and everyone knew within the week."
Why this works
A real, fast, numbered failure shows what happens the moment an applied PM starts quietly building for engineers instead of the person on the other end of the screen.
5
Prove the platform side with the real failure
Say it like this
"Now the other direction. Torvin built Lanecast tight to exactly what Boundfare needed: one lane at a time, one point estimate, a weekly retrain that matched Boundfare's own release cycle. Five months in, Nuray's network-planning team needed rates across 3,800 lanes at once, for capacity planning. Lanecast couldn't do batch, so instead of waiting on a roadmap slot, they built their own version. Different retrain schedule, slightly different data. Within three months the two models disagreed by 9 percent on the same lane, on average, spiking to 21 percent on the volatile ones. Nobody noticed until a customer's own logistics manager saw a $2,150 quote from Boundfare and a $2,390 number in one of our own capacity reports, same lane, same week."
Why this works
This is the expensive mistake, and it proves the platform PM's error was never about Lanecast's accuracy. It was about who else needed to build on it.
6
Name the cost asymmetry and the mechanism, out loud
Say it like this
"Here's the asymmetry. The broker-screen mistake is loud and cheap: nine days, $96,000, one screen change and it's fixed. The narrow-model mistake is quiet and expensive: five months to surface, about $310,000 once it reached a renewal conversation, plus a shadow model that has to get torn out and reconciled. If I only have attention for one of these, I watch the quiet one. It's silent drift between two copies of what's supposed to be the same number, and nothing about it sends anyone an alert."
Why this works
Naming which failure hides is the actual hard part of PICK. It's what makes this a real decision instead of a preference.
7
Give the kill criteria and close on one line
Say it like this
"Last thing. I'd track one number: how many real internal teams actually call Lanecast in production, not teams that say they might someday. If that number sits at one for two straight quarters, the platform role isn't earning its keep yet, fold Torvin and Isaure back into one AI PM. If it's two or more, with genuinely different needs, the split is real, and the fix is a shared roadmap between them, not fewer people. So: the applied PM's customer is the broker about to bind a load. The platform PM's customer is every team building on the same model. And I'd know the split at Drossel specifically was wrong the day Lanecast still had exactly one caller after six months."
Why this works
A position that names its own falsifying evidence, before anyone asks for it, is what separates a real answer from a job-title opinion.

Let's learn

Lanecast is a guess. Boundfare is a price a person can actually click "book" on.

Hand sketched two panel metaphor scene titled One rate, two different customers. Left panel, a person icon labeled The broker, caption needs one price and needs to trust it fast. A VS mark between the panels. Right panel, a funnel icon labeled Three internal teams, caption need the same number shaped a different way each.
Same number, two different people asking two different questions of it.

Drossel Freight built a tool that predicts what a truckload should cost to move from one place to another, and a screen where a broker can see that price and book the load right there.

Before Boundfare, a broker rang around: four or five calls or emails to different carriers, then waited for someone to call back. About 35 minutes for one lane, and that's if nobody was at lunch. With Boundfare, that same lane comes back priced in about three seconds.

Hand sketched left to right flow diagram titled How one freight quote moves through Drossel. Five connected boxes reading Lane request, Rate prediction, Priced quote, this box emphasized in steel blue, Broker sees it, Books or walks.
Five steps. The middle one, turning a raw prediction into a price a stranger can trust, is the step this whole answer turns on.
Knowledge spark: what's a settled rate? The price a lane actually cleared at once the load moved and got paid for. Lanecast counts as accurate when its prediction lands inside a calibrated band of that number, checked against last week's lanes, most of the time. Not every time. It's a guess, a good one, not a number pulled from a rulebook.

Here's the turn. The model's own accuracy was never the real problem, on either side. Isaure's mistake wasn't that Lanecast was wrong. It was that Boundfare's screen got built for someone who already understood the model, not for the person actually staring at it. Torvin's mistake wasn't that Lanecast was inaccurate either. It was that Lanecast got built for exactly one caller, and a second one showed up anyway.

We didn't build two broken products. We built two products that each knew one customer, and forgot there was a second one waiting.
Hand sketched comparison diagram titled The broker screen Isaure shipped by accident. Left panel, a document icon labeled Before, caption one price, one button, Book this load. Right panel, a red question mark icon labeled After, caption price, an 80 percent band, a model version tag.
Same price underneath. One version a broker can act on in a glance, one version that stops him mid-click to wonder what he's looking at.

At its worst, Isaure's mistake sends a broker straight to a competitor's app mid-quote. At its worst, Torvin's mistake puts two different prices in front of the same customer, in the same week, from the same company.

Hand sketched labeled parts diagram titled What Lanecast v1 was actually built to do. A gauge icon at the center labeled Lanecast v1, with four callouts arranged around it reading one lane at a time, weekly retrain matches Boundfare, one point estimate out, no batch endpoint.
Every one of these four choices was reasonable on its own. Together, they built a model that could only ever answer to one caller.
The choice I would take back When Lanecast first shipped, it was built to serve exactly one caller: Boundfare. Single lane in, one point estimate out, a retrain cadence tuned to Boundfare's own release schedule, no batch endpoint. That made sense when Boundfare was the only thing asking. It stopped making sense the day the network-planning team needed the same model and Lanecast had no way to give it to them.

What I would leave alone: a rough range on Boundfare's screen for a broker who's still browsing lanes, not booking one yet, doesn't need any of this. Without a confidence band or a model tag, a loose number is fine there, because nobody's about to click "book" on it.

The lesson: one model can answer two different questions honestly, and still fail both people asking them, if nobody ever writes down which question belongs to which person.

Now here is the same thing as a story

The short version above is what you actually say in the room. Read this one when you want to feel exactly what a narrow model and a cluttered screen cost, and how long it took anyone to notice either one.

Every Monday morning, before the standup, Torvin Bramlett checks one number: how many lanes Lanecast is least sure about that week. He's owned the rate-prediction engine at Drossel Freight for two years, since before Boundfare had a single broker using it.

Boundfare came out of his model. Isaure Callary built the screen: a broker types in a lane, origin and destination, trailer type, ready date, and in about three seconds sees one price and one button, Book this load. For most of a year, that was the whole relationship between the two of them. Torvin kept the number accurate. Isaure kept the screen simple. Neither had much reason to think about the other's job.

Then, eight months in, growth wanted more trust on the screen. Someone had noticed brokers sometimes hesitated on a quote before booking, and the fix that got greenlit was to show more of what Lanecast already knew: predicted rate per mile, an 80 percent confidence band around it, even a small model-version tag in the corner, so a broker could see the number was fresh.

It shipped on a Tuesday. Isaure watched book-through rate all week the way she always did. By Thursday it had gone from 54 percent to 41. By the following Thursday, 29.

Zsolt Nazarov, an independent broker who had used Boundfare almost daily since it launched, was one of the ones who stopped. He didn't stop because the price was wrong. He stopped because he didn't know what an 80 percent band meant, and a screen that used to answer his question in one glance now made him feel like he needed to ask someone first. So he went back to calling around, the way he used to, and Drossel lost the load.

We didn't make the price less accurate. We made the person reading it feel less sure of it.

Isaure pulled the screen back to one number and one button within the week. Book-through recovered to 51 percent inside ten days, and the whole thing cost about $96,000 in lost margin, mostly absorbed in that first week before anyone even asked about the pattern. Loud, fast, and by the time it came up in a review meeting, it was already fixed.

Torvin's mistake took longer to find, because nobody was watching for it at all.

When Lanecast first shipped, it was built to answer exactly one kind of question, fast: what should this one lane cost, right now, for the broker looking at Boundfare. Single lane in, one point estimate out, retrained every week to match Boundfare's own release cycle. Nobody built a batch mode, because nobody had asked for one.

Hand sketched horizontal timeline titled Five months to the mismatch. Four milestones. Boundfare launches, caption month 0. Team wants batch, caption month 2. They build their own, this milestone emphasized in rust orange, caption month 3. Two prices collide, caption month 5.
Nobody at Drossel decided, on any single day, to let a second version of Lanecast exist. It just got built, the way a workaround does.

Five months after Boundfare launched, Nuray Endale's network-planning team wanted something different: not one lane, fast, but 3,800 lanes, once a week, to plan where Drossel needed more trucks the following month. Lanecast could not do that. It would have taken the model over six hours to answer 3,800 questions built for answering one at a time, and Nuray's team had a Thursday deadline.

So they didn't wait. One of Nuray's engineers, who had worked with Torvin's team before, built a lighter version of the same idea: similar features, a monthly retrain instead of weekly, tuned to run fast across thousands of lanes overnight instead of one lane in real time. It worked. Nobody thought of it as a second model. They thought of it as a spreadsheet with some machine learning in it, built to hit a deadline.

For three months, both numbers existed, quietly disagreeing a little more each month. Average gap of 9 percent by month three. On the lanes that moved a lot, produce season, holiday freight, the gap hit 21 percent.

It surfaced on a Wednesday, in a renewal call Isaure wasn't even on. A customer's logistics manager pulled up a live Boundfare quote for a lane, $2,150, at the same time an account manager was sharing a capacity report built off the network team's model, which priced that same lane at $2,390 for the same week. The customer asked, reasonably, which number Drossel actually believed.

We didn't lose the account over a wrong number. We nearly lost it over whether we had one number at all.

The decision Torvin would take back sits fourteen months earlier, the week Lanecast first shipped. Someone asked, in passing, whether the model should be built to serve more than Boundfare eventually. The honest answer at the time was no, not yet: Boundfare was the only thing calling it, and building batch support nobody had asked for felt like the wrong thing to spend a sprint on. Nobody wrote down that the question should get asked again once a second team actually needed it.

Run the five months again with one change: Lanecast ships with a batch endpoint from day one, even a slow, overnight one, and Torvin's team owns a standing answer for "who else might call this." Nuray's team never builds their own copy, because they don't need to. There's one number, computed twice, updated on two different schedules, but from the same model, the same training run, the same eval set. The renewal call still happens. Both numbers on the screen agree, because they were always the same number.

What I would tell myself, back in that first sprint planning meeting: "nobody's asking for it yet" was true, and it was never a reason to stop asking whether somebody would.

Four checks for telling two customers apart

Not a way to prove either PM was careless. PICK is what forces you to commit to who each role actually serves, then say which kind of "wrong customer" mistake costs you more.

PPosition. The real claim, in one sentence.
Isaure's customer is the broker staring at a bindable price on Boundfare, in the three seconds before he clicks book or closes the tab. Torvin's customer is every internal team pulling from Lanecast to build something of their own: Boundfare today, the network-planning team and whoever calls it next tomorrow. The model itself is not the customer for either of them.
State this before any story, or the rest sounds like two people arguing about whose job is harder.
IImpact. Who feels each kind of wrong.
If Drossel builds Boundfare's screen like it's for engineers, Zsolt and every broker like him feels it inside a week: a stalled quote, a load booked somewhere else. If Drossel builds Lanecast like it only ever needs to serve Boundfare, Nuray's team feels it five months later as duplicated work, and a customer feels it as two different prices for the same lane in the same renewal call.
Naming a real person on each side keeps this from turning into an abstract fight about two job titles.
Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a document icon labeled Wrong screen for a broker, caption loud fast, caught in 9 days, cost stays put. Right panel, a red question mark icon labeled Wrong API for a second team, caption quiet for months, cost keeps compounding.
One mistake announces itself on a dashboard within the week. The other one hides inside two numbers that both look fine on their own.
CCost asymmetry. The heart of it.
The applied mistake is cheap and loud: nine days to notice, about $96,000, fixed by reverting one screen. The platform mistake is hidden and expensive: about five months to surface, roughly $310,000 once it reached a renewal negotiation, plus a shadow model that had to be torn out and reconciled by hand. When I only have attention for one of these, I spend it watching for silent drift between two copies of what's supposed to be the same number, because nothing about it sends an alert on its own.
This is the hardest step in PICK. Naming which mistake hides is what turns a preference into an actual decision.
Cost of the two mistakes, in dollars
$350k $175k $0 $96k Applied mistake loud, fixed in days $310k Platform mistake hidden, found in months
Applied mistake, caught in one weekPlatform mistake, caught in a renewal call
The loud mistake cost less and got fixed inside the same week it was noticed. The quiet one cost more than three times as much and took a customer's own logistics manager to surface it.
KKill criteria. What evidence flips the position.
Track one number: how many real internal teams call Lanecast in production, not teams that say they might. If it sits at one caller for two straight quarters, the platform role isn't earning its keep at Drossel specifically, fold it back into one AI PM who owns the model and the broker screen together. Right now it's two, Boundfare and the network-planning team, with genuinely different needs, batch against single lane, monthly against weekly, so the split is real.
A position with no way to be proven wrong is just an opinion about an org chart. Naming the actual bar is what makes this a real case.
The kill line, charted: real teams calling Lanecast, by month
4 3 2 1 0 kill line: 2 real committed callers needed 1 1 1 1, incident 2 3 now month 0 month 2 month 4 month 5 month 6 month 8
Real production callers of LanecastKill line, 2 needed to justify a split role
For five months the count sat at one, the line nowhere near where it needed to be. The incident at month five is what finally forced the network team onto the real model instead of their own copy of it.

Three things worth naming directly, since this is where the real judgment sits. Torvin's team considered merging the two roles back into one AI PM who would own both Lanecast and Boundfare, and rejected it: a synchronous, three-second, single-lane call for a broker and a batch, overnight, thousands-of-lanes forecast for a planning team pull a roadmap in two directions at once, and one person cannot protect both without one of them quietly losing every sprint. The AI-specific failure worth naming by name is silent drift: two independently retrained copies of what is supposed to be the same prediction, disagreeing a little more each month, with nothing that alerts anyone because neither model is technically wrong, they are just no longer the same guess. The guardrail is a weekly reconciliation job, not a smarter model: both versions get scored against the same settled-rate eval set every week, and any lane where they disagree by more than 5 percent gets held from any customer-facing number until someone checks it. And the trade-off is real, and accepted on purpose: the batch endpoint Torvin eventually built prices 3,800 lanes overnight for a fraction of the cost of 3,800 live calls, but every number in it can be up to eighteen hours old, which is fine for a Thursday planning meeting and would be the wrong choice entirely for a broker about to click book.

And if you want to be sure it really works, try it somewhere else

Same four letters, an insurance claims desk instead of a dispatch floor, and this time the shared number isn't a freight rate. It's a fraud-risk score.

Wyndham Mutual built Brackendale, a model that scores every incoming auto claim for fraud risk, and Beaufort, the screen an adjuster uses to actually work a claim. Colwyn Bertoldi owns Brackendale. Yevheniy Ottoway owns Beaufort. Two other teams pull from Brackendale too: subrogation, which chases claims tied to known fraud rings, and underwriting, which uses the same risk signal in aggregate to reprice next year's policies in a region.

Hand sketched decision tree titled One score, four different next steps. Root box reads Brackendale scores a claim, branching into four outcomes. Routine, low risk leads to Beaufort auto closes it. Flags fraud risk leads to Adjuster reviews in Beaufort. Matches a known ring leads to Subrogation investigates. Portfolio level signal leads to Underwriting re-prices next term.
One score, walking out four different doors, to four different customers of the same model.

The same confusion showed up here, in a different shape. Beaufort briefly showed adjusters Brackendale's raw risk score, a number from 0 to 100, instead of a plain reason. Adjusters started treating any claim under 20 as automatically safe and anything over 80 as automatically suspicious, skipping their own judgment on the ones in between, because a number felt more official than it was. Yevheniy's team pulled the raw score and replaced it with three lines a person actually reads: what looked odd, what looked normal, and what to check first.

The decision Colwyn would take back Brackendale launched serving one team, the claims desk, with a single risk score and no separate cut for portfolio-level patterns. That was fine while Beaufort was the only caller. It stopped being fine the day underwriting wanted the same signal aggregated by region, and subrogation wanted it filtered to only ring-pattern claims, two different shapes of the same score, from a model that only ever output one.

Mapped straight onto PICK: the position holds the same shape. Yevheniy's customer is the adjuster deciding whether to pay a claim today. Colwyn's customer is every team, claims, subrogation, underwriting, pulling a different cut of the same risk signal. The impact splits the same way too: a raw score on an adjuster's screen gets misread inside days; a model tuned only for one team's shape of the question gets duplicated in the shadows and drifts, the same way Lanecast did. The cost asymmetry lands the same way: the adjuster-screen fix was loud and cheap, the narrow-model fix was quiet and took a full underwriting cycle to notice. And the kill criteria transfer directly: track how many real teams call Brackendale in production, and fold the roles back together the day that number sits at one for two straight quarters.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: the applied PM answers to the person reading the screen, the platform PM answers to every team calling the model, and the two roles fight if nobody names which mistake to watch for.
Cost: no budget to build the batch endpoint or the region-level cut this quarter. Ship the narrow version anyway, but write down, on day one, which second team is most likely to need it, so the rebuild-versus-reuse decision isn't a surprise later.
The model got better, for real: say Lanecast's accuracy climbs another five points next quarter. It doesn't touch this problem at all. A more accurate single-lane model is still the wrong shape for a team that needs 3,800 lanes at once.

Where people run it wrong.
They answer with "the platform PM is more technical," instead of naming who each one actually serves.
They let the first team that asks for a model define its shape forever, instead of asking who else might need it.
They fix a trust problem by hiding the model's own number entirely, instead of translating it into something the actual customer can use.

How to use it live. Before answering, ask yourself one question: does this person ever see the model directly, or only the product built on top of it? Whoever only sees the product built on top of it is the applied PM's customer. Everyone else, calling the same model for their own reasons, is the platform PM's customer.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "describe the difference, in terms of who's the customer" question, and what's its job?
Tap to flip
ANSWER
PICK: commit to a position on who the customer is, name who's hurt by each kind of wrong, find the asymmetry between the cheap mistake and the hidden one, then say what evidence would flip you.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Isaure Callary, who owns Boundfare, and Torvin Bramlett, who owns Lanecast, at Drossel Freight. Also Zsolt Nazarov, the broker who stopped trusting the screen, and Nuray Endale, who leads network planning.
3 · THE POSITION
What's the actual difference between the two customers?
Tap to flip
ANSWER
The applied PM's customer is the person acting on a single decision the model just made for them, in about three seconds. The platform PM's customer is every internal team building something of their own on the same model.
4 · THE APPLIED MISTAKE
What happens when an applied PM treats internal engineers as the real audience?
Tap to flip
ANSWER
Boundfare showed a broker the raw model fields, a predicted rate, an 80 percent confidence band, a model-version tag. Book-through rate dropped from 54 to 29 percent in nine days, about $96,000 in lost margin.
5 · THE PLATFORM MISTAKE
What happens when a platform PM builds only for the first team that asked?
Tap to flip
ANSWER
Lanecast was built single-lane and weekly, only for Boundfare. A second team built a shadow copy instead of waiting, and the two drifted 9 percent apart on average, 21 percent on volatile lanes, until a customer saw two different prices for the same lane.
6 · THE COST ASYMMETRY
Which mistake is cheap and loud, and which one hides?
Tap to flip
ANSWER
The applied mistake is cheap and loud: nine days, about $96,000, fixed with one screen change. The platform mistake is hidden and expensive: about five months to surface, roughly $310,000, once it reached a customer's renewal call.
7 · THE KILL CRITERIA
Fill in the blank: if Lanecast still has just ___ real internal caller after ___ straight quarters, fold the two roles back into one.
Tap to flip
ANSWER
One caller, two quarters. At Drossel, it's already at two real callers, Boundfare and the network-planning team, so the split holds.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent split?
Tap to flip
ANSWER
Wyndham Mutual's claims tools. Colwyn Bertoldi owns Brackendale, the fraud-risk model, called by claims, subrogation, and underwriting. Yevheniy Ottoway owns Beaufort, the adjuster's screen, the applied side an adjuster actually sees.

Check yourself Score: 0 / 0

Multiple choice
1. Why doesn't showing Boundfare's raw model fields (predicted rate, confidence band, model-version tag) actually serve the broker better, even though it's more information?
  • A. Because brokers cannot read numbers quickly under time pressure.
  • B. Because a raw model field answers a question an engineer debugging Lanecast would ask, not the question a broker booking a load actually has.
  • C. Because Lanecast's confidence bands were usually wrong.
  • D. Because Boundfare's screen did not have room to show extra text.
Show hint
Look at stage 4 of the walkthrough, where the applied-side failure gets proven.
Show answer
B. The information was accurate. It just answered the wrong person's question. The broker needed a yes-or-no on price, not a peek at the model's own uncertainty.
True or false
2. True or false: the network-planning team building their own copy of Lanecast was a bigger mistake than Torvin building Lanecast without a batch endpoint in the first place.
  • True
  • False
Show hint
Check "The choice I would take back," in Let's learn.
Show answer
False. Building Lanecast narrow at launch made sense when Boundfare was the only caller. The real reversal is never revisiting that choice once a second real team needed it, not the network team's workaround, which was the sensible move given no other option was offered.
Fill in the blank
3. Book-through rate on Boundfare dropped from ___ percent to ___ percent in ___ days after the raw model fields shipped.
Show hint
Look in the story section, right before the highlight about the price feeling less sure.
Show answer
54 percent, 29 percent, 9 days. The number moved fast because the mistake was loud. That's exactly why it was the cheaper one to fix.
Short answer, name the rejected alternative
4. What alternative did Drossel consider instead of keeping Torvin and Isaure as two separate AI PMs, and why did it lose?
Show hint
Look for the "three things worth naming directly" paragraph near the end of the PICK recap.
Show answer
Model answer: Merging both roles into one AI PM owning Lanecast and Boundfare together. It lost because a synchronous, three-second, single-lane call and a batch, overnight, thousands-of-lanes forecast pull one roadmap in two directions, and one person cannot protect both without one of them starving.
Short answer, apply it yourself
5. Think of an AI product you use that has a customer-facing side and an internal model or API behind it. Name one thing the two sides would need from that same prediction that pulls in different directions.
Show hint
Look for a case where the fast, simple version a user sees and the detailed version an internal team needs cannot both come from the exact same shape of output.
Show answer
Model answer: A spam filter's inbox view just needs a fast yes-or-no, because a wrongly-hidden email feels much worse than a slow one. An internal trust-and-safety dashboard built on that same spam model needs a graded risk score and reasons, not a yes-or-no, because someone there is deciding whether to change the whole model's threshold.
Short answer, work the number
6. If the network-planning team's shadow model had drifted only 2 percent by month three instead of 9 percent, would it still have been worth catching before the renewal call?
Show hint
Check what the drift did between month three and the incident at month five in the kill-line chart's note.
Show answer
Probably yes. At $2,150 for an average lane, even a 2 percent gap is real money, and it's the compounding pattern that matters, not the size of one gap. This drift widened from 9 to 21 percent on volatile lanes rather than settling on its own.
Before you close the answer
Why this works
Tests whether you can name a real, checkable difference grounded in what each customer actually does with the model's output, not just repeat two job titles from a job posting. Most candidates say "platform works on infra, applied works on features" and never say why that distinction matters when something breaks.
Follow-up traps
"Isn't calling an internal team a 'customer' a stretch? Isn't the platform PM's real customer just the model itself?" Response: no. A customer is whoever has to accept or reject the model's output for their own use, and Nuray's team decides what "done" means for Lanecast the same way Zsolt decides what "done" means for Boundfare.

"Why not just give every internal team direct access to the raw model and skip having a platform PM at all?" Response: that is exactly what happened by accident when the network team built their own copy, and it is the failure this whole answer exists to name. Unmanaged shared access does not remove the need for someone to own consistency, it just hides who is supposed to.
If pressed
The reconciliation job doesn't just diff two numbers. It re-scores both models against the same held-out set of last week's settled rates and only flags a lane if both models sit outside their own calibrated band on that lane, so ordinary week-to-week noise on one lane doesn't trigger a false alarm every Monday.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more