CaseAdvancedModel Fluency & the AI PM Role / AI PM vs traditional PM vs technical PM / #11

An engineering manager says the AI PM is doing their job. How do you draw the boundary?

PICK · an AI cost-estimating and scheduling assistant for construction bids

Formwork is Underpin Systems' tool for construction bids: it reads a subcontractor's quote PDF, matches each cost line against historical prices, and either fills in the number on its own or flags it for a person. Sindri Loewen is the AI PM. Ferghal Achterhoff is the engineering manager. Neither of them had ever written down, in dollars, how wrong Formwork was allowed to be before a person had to look.

The direct answer
I own the threshold that decides when Formwork's cost estimate ships without a person looking at it, and what happens to anything below that line. The engineering manager owns how the model gets to that number: which model, how it's trained, what infrastructure serves it. I draw the line by asking which kind of mistake is cheap and visible and which is hidden and expensive, and I keep the hidden one, because a wrong threshold never announces itself. It just quietly costs a client money until someone finds the invoice.
Do this, in order
  1. State your position first: you own the threshold and what happens below it, the EM owns how the model gets there.Why: without a stated line, every disagreement becomes a fresh turf fight instead of a decision anyone can point back to.
  2. Name who feels each kind of error, in real terms, not in feelings.Why: "I feel micromanaged" and "a client lost margin" are not the same size of problem, and saying so out loud is what makes the boundary defensible.
  3. Optimise against the mistake that stays hidden, not the one that gets argued about in the room.Why: a loud pushback in a meeting corrects itself the same day; a silent bad default doesn't correct itself at all, someone has to go looking.
  4. Put the threshold in writing, owned by product, built however engineering wants.Why: a number nobody explicitly owns defaults to whoever tuned the eval curve, and an eval curve was never asked to price a client's risk.
  5. Set kill criteria for the boundary itself, not just for today's argument.Why: without a number to check against, the next disagreement starts from zero instead of from agreed evidence.
  6. Leave the model choice, the retrain schedule, and the infrastructure alone.Why: none of that touches the dollar risk the boundary exists to protect, so contesting it spends trust for nothing.

How to answer this, stage by stage

Nobody is grading whether you can say "communication is important." They're grading whether you can turn a fight about job titles into a rule that survives the next ten fights.

1
Scope it to one disputed call
Say it like this
"Let's make this concrete. Formwork reads a subcontractor's bid PDF and either fills in a cost line on its own or flags it for a person. The disputed call is: who decides where 'automatic' stops and 'flagged' starts."
Why this works
Gives the interviewer one real decision to hold onto instead of an abstract question about org charts.
2
Say the structure out loud
Say it like this
"I'll run this as PICK. Position, then who feels each kind of error, then the cost asymmetry, then the evidence that would change my mind."
Why this works
Two seconds of structure tells the interviewer you have a method, not an opinion forming as you go.
3
Give your position first, no hedging
Say it like this
"My position: I own the auto-approve threshold and what happens below it. Ferghal owns how the model gets to that number, which model runs, how often it retrains, what serves it."
Why this works
Answers the literal question in one breath. Everything after this is defense, not discovery.
4
Name who feels each kind of error
Say it like this
"If I start telling Ferghal which embedding model to use, he feels micromanaged, and he's right to push back, out loud, in that meeting. If I back off setting the threshold because it 'sounds technical,' nobody feels anything, until a client's bid is short forty-one thousand dollars and nobody signed off on the number that let it through."
Why this works
Turns an abstract fight into two named, comparable outcomes an interviewer can actually weigh against each other.
5
Name the cost asymmetry, and prove it with the failure
Say it like this
"One of these mistakes is loud, cheap, and self-correcting. The other is quiet and expensive. Formwork matched a rebar line at eighty-nine percent confidence, above our eighty-five percent line, so it auto-approved. The price behind that match was three months stale, steel had moved eighteen percent in that window. Nobody caught it until close-out, six weeks later, forty-one thousand dollars into a sixty-thousand-dollar margin. That's the mistake I optimize against."
Why this works
A real number with a real timeline is what separates a decision from a preference.
6
Give the kill criteria, and close on the line
Say it like this
"I'd watch two things. If I'm reversing Ferghal's implementation calls more than about once a quarter with no new product reason behind it, I've drifted into his job, and the line moves toward him. If a threshold I wrote down keeps getting quietly changed during a retrain, he's drifted into mine, and we need a firmer process, not another conversation. So: I own the threshold and the dollar risk under it, he owns how we get there, and I watch which mistake is actually happening to know if that's still true."
Why this works
Leaves the interviewer with a rule that outlives this one disagreement, which is what "how do you draw the boundary" is actually asking.

Let's learn

Formwork lives inside one screen: a bid, line items down the left, and next to each one either a quiet green check or a flag. The check means Formwork matched the line against its price history with enough confidence to fill it in on its own. The flag means a person looks first.

Before Formwork, a Slateford Construction estimator re-keyed every subcontractor quote by hand into the master estimate, cross-checking each price against a binder of past jobs. For a normal mid-size commercial bid, a few hundred line items, that took a senior estimator about three hours, usually the night before the bid was due.

Hand sketched flow diagram titled Before Formwork: one bid, by hand. Four steps in sequence connected by arrows. Quotes arrive as PDFs. Estimator re-keys every line. Checks the price book by hand. Bid ready, 3 hours later, this final step outlined in amber.
This is the three hours Formwork was built to give back. It mostly did.

With Formwork, that estimator now spends about 25 minutes on a bid: reviewing whatever got flagged, and glancing at the rest. The auto-approve cutoff sits at 85 percent confidence, a number Underpin's engineering team picked because that's roughly where their internal test set balanced catching real errors against flagging too much.

Knowledge spark: what's an auto-approve threshold? The cut-off point where a model's own guess is trusted enough to act without a person checking first. Above the line, Formwork fills the number in. Below it, someone has to look before the bid goes out.

Here's the turn. The argument between Sindri and Ferghal was never really about who was "doing whose job." It sounded that way, because Sindri kept weighing in on things that looked technical, and Ferghal kept feeling like his calls were being taken from him. But underneath it, the real question was narrower: which kind of mistake are we going to let hide.

One of these mistakes gets argued about out loud, the same day. The other one gets found six weeks later, in an invoice.
Hand sketched comparison diagram titled Two ways this boundary gets it wrong. Left panel, a person icon in green, labeled PM oversteps into architecture, caption EM pushes back in the meeting, fixed the same day. Right panel, a scale icon in red orange, labeled PM backs off the threshold call, caption bad default ships quiet, found six weeks later.
Both are real mistakes. Only one of them costs a client money before anyone notices.

At its worst, this cost Slateford real margin. A six-story mixed-use bid carried a line for grade 60 rebar, 380 tons of it. Formwork matched that line to a similar job in its history at 89 percent confidence, comfortably above the 85 percent cutoff, so it filled the number in and moved on. What Formwork didn't carry was that steel prices had climbed about 18 percent in the months since that historical price was last refreshed. The match was confident. The price behind the match was stale.

Customer dollars at risk, by mistake type
$45k $22k $0 $0 PM oversteps into architecture $41,000 EM sets the threshold alone
Caught before it reached a customerReached a real bid, found six weeks later
The left bar isn't zero because that mistake doesn't matter. It's zero because it never left the building. The right one did.
The choice I would take back Underpin let engineering set the 85 percent threshold purely from where their eval curve balanced precision and recall. Nobody in product ever translated that number into what it was allowed to cost a client. That made sense when Formwork mostly auto-filled small hardware lines. It stopped making sense the moment it started auto-filling six-figure material orders on real bids.
Hand sketched labeled parts diagram titled Where the line actually sits. A balance scale icon at the center labeled The Boundary, with four callouts around it. PM: the auto-approve threshold. PM: what happens below it. EM: how confidence is computed. EM: retrain cadence, infra.
Two callouts are a dollar-risk question. Two are an implementation question. That split is the whole boundary.

What I would leave alone: which model Formwork uses to match a line item, how often it retrains, what infrastructure serves it. None of that changes the dollar risk the threshold protects, and Sindri had already tried to weigh in on one of those, the retrain schedule, and lost that argument fairly. It stayed lost.

The lesson: the boundary was never really a map of two job titles. It's a line drawn by which mistake is allowed to stay invisible until it's expensive, and that line has to be checked, not just declared once and left alone.

Boundary overrides per month, before and after the line got written down
8/mo 4/mo 0 Kill line: 3/month Boundary written down M1 M3 M5 M7
Before the boundary was written downAfter
This is the number I'd actually watch afterward. It's the leading signal for whether the line needs to move again.

Now here is the same thing as a story

Read this one when you want to feel why 89 percent seemed safe enough to auto-approve, and why nobody thought to ask what it was allowed to cost.

Sindri Loewen had been the one who caught, a year before any of this, that Formwork's first version would have silently rounded every unit price to the nearest dollar, fine for a $40 fixture, quietly wrong for a $2,400-a-ton material order. She found it reading a spec doc, before a line of code existed. Ferghal remembered that catch. It's part of why, for the first eight months after launch, the two of them worked well together without anyone drawing a line at all.

Those were good months. Sindri would bring Ferghal a product problem, "estimators don't trust a flag with no reason attached," and he'd come back with an interface for it inside a sprint. Ferghal would flag a modeling tradeoff, "faster matching costs us some accuracy," and Sindri would tell him which side of that tradeoff Slateford actually needed. Neither of them thought about whose job anything was. It just worked.

It thinned in three small moves, none of them a fight. First, Sindri stopped asking why the confidence cutoff sat at 85 percent, because Ferghal had a chart showing that's where the model's own numbers looked best, and a chart felt like an answer. Second, when a new hire asked in a standup who actually owned that number, Sindri said "engineering, it's a model thing," and nobody corrected her. Third, when Ferghal quietly moved the cutoff from 83 to 85 during a retrain because it tested better, nobody from product was in the room, and nobody asked to be.

Then, in a planning meeting, Sindri wrote into a spec that Formwork should use "a fine-tuned classifier retrained every Friday night." Ferghal pushed back hard, right there: "That's not your call. That's architecture." He was right, and Sindri knew it before he finished the sentence. She rewrote the spec that afternoon to state the outcome she actually needed, accurate matches, fresh prices, and left the how to him. It cost about 45 minutes of friction. Everyone in the room saw it happen, and it was over by lunch.

Hand sketched flow diagram titled How $41,000 slipped through quiet. Five steps connected by arrows: Rebar line in the bid PDF, Formwork matches historical price, 89 percent confidence above the 85 percent line, this step outlined in red orange, Auto approved no flag shown, Found 6 weeks later at close out.
Every single step here looked reasonable on its own. That's what made it hard to catch.

Three weeks after that, the Slateford bid went out with the rebar line auto-approved at 89 percent confidence. Nobody argued about that one, because nobody saw it happen. It just cleared the line, the way hundreds of lines had cleared it before, correctly, every time the underlying price data was current.

We didn't lose a rebar order. We lost the ability to say, in writing, whose job it was to have caught it.

Sindri thought back to the kickoff, over a year earlier, where 85 percent got chosen. Someone had asked, mildly, whether product should sign off on the cutoff. The answer at the time was reasonable: the line items were small then, a mismatch cost at most a few hundred dollars, and slowing down a fast-moving build for a sign-off process felt like the wrong tradeoff. Nobody ever circled back, because nothing forced the question until Formwork was auto-filling material orders worth tens of thousands of dollars.

Hand sketched comparison diagram titled One shared call, drawn as two people. Left figure in green labeled THE PM, caption sets the bar and what happens below it. Right figure in warm grey labeled THE EM, caption builds whatever clears that bar.
Not two job descriptions. One decision, split cleanly down the middle.

What changed after the rebar line wasn't a new rule about who could talk to whom. Sindri wrote the threshold down as a product artifact: a dollar-risk cap per material category, and a staleness check that has to pass alongside the confidence score before anything auto-approves. Ferghal built it however he wanted, and moved the actual number twice since, both times with a written note back to Sindri explaining why. Overrides, the number of times either of them second-guessed the other's call outright, dropped from an average of five or six a month to under one. The replay isn't a better argument. It's a Tuesday, six months later, where nobody argues about it at all.

What I'd tell my past self, at that first kickoff: skipping the sign-off wasn't careless. It was the sensible call when the numbers were small. Nobody ever agreed to revisit it once the numbers stopped being small, and by then the line had been quietly deciding itself for months.

PICK, or pricing the boundary instead of guessing at it

Hand sketched quadrant diagram titled Sounds technical, is it. Y axis, how technical it sounds, from sounds like product to sounds like engineering. X axis, what it actually decides, from how it's built to what ships to a customer. Which embeddings model and Retrain cadence plotted top left, sounds technical and is technical. Auto approve threshold and What happens below the line plotted top right, sounds technical but is actually a risk call.
The whole answer in one picture. Two of these are engineering's call. Two only sound that way.
PPosition. Say the line before any reasoning.
I own Formwork's auto-approve threshold and what happens below it. Ferghal owns how the model gets to that number, which model, how it retrains, what runs it.
Everything else in this answer defends this one sentence. It doesn't build up to it.
IImpact. Name who feels each kind of error.
Ferghal feels micromanaged when I weigh in on the model. Slateford feels a silent margin hit when the threshold is set without pricing what it's allowed to cost.
Naming both sides, not just the one that's easier to defend, is what keeps this from sounding like a land grab.
CCost asymmetry. One error is cheap and visible. One is hidden and expensive.
An overstep into architecture gets caught in the room, same day, cost about 45 minutes. A threshold nobody priced cost Slateford $41,000, found six weeks later.
This is the actual answer to "how do you draw the boundary." Draw it around whichever mistake doesn't announce itself.
KKill criteria. What would tell you the line itself needs to move.
If I'm reversing Ferghal's technical calls more than once a quarter with no new product reason, the line has drifted toward me and should move back. If a threshold I own keeps getting quietly changed during a retrain, it's drifted toward him.
Without this, every future disagreement restarts from zero instead of checking against agreed evidence.

The recap, one line per letter: say the line first, name who feels each kind of mistake, optimize against the one that hides, and watch a number that would tell you the line itself is wrong, not just this one argument.

One alternative we actually considered and rejected: writing a fixed decision-rights document up front, a RACI-style list covering every possible call between product and engineering. It lost, because Formwork's decisions change shape every few months as the model and the bid data change, and a fixed list goes stale and stops getting opened. A live cost-asymmetry test travels better than a document nobody reads. The AI-specific failure worth naming directly is confident-but-stale matching: a text match on a line item can score high on wording similarity while the price behind it has quietly drifted, because the drift lives in the world's prices, not in the words being matched. The guardrail is a second signal, a staleness check per material category, required to pass alongside the confidence score before anything auto-approves. And the trade-off is real and accepted on purpose: a lower, more conservative dollar-risk cap catches more of these stale-price misses, but it means more lines get flagged, which slows down a bid that's often due the next morning. Underpin accepts a few more minutes of review time as the price of not silently eating a client's margin.

And if you want to be sure it really works, try it somewhere else

Same four letters, an insurance claims tool instead of a construction bid, and the disputed number is a dollar cap on a photo, not a price on a ton of steel.

Wexmoor is Cravensworth Mutual's tool for property claims: it reads a submitted claim and either auto-approves a payout or flags it for a human adjuster. Ombretta Castelnuovo is the AI PM. Torbjorn Manx is the engineering manager. Their version of the fight was the same shape as Sindri and Ferghal's, over a different number: auto-approve any claim under $2,000 with damage-assessment confidence above 90 percent.

Mapped onto PICK: Ombretta's position is that she owns the dollar cap and the confidence bar for auto-approval, and what happens to anything under it; Torbjorn owns the computer-vision model reading the damage photos, its training data, and the infrastructure serving it. The impact: Torbjorn feels overruled if Ombretta specifies which model to fine-tune; Cravensworth's loss ratio feels it, quietly, if the dollar cap is set purely off model accuracy instead of actual fraud and misestimate exposure. The cost asymmetry is identical in shape: a pushback in planning is loud and fixed same day; a miscalibrated cap is invisible until a quarter's loss numbers come in high, or a pattern of claims that needed a human eye got waved through instead. The kill criteria transfer directly: too many product reversals of technical calls with no new evidence, move the line toward engineering; a written product threshold getting silently redefined during a model update, tighten the process instead of re-arguing it.

Swap the trigger and it still runs.
Speed: an interviewer caps you at 90 seconds. Skip straight to it: state your position, then name the one cheap-versus-expensive mistake pair, done.
Cost: no time to write a formal threshold document this sprint. Add one required field to the ticket system you already use, who owns this number, that's nearly free and gets most of the benefit.
The model got better, for real: say Formwork's matching accuracy climbs to 99 percent. Keep the dollar-risk cap process anyway, because "how much are we willing to auto-approve without a look" doesn't stop being a risk question just because the model got better at the wording part.

Where people run it wrong.
They write a permanent decision-rights document instead of a live, checkable cost test.
They let whoever sounds more technical win the argument, instead of asking what the number actually decides.
They treat "the threshold" as settled forever once it's written down, instead of watching the override rate for a sign it needs to move again.

How to use it live. When an interviewer says "so whose call is this, really," ask one thing back before answering: "which mistake here is expensive if it ships quiet, and which one just gets argued about?" That's usually the exact test being run.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "who owns this call" tradeoff question like this one?
Tap to flip
ANSWER
PICK: state your Position first, name the Impact on both sides, find the Cost asymmetry, and give Kill criteria for when the boundary itself should move.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Sindri Loewen, the AI PM on Formwork, and Ferghal Achterhoff, the engineering manager, both at Underpin Systems.
3 · THE POSITION
What's the one-sentence position?
Tap to flip
ANSWER
The PM owns the auto-approve threshold and what happens below it. The EM owns how the model reaches that number: which model, retrain cadence, infrastructure.
4 · THE ASYMMETRY
What's the cost asymmetry in this story?
Tap to flip
ANSWER
PM oversteps into architecture: caught in the meeting, fixed the same day. PM backs off the threshold call: ships quiet, found six weeks later, $41,000 gone from a client's margin.
5 · THE OLD DECISION
What decision would this answer take back?
Tap to flip
ANSWER
Letting engineering set the 85 percent auto-approve threshold purely from where their eval curve balanced precision and recall, with no product sign-off on what that number was allowed to cost a client.
6 · THE NUMBER
Fill in the blank: the rebar line matched at ___ percent confidence, above the ___ percent cutoff, and cost Slateford $___ on a bid with $___ of margin built in.
Tap to flip
ANSWER
89 percent, 85 percent, $41,000, $60,000.
7 · THE KILL CRITERIA
What evidence would say the boundary itself needs to move, not just settle this one argument?
Tap to flip
ANSWER
The PM reversing the EM's implementation calls more than about once a quarter with no new product reason (move the line toward the EM), or a PM-owned threshold getting quietly changed during a retrain with no product sign-off (tighten the process, don't just re-argue it).
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent threshold there?
Tap to flip
ANSWER
Wexmoor, Cravensworth Mutual's property-claims tool. The equivalent threshold is a dollar cap and confidence bar for auto-approving a claim payout without an adjuster's review.

Check yourself Score: 0 / 0

True or false
1. True or false: what actually cost Slateford money was that Formwork's matching model was too inaccurate.
  • True
  • False
Show hint
Look at what was actually stale: the match, or the price behind it.
Show answer
False. The match was confident and correct on wording. The price data behind that match was three months stale. Nobody had a rule requiring a freshness check alongside the confidence score.
Multiple choice
2. Which of these is the actual cost asymmetry the C step finds in this story?
  • A. Both mistakes cost the company about the same, so it comes down to whoever argues harder.
  • B. PM overstepping into architecture is cheap and visible; PM backing off the threshold call is hidden and expensive.
  • C. Ferghal's calls are always cheaper to get wrong than Sindri's.
  • D. The only real cost that matters is the model's raw error rate.
Show hint
Compare how each mistake was actually discovered, and how long that took.
Show answer
B. One mistake self-corrects in the room the same day. The other doesn't announce itself at all, and cost real money before anyone found it.
Fill in the blank
3. The rebar line auto-approved at ___ percent confidence, above the ___ percent line, three months after the price it relied on had moved ___ percent.
Show hint
Check the paragraph describing the six-story mixed-use bid, in "Let's learn."
Show answer
89 percent, 85 percent, 18 percent. A confident match on wording said nothing about whether the price underneath it was still current.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Letting engineering set the 85 percent threshold purely off its own eval curve, with no product sign-off on the dollar risk. It made sense early, when Formwork mostly auto-filled small hardware lines and a mismatch cost a few hundred dollars at most.
Short answer, apply it yourself
5. Think of a product you use where a person and a model both have a hand in a decision. Name one setting that sounds like an engineering knob but is actually a risk call nobody in product ever explicitly owns.
Show hint
Look for a cut-off or a cap the team probably picked by watching a model metric, not by pricing what a wrong call costs.
Show answer
Model answer: A photo backup app's setting for auto-deleting photos it scores as "blurry." It sounds like a computer-vision confidence threshold, but it's really a risk call about how much irreplaceable footage someone is willing to lose without ever seeing it first.
Short answer, work the kill criteria
6. Say Underpin left the threshold unowned for another year and overrides climbed back to 5 a month. Per the kill criteria in this answer, does the boundary need to move, and toward whom?
Show hint
Look at the line chart's kill line, and what it's actually a signal for.
Show answer
Model answer: Five a month is above the 3-a-month kill line, so the boundary itself needs a look, not just whoever's disagreement is loudest that week. Whether it moves toward the EM or toward a firmer process depends on which direction the overrides are running: new legitimate product evidence points toward the EM, engineering quietly redeciding the threshold points toward tightening the process.
Before you close the answer
Why this works
Tests whether you can turn an org-chart argument into an actual decision rule, instead of either capitulating (the EM decides anything that sounds technical) or empire-building (the PM decides everything). Most candidates reach for generic collaboration advice; the strong answer prices the boundary in dollars.
Follow-up traps
"Isn't the threshold value literally a hyperparameter? That's an ML decision." Response: only until someone translates it into what it's allowed to cost a customer. Once it's priced, it's a risk decision with an implementation underneath it, and risk is product's call to own.

"What if the EM just refuses to hand that number over?" Response: point at the $41,000 case. The ask was never "give me the number," it was "let's agree, in writing, on a number we're both willing to defend to Slateford," which makes it shared accountability, not a power grab.
If pressed
Formwork now carries a second signal alongside text-match confidence: a staleness count, days since that material category's price was last refreshed. Auto-approval requires match confidence above 85 percent AND staleness under 60 days. Either one failing routes the line to a person, no matter how confident the wording match looks.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more