An engineering manager says the AI PM is doing their job. How do you draw the boundary?
Formwork is Underpin Systems' tool for construction bids: it reads a subcontractor's quote PDF, matches each cost line against historical prices, and either fills in the number on its own or flags it for a person. Sindri Loewen is the AI PM. Ferghal Achterhoff is the engineering manager. Neither of them had ever written down, in dollars, how wrong Formwork was allowed to be before a person had to look.
- State your position first: you own the threshold and what happens below it, the EM owns how the model gets there.Why: without a stated line, every disagreement becomes a fresh turf fight instead of a decision anyone can point back to.
- Name who feels each kind of error, in real terms, not in feelings.Why: "I feel micromanaged" and "a client lost margin" are not the same size of problem, and saying so out loud is what makes the boundary defensible.
- Optimise against the mistake that stays hidden, not the one that gets argued about in the room.Why: a loud pushback in a meeting corrects itself the same day; a silent bad default doesn't correct itself at all, someone has to go looking.
- Put the threshold in writing, owned by product, built however engineering wants.Why: a number nobody explicitly owns defaults to whoever tuned the eval curve, and an eval curve was never asked to price a client's risk.
- Set kill criteria for the boundary itself, not just for today's argument.Why: without a number to check against, the next disagreement starts from zero instead of from agreed evidence.
- Leave the model choice, the retrain schedule, and the infrastructure alone.Why: none of that touches the dollar risk the boundary exists to protect, so contesting it spends trust for nothing.
How to answer this, stage by stage
Nobody is grading whether you can say "communication is important." They're grading whether you can turn a fight about job titles into a rule that survives the next ten fights.
Let's learn
Formwork lives inside one screen: a bid, line items down the left, and next to each one either a quiet green check or a flag. The check means Formwork matched the line against its price history with enough confidence to fill it in on its own. The flag means a person looks first.
Before Formwork, a Slateford Construction estimator re-keyed every subcontractor quote by hand into the master estimate, cross-checking each price against a binder of past jobs. For a normal mid-size commercial bid, a few hundred line items, that took a senior estimator about three hours, usually the night before the bid was due.
With Formwork, that estimator now spends about 25 minutes on a bid: reviewing whatever got flagged, and glancing at the rest. The auto-approve cutoff sits at 85 percent confidence, a number Underpin's engineering team picked because that's roughly where their internal test set balanced catching real errors against flagging too much.
Here's the turn. The argument between Sindri and Ferghal was never really about who was "doing whose job." It sounded that way, because Sindri kept weighing in on things that looked technical, and Ferghal kept feeling like his calls were being taken from him. But underneath it, the real question was narrower: which kind of mistake are we going to let hide.
At its worst, this cost Slateford real margin. A six-story mixed-use bid carried a line for grade 60 rebar, 380 tons of it. Formwork matched that line to a similar job in its history at 89 percent confidence, comfortably above the 85 percent cutoff, so it filled the number in and moved on. What Formwork didn't carry was that steel prices had climbed about 18 percent in the months since that historical price was last refreshed. The match was confident. The price behind the match was stale.
What I would leave alone: which model Formwork uses to match a line item, how often it retrains, what infrastructure serves it. None of that changes the dollar risk the threshold protects, and Sindri had already tried to weigh in on one of those, the retrain schedule, and lost that argument fairly. It stayed lost.
The lesson: the boundary was never really a map of two job titles. It's a line drawn by which mistake is allowed to stay invisible until it's expensive, and that line has to be checked, not just declared once and left alone.
Now here is the same thing as a story
Read this one when you want to feel why 89 percent seemed safe enough to auto-approve, and why nobody thought to ask what it was allowed to cost.
Sindri Loewen had been the one who caught, a year before any of this, that Formwork's first version would have silently rounded every unit price to the nearest dollar, fine for a $40 fixture, quietly wrong for a $2,400-a-ton material order. She found it reading a spec doc, before a line of code existed. Ferghal remembered that catch. It's part of why, for the first eight months after launch, the two of them worked well together without anyone drawing a line at all.
Those were good months. Sindri would bring Ferghal a product problem, "estimators don't trust a flag with no reason attached," and he'd come back with an interface for it inside a sprint. Ferghal would flag a modeling tradeoff, "faster matching costs us some accuracy," and Sindri would tell him which side of that tradeoff Slateford actually needed. Neither of them thought about whose job anything was. It just worked.
It thinned in three small moves, none of them a fight. First, Sindri stopped asking why the confidence cutoff sat at 85 percent, because Ferghal had a chart showing that's where the model's own numbers looked best, and a chart felt like an answer. Second, when a new hire asked in a standup who actually owned that number, Sindri said "engineering, it's a model thing," and nobody corrected her. Third, when Ferghal quietly moved the cutoff from 83 to 85 during a retrain because it tested better, nobody from product was in the room, and nobody asked to be.
Then, in a planning meeting, Sindri wrote into a spec that Formwork should use "a fine-tuned classifier retrained every Friday night." Ferghal pushed back hard, right there: "That's not your call. That's architecture." He was right, and Sindri knew it before he finished the sentence. She rewrote the spec that afternoon to state the outcome she actually needed, accurate matches, fresh prices, and left the how to him. It cost about 45 minutes of friction. Everyone in the room saw it happen, and it was over by lunch.
Three weeks after that, the Slateford bid went out with the rebar line auto-approved at 89 percent confidence. Nobody argued about that one, because nobody saw it happen. It just cleared the line, the way hundreds of lines had cleared it before, correctly, every time the underlying price data was current.
Sindri thought back to the kickoff, over a year earlier, where 85 percent got chosen. Someone had asked, mildly, whether product should sign off on the cutoff. The answer at the time was reasonable: the line items were small then, a mismatch cost at most a few hundred dollars, and slowing down a fast-moving build for a sign-off process felt like the wrong tradeoff. Nobody ever circled back, because nothing forced the question until Formwork was auto-filling material orders worth tens of thousands of dollars.
What changed after the rebar line wasn't a new rule about who could talk to whom. Sindri wrote the threshold down as a product artifact: a dollar-risk cap per material category, and a staleness check that has to pass alongside the confidence score before anything auto-approves. Ferghal built it however he wanted, and moved the actual number twice since, both times with a written note back to Sindri explaining why. Overrides, the number of times either of them second-guessed the other's call outright, dropped from an average of five or six a month to under one. The replay isn't a better argument. It's a Tuesday, six months later, where nobody argues about it at all.
What I'd tell my past self, at that first kickoff: skipping the sign-off wasn't careless. It was the sensible call when the numbers were small. Nobody ever agreed to revisit it once the numbers stopped being small, and by then the line had been quietly deciding itself for months.
PICK, or pricing the boundary instead of guessing at it
The recap, one line per letter: say the line first, name who feels each kind of mistake, optimize against the one that hides, and watch a number that would tell you the line itself is wrong, not just this one argument.
One alternative we actually considered and rejected: writing a fixed decision-rights document up front, a RACI-style list covering every possible call between product and engineering. It lost, because Formwork's decisions change shape every few months as the model and the bid data change, and a fixed list goes stale and stops getting opened. A live cost-asymmetry test travels better than a document nobody reads. The AI-specific failure worth naming directly is confident-but-stale matching: a text match on a line item can score high on wording similarity while the price behind it has quietly drifted, because the drift lives in the world's prices, not in the words being matched. The guardrail is a second signal, a staleness check per material category, required to pass alongside the confidence score before anything auto-approves. And the trade-off is real and accepted on purpose: a lower, more conservative dollar-risk cap catches more of these stale-price misses, but it means more lines get flagged, which slows down a bid that's often due the next morning. Underpin accepts a few more minutes of review time as the price of not silently eating a client's margin.
And if you want to be sure it really works, try it somewhere else
Same four letters, an insurance claims tool instead of a construction bid, and the disputed number is a dollar cap on a photo, not a price on a ton of steel.
Wexmoor is Cravensworth Mutual's tool for property claims: it reads a submitted claim and either auto-approves a payout or flags it for a human adjuster. Ombretta Castelnuovo is the AI PM. Torbjorn Manx is the engineering manager. Their version of the fight was the same shape as Sindri and Ferghal's, over a different number: auto-approve any claim under $2,000 with damage-assessment confidence above 90 percent.
Mapped onto PICK: Ombretta's position is that she owns the dollar cap and the confidence bar for auto-approval, and what happens to anything under it; Torbjorn owns the computer-vision model reading the damage photos, its training data, and the infrastructure serving it. The impact: Torbjorn feels overruled if Ombretta specifies which model to fine-tune; Cravensworth's loss ratio feels it, quietly, if the dollar cap is set purely off model accuracy instead of actual fraud and misestimate exposure. The cost asymmetry is identical in shape: a pushback in planning is loud and fixed same day; a miscalibrated cap is invisible until a quarter's loss numbers come in high, or a pattern of claims that needed a human eye got waved through instead. The kill criteria transfer directly: too many product reversals of technical calls with no new evidence, move the line toward engineering; a written product threshold getting silently redefined during a model update, tighten the process instead of re-arguing it.
Swap the trigger and it still runs.
Speed: an interviewer caps you at 90 seconds. Skip straight to it: state your position, then name the one cheap-versus-expensive mistake pair, done.
Cost: no time to write a formal threshold document this sprint. Add one required field to the ticket system you already use, who owns this number, that's nearly free and gets most of the benefit.
The model got better, for real: say Formwork's matching accuracy climbs to 99 percent. Keep the dollar-risk cap process anyway, because "how much are we willing to auto-approve without a look" doesn't stop being a risk question just because the model got better at the wording part.
Where people run it wrong.
They write a permanent decision-rights document instead of a live, checkable cost test.
They let whoever sounds more technical win the argument, instead of asking what the number actually decides.
They treat "the threshold" as settled forever once it's written down, instead of watching the override rate for a sign it needs to move again.
How to use it live. When an interviewer says "so whose call is this, really," ask one thing back before answering: "which mistake here is expensive if it ships quiet, and which one just gets argued about?" That's usually the exact test being run.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the EM just refuses to hand that number over?" Response: point at the $41,000 case. The ask was never "give me the number," it was "let's agree, in writing, on a number we're both willing to defend to Slateford," which makes it shared accountability, not a power grab.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on AI PM vs traditional PM vs technical PM
- #1 List four responsibilities an AI PM holds that a traditional PM does not.
- #2 Which parts of the classic PM toolkit transfer unchanged to AI products, and which do not?
- #3 Explain why an AI PM often owns the evaluation set while a traditional PM would not own a test plan.
- #4 How does the discovery phase differ when feasibility is genuinely unknown until you build?
- #5 Describe the difference between an AI PM and an ML PM at a company that has both.
- #6 Why does the AI PM role pull the PM further into the technical stack than most PM roles?