Explain why velocity-based planning breaks down on AI projects.
Openline is Duskglass Labs' tool for writing dating-app bios and opening lines from a few facts someone types in. Loveday Hallanby runs sprint planning for the squad that builds it. For thirteen sprints the backlog was almost entirely account and screen work, and the team's velocity predicted ship dates within two or three days, every single time. Then AI-behavior tickets started entering the backlog, and by sprint twenty the same velocity number was missing dates by weeks, with nobody able to say when it had actually started going wrong.
- Recut the backlog by ticket type, CRUD or UI against AI-behavior, before trusting one blended velocity number.Why: a healthy-looking average can hide a slice that has already cratered.
- Track estimate accuracy, actual effort against points, separately for AI-tagged tickets.Why: this is the one check that shows the gap, months before a missed release date makes it obvious.
- Rule out capacity and process first: headcount, meeting load, new tools.Why: blaming the AI work without checking wastes the fix if something more ordinary actually moved.
- Stop sizing "does the model clear a bar" tickets on the same scale as "build the feature" tickets.Why: one assumes solvable effort, the other assumes unknown feasibility, and one number can't hold both.
- Give tickets with no natural finish line a timebox, not a point estimate.Why: a subjective quality bar will eat any number of points you hand it and still not be done.
- Leave normal CRUD and UI tickets on the old velocity system.Why: their points were never broken, only the newly mixed-in AI tickets were.
How to answer this, stage by stage
Nobody is grading whether you can say "AI work is hard to estimate." They're grading whether you can name the actual reason a point stops meaning one thing, and the one check that confirms it before a launch date gets missed in public.
Let's learn
What happens when the number a team trusts to predict a ship date keeps working, right up until it doesn't, and nobody can say when it stopped.
Openline reads a few facts someone types in, their job, a couple of hobbies, a tone they'd like, and drafts three candidate bios and an opening line. Before it existed, most people spent twenty to thirty minutes staring at a blank profile field, or copied a bio from a friend and hoped it fit. With Openline, that dropped to about two minutes: type the facts, pick a draft, adjust a line, done.
For thirteen sprints, the backlog was almost entirely CRUD and UI work: the onboarding quiz, a tone selector screen, a save-for-later button, an account settings page, a referral code flow. Velocity sat at 41 points a sprint, steady, rarely moving more than two points either way. Commit-hit-rate, the share of committed points that actually shipped inside the sprint they were committed to, held at 91 percent. Release forecasts built from that velocity landed within two or three days of the real ship date, every time, going back over a year.
Sprint fourteen, the first AI-behavior ticket entered the backlog: "Make the playful-tone bios sound less generic." It was sized at five points, the same size as a save-for-later button that had shipped exactly on schedule three sprints earlier.
The tone ticket didn't close in sprint fourteen. It reopened in fifteen, reopened again in sixteen, and finally reached a "good enough" bar in seventeen, three extra sprints for a ticket sized to fit inside one. Around the same time, a second AI ticket landed: "Stop Openline inventing hobbies the user never typed in." Sized at eight points, it dropped the invented-detail rate from 6 percent to 2 percent in its first pass, then reopened in sprint eighteen to push further, to 0.4 percent, while the team argued whether a number that still wasn't zero counted as finished.
By sprint twenty, about ten months in, AI-tagged tickets made up 35 percent of committed sprint points. The team's overall commit-hit-rate had fallen from 91 percent to 54 percent. Forecasts that used to land within two or three days were now missing by three weeks or more. Concretely: Duskglass had told marketing a tone-matching relaunch would ship by a fixed date for a public campaign. It shipped 24 days late.
Here's the turn. The extra weeks were never really the story. The real problem was that the one number the whole team trusted to predict a ship date had quietly stopped measuring one thing, and nothing on the sprint board said so.
What I would leave alone: the CRUD and UI backlog, settings, onboarding, the referral flow, never needed anything different. Their points stayed accurate the entire time, right alongside the AI tickets falling apart next to them on the same board.
The lesson: a team doesn't need to get worse at estimating for velocity to stop working. It just needs a second kind of ticket to start sharing a scale that was only ever calibrated for the first kind.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one when you want to feel exactly what "5 points" stopped meaning, and why nobody caught it sooner.
Loveday Hallanby has run sprint planning for Duskglass's bio-quality squad for three years, and the board she keeps is the kind other teams point to. Committed points and delivered points land on top of each other, sprint after sprint. Ask her for a ship date in October and she'll give you a day, not a month, and she'll be right.
For thirteen sprints, that reputation held without her having to think about it. The backlog was onboarding screens, settings pages, a referral flow, a save-for-later button. She'd estimate a ticket, watch it come in close, and move on. When a ticket ran a little long, another ran a little short, and the two canceled out the way they always had. She stopped double-checking individual misses around sprint nine, because for months the misses were noise, never a pattern.
Sprint fourteen, a ticket came in from the roadmap: "Make the playful-tone bios sound less generic." Loveday sized it the way she sized everything, against the last comparable thing on the board. A save-for-later button, also five points, had shipped exactly on schedule in sprint eleven. Five felt right. She wrote it in the box.
It didn't close. It came back in the sprint fifteen retro, half finished, and again in sixteen. Loveday wasn't alarmed. Tickets slipped sometimes; she'd seen it before, always for a reason that resolved itself in a sprint or two. This one resolved in three, finally reaching "good enough" in sprint seventeen. By then a second AI ticket, "stop inventing hobbies the user never typed in," was already open, and it wasn't closing cleanly either, dropping from 6 percent to 2 percent and then getting reopened to chase 0.4, with nobody quite agreeing on whether 0.4 was the finish line or just the next stop.
Loveday didn't connect it yet. A slipped ticket here, a reopened one there, nothing that looked, sprint to sprint, like anything other than ordinary variance. By sprint twenty, roughly a third of the board's points sat in tickets like these, and commit-hit-rate had drifted from 91 percent down to 54, a little at a time, never with a single sprint bad enough on its own to demand an explanation.
Jorunn Sennwright, who owns the roadmap, had already told marketing a relaunch date for Openline's tone-matching feature, built off the velocity number Loveday's board had always delivered on. The date came and went. It shipped 24 days later, into a campaign that had already run.
The moment that actually cracked it open wasn't the missed date. It was smaller than that. Tsering Oakburn, three weeks into the job, sat in sprint twenty one planning and asked a plain question nobody senior had thought to ask in months: why was the "stop inventing hobbies" ticket, still open after three sprints, sized the same as a button that had shipped in one. Loveday started to answer with the usual line, estimates are just estimates, and stopped halfway through the sentence, because she didn't actually have a better answer than that.
She pulled the last six sprints of ticket data that night. Split by type: CRUD and UI tickets, burn ratio 1.04, tight, barely moving. AI-tagged tickets, burn ratio 2.3 on average, but ranging from 0.9 to 4.1, a spread wide enough that no single number could ever have represented it honestly.
She thought back to the kickoff meeting, over a year earlier, where the point scale had been set. Someone had asked whether AI-behavior work should be estimated differently from day one. The answer, reasonable at the time, was that there wasn't enough of it yet to bother, one shared scale was simpler, and they'd revisit if it ever became a real share of the board. Nobody ever circled back, because nothing forced the question until a third of the board was AI-tagged and the board had already stopped telling the truth.
What Loveday would tell herself, back at that kickoff: skipping a separate estimation practice for AI work wasn't careless. It was the sensible call when AI-behavior tickets were one a quarter. Nobody ever agreed to revisit it once they became one in three, and by then the board had been quietly lying for months, in a language that looked exactly like ordinary variance right up until it didn't.
TRACE, and why a ticket's own shape decides whether a point means anything
Not a way to spot a bad estimate. TRACE is what you run when the estimates all looked reasonable individually, because that's exactly how this kind of drift hides.
The recap, one line per letter: thirteen clean sprints, then a ticket type nobody re-scoped the ruler for. Recut by type, not sprint, and the average splits cleanly into two very different stories. Capacity and process never moved. Three habits, not one mystery, explain the gap. One burn-ratio check, run on data the team already had, would have shown it seven sprints earlier.
Three things worth stating directly, since this is where the real judgment sits. The alternative Loveday's team considered, and rejected, was multiplying every AI-tagged ticket's estimate by a fixed factor, say three times its CRUD-equivalent size, instead of pulling AI tickets out of blended velocity entirely. It lost, because the burn ratio itself ranged from 0.9x to 4.1x across six sprints; a fixed multiplier would just have been a different wrong number wearing a fix's clothes. The AI-specific failure worth naming by name is hallucination: Openline sometimes filled a bio with a hobby, an employer, or a detail the user never typed in, because a tidy, plausible story was closer to what its best-tested prompts had always produced. The guardrail is a grounding check: before any draft reaches a user, flag any concrete noun phrase in the bio that doesn't trace back to something the user actually entered. And the trade-off is real, and accepted on purpose: timeboxing an AI-behavior ticket to two sprints instead of a point estimate trades away a perfect, undefined bar for a schedule the team can actually promise to marketing, and Duskglass accepts shipping "good enough, measured" bios over chasing a target that was never going to hold still.
And if you want to be sure it really works, try it somewhere else
Same five letters, a legal-translation tool instead of a dating app, and this time the ticket with no natural finish line isn't about tone. It's about whether a contract clause survives translation with its meaning intact.
Faithword, built by Nettlebridge Linguistics, drafts and reviews translations of contracts and filings for translation agencies, flagging risky or ambiguous renderings for a human reviewer. Zephira Castelic leads engineering there, and Faithword's backlog broke the same way Openline's did, on a different kind of sentence.
CRUD tickets there, an upload flow, a reviewer assignment queue, a billing export, held a burn ratio near 1.06 across six sprints, steady the whole way. One AI-tagged ticket, "reduce mistranslation of standard indemnity clause boilerplate," was sized at five points, matched against a CRUD ticket of the same size that had shipped on schedule. It took two extra sprints to reach a reviewer-approved bar, and the AI-tagged slice across six sprints averaged a burn ratio of 2.1, ranging from 0.8 to 3.6.
Mapped straight onto TRACE: the timeline is a five-point ticket that entered looking ordinary and took two extra sprints before anyone flagged the pattern. The recut is CRUD tickets near 1.06 against AI-tagged tickets averaging 2.1, spread from 0.8 to 3.6. Assume nothing rules out headcount and process, both held steady while the mix of ticket types shifted. The cause candidates are the same three habits in new clothes: whether the model can preserve legal meaning across two languages' idioms is a feasibility unknown, not an effort estimate; "close enough" for a legal rendering is a judgment call with no binary check; and a fine-tune pass touches every clause type at once, so the work can't be split across engineers the way a queue screen can. The evidence test is identical in shape: burn ratio by ticket type, over six sprints, no argument needed once it's on the table.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: recut the backlog by ticket type, and track burn ratio for AI-tagged tickets on its own, full stop.
Cost: no time this sprint to build new tracking. Add a single dropdown to the ticket tool you already use, AI-tagged or not, that's nearly free, then compute the split from data you already have.
The model got better, for real: say Openline's hallucination rate drops close to zero after a base-model upgrade. Keep timeboxing AI-behavior tickets anyway, because "no natural finish line" doesn't go away just because today's bar got easier to clear. The next quality target will have the same shape.
Where people run it wrong.
They add a fixed story-point multiplier for AI work instead of tracking the real, unstable ratio.
They blame the individual engineer's estimate instead of the ticket type.
They keep chasing a subjective quality bar past the timebox because stopping feels like giving up, instead of shipping "good enough, measured" and moving on.
How to use it live. When an interviewer says "your team's velocity has gotten unreliable, what do you do," ask one thing back before answering: "is the backlog still one kind of work, or did a second kind quietly get mixed into the same points scale?" That question alone is usually the exact distinction being tested.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't this just proof AI tickets should never share a board with everything else?" Response: no, they stay on the same board. Only the estimation and tracking method splits: CRUD keeps its points, AI-behavior tickets get a timebox and their own burn-ratio tracking.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on AI PM vs traditional PM vs technical PM
- #1 List four responsibilities an AI PM holds that a traditional PM does not.
- #2 Which parts of the classic PM toolkit transfer unchanged to AI products, and which do not?
- #3 Explain why an AI PM often owns the evaluation set while a traditional PM would not own a test plan.
- #4 How does the discovery phase differ when feasibility is genuinely unknown until you build?
- #5 Describe the difference between an AI PM and an ML PM at a company that has both.
- #6 Why does the AI PM role pull the PM further into the technical stack than most PM roles?