ConceptIntermediateModel Fluency & the AI PM Role / AI PM vs traditional PM vs technical PM / #10

Explain why velocity-based planning breaks down on AI projects.

TRACE · a sprint-planning breakdown on Openline, an AI bio and opening-line writer for daters

Openline is Duskglass Labs' tool for writing dating-app bios and opening lines from a few facts someone types in. Loveday Hallanby runs sprint planning for the squad that builds it. For thirteen sprints the backlog was almost entirely account and screen work, and the team's velocity predicted ship dates within two or three days, every single time. Then AI-behavior tickets started entering the backlog, and by sprint twenty the same velocity number was missing dates by weeks, with nobody able to say when it had actually started going wrong.

The direct answer
Velocity breaks the moment two different kinds of work share one points scale. A CRUD ticket's point measures how many hours a solvable task takes; an AI-behavior ticket's real unknown is whether the model can clear a subjective quality bar at all, which no point can represent. Recut the backlog by ticket type the moment AI-behavior tickets appear, track their estimate accuracy on its own, and give them a timebox instead of a point, because a blended average can look healthy for months while that one slice is already failing underneath it.
Do this, in order
  1. Recut the backlog by ticket type, CRUD or UI against AI-behavior, before trusting one blended velocity number.Why: a healthy-looking average can hide a slice that has already cratered.
  2. Track estimate accuracy, actual effort against points, separately for AI-tagged tickets.Why: this is the one check that shows the gap, months before a missed release date makes it obvious.
  3. Rule out capacity and process first: headcount, meeting load, new tools.Why: blaming the AI work without checking wastes the fix if something more ordinary actually moved.
  4. Stop sizing "does the model clear a bar" tickets on the same scale as "build the feature" tickets.Why: one assumes solvable effort, the other assumes unknown feasibility, and one number can't hold both.
  5. Give tickets with no natural finish line a timebox, not a point estimate.Why: a subjective quality bar will eat any number of points you hand it and still not be done.
  6. Leave normal CRUD and UI tickets on the old velocity system.Why: their points were never broken, only the newly mixed-in AI tickets were.

How to answer this, stage by stage

Nobody is grading whether you can say "AI work is hard to estimate." They're grading whether you can name the actual reason a point stops meaning one thing, and the one check that confirms it before a launch date gets missed in public.

1
Scope it to one team and one sprint board
Say it like this
"Let's make this concrete. Openline drafts dating bios from a few facts someone types in. For thirteen sprints, about six and a half months, Duskglass's velocity sat at 41 points, steady, and release forecasts landed within two or three days. Then AI-behavior tickets started entering the backlog."
Why this works
A named product with a real number is something an interviewer can follow. "Velocity gets unreliable" on its own is not a story yet.
2
Say the structure out loud
Say it like this
"I'll run this as TRACE. Lay out the timeline, recut the backlog by ticket type, rule out capacity and process, name the real causes, then give the one check that confirms it."
Why this works
Two seconds of structure tells the interviewer you have a method, not five scattered guesses arriving as they occur to you.
3
Reframe the question before naming a single fact
Say it like this
"Velocity only works if a point means roughly the same thing every time you use it. The question isn't whether the team got worse at estimating. It's whether a ticket that asks 'can the model clear this bar at all' can be sized with the same ruler as a ticket that just asks 'how long does this take to wire up.'"
Why this works
This line is the whole answer in miniature. Skip it and the rest sounds like a complaint about one bad quarter instead of a real pattern.
4
Lay out the timeline
Say it like this
"Sprint fourteen, the first AI-behavior ticket enters, sized five points, the same size as a save-for-later button that had shipped on schedule three sprints earlier. It doesn't close in fourteen. It reopens in fifteen, sixteen, and finally clears in seventeen. Nobody names the pattern until sprint twenty one, when a new hire asks why."
Why this works
Naming exactly when the first AI ticket landed, and how much later anyone noticed, stops the story from sounding like a sudden, unexplainable slide.
5
Recut the backlog by type, not by sprint
Say it like this
"Split it by ticket type, over the same six sprints. CRUD and UI tickets: burn ratio right around one point oh four, barely off. AI-tagged tickets: two point three on average, and anywhere from point nine to four point one. That spread is the whole diagnosis in two numbers."
Why this works
This is the move most candidates skip. A blended average hides exactly the split that explains everything.
6
Rule out the boring explanations first
Say it like this
"Before blaming the AI tickets, rule out the ordinary suspects. Nobody left the team. No new process slowed anyone down. Meeting load is flat, sprint over sprint. None of that moved. Only the mix of ticket types moved, starting sprint fourteen."
Why this works
Naming this out loud shows you're not reaching for the flashiest explanation before checking the plain ones.
7
Name the three real causes
Say it like this
"Three things, all real, all at once. One, a ticket like 'make the tone less generic' isn't unsure how long it takes, it's unsure whether the model can hit the bar at all. Two, 'stop inventing hobbies' has no natural finish line, it goes from six percent to two to point four and the team can argue forever about what counts as done. Three, you can't split a prompt fix across three engineers the way you split a button into three sub-tasks, because one change touches every bio Openline writes at once."
Why this works
Three named, checkable reasons beat one vague line about "AI work being unpredictable."
8
Give the evidence test, and close on the one line
Say it like this
"The check that confirms it: pull six sprints, compute burn ratio by ticket type. If CRUD stays near one and AI-tagged tickets spread wide, that's the finding, no argument needed. So: velocity breaks the moment two different kinds of work share one points scale, and the fix is splitting the ruler, not blaming the team."
Why this works
Leaves the interviewer with a concrete, runnable check, not just a warning to plan for uncertainty.

Let's learn

What happens when the number a team trusts to predict a ship date keeps working, right up until it doesn't, and nobody can say when it stopped.

Openline reads a few facts someone types in, their job, a couple of hobbies, a tone they'd like, and drafts three candidate bios and an opening line. Before it existed, most people spent twenty to thirty minutes staring at a blank profile field, or copied a bio from a friend and hoped it fit. With Openline, that dropped to about two minutes: type the facts, pick a draft, adjust a line, done.

Hand sketched horizontal timeline titled Openline's sprint board, 13 to 21. Four milestones. Sprint 13, caption 91 percent commit-hit-rate. Sprint 14, this milestone emphasized in purple, caption first AI ticket, sized 5. Sprint 17, caption clears 3 sprints late. Sprint 21, caption Tsering asks why.
Sprint fourteen looked like an ordinary Tuesday. Nobody circled it until sprint twenty one.

For thirteen sprints, the backlog was almost entirely CRUD and UI work: the onboarding quiz, a tone selector screen, a save-for-later button, an account settings page, a referral code flow. Velocity sat at 41 points a sprint, steady, rarely moving more than two points either way. Commit-hit-rate, the share of committed points that actually shipped inside the sprint they were committed to, held at 91 percent. Release forecasts built from that velocity landed within two or three days of the real ship date, every time, going back over a year.

Sprint fourteen, the first AI-behavior ticket entered the backlog: "Make the playful-tone bios sound less generic." It was sized at five points, the same size as a save-for-later button that had shipped exactly on schedule three sprints earlier.

Knowledge spark: what's a burn ratio? Actual effort a ticket ate, divided by the points it was estimated at. A ticket sized five that took five points' worth of real work has a burn ratio of 1.0. A ticket sized five that ate more than double that has a burn ratio over 2.0.

The tone ticket didn't close in sprint fourteen. It reopened in fifteen, reopened again in sixteen, and finally reached a "good enough" bar in seventeen, three extra sprints for a ticket sized to fit inside one. Around the same time, a second AI ticket landed: "Stop Openline inventing hobbies the user never typed in." Sized at eight points, it dropped the invented-detail rate from 6 percent to 2 percent in its first pass, then reopened in sprint eighteen to push further, to 0.4 percent, while the team argued whether a number that still wasn't zero counted as finished.

Hand sketched comparison diagram titled Two tickets, sized 5, that were not the same job. Left panel, a box icon labeled Save-for-later button, caption 3 clean sub-tasks, 3 engineers, in parallel. Right panel, a question mark icon labeled Make tone less generic, caption 1 prompt change, touches every bio at once.
Same number in the estimate column. Not close to the same shape of work underneath it.

By sprint twenty, about ten months in, AI-tagged tickets made up 35 percent of committed sprint points. The team's overall commit-hit-rate had fallen from 91 percent to 54 percent. Forecasts that used to land within two or three days were now missing by three weeks or more. Concretely: Duskglass had told marketing a tone-matching relaunch would ship by a fixed date for a public campaign. It shipped 24 days late.

Burn ratio by ticket type, sprints 15 to 20
3.0x 1.5x 0x 1.04x CRUD / UI tickets 2.3x AI-tagged tickets
CRUD / UI tickets, average burn ratioAI-tagged tickets, average burn ratio
The AI-tagged bar isn't just higher. Across the six sprints it ranged from 0.9x to 4.1x, wildly inconsistent, which is a different, worse problem than "runs a bit long."

Here's the turn. The extra weeks were never really the story. The real problem was that the one number the whole team trusted to predict a ship date had quietly stopped measuring one thing, and nothing on the sprint board said so.

We didn't lose two weeks of accuracy all at once. We lost it exactly where the backlog stopped being one kind of thing.
Commit-hit-rate by sprint, sprint 10 to 20
100% 50% 0% Sprint 14: first AI ticket enters Sp10 Sp13 Sp15 Sp17 Sp19 Sp20
Commit-hit-rate, by sprintSprint the first AI ticket entered
The number started sliding at sprint fourteen. Nobody named the pattern until sprint twenty one, seven sprints after it began.
The decision I would take back Duskglass estimated every ticket on one shared point scale, calibrated entirely against CRUD and UI work, and never built a separate practice for AI-behavior tickets. That made complete sense when Openline's roadmap was 95 percent forms, flows, and screens. It stopped making sense the moment "does the model clear a quality bar" tickets started showing up in the same backlog, sized with the same ruler as everything else.

What I would leave alone: the CRUD and UI backlog, settings, onboarding, the referral flow, never needed anything different. Their points stayed accurate the entire time, right alongside the AI tickets falling apart next to them on the same board.

The lesson: a team doesn't need to get worse at estimating for velocity to stop working. It just needs a second kind of ticket to start sharing a scale that was only ever calibrated for the first kind.

Now here is the same thing as a story

The short version above is what you actually say in the room. Read this one when you want to feel exactly what "5 points" stopped meaning, and why nobody caught it sooner.

Loveday Hallanby has run sprint planning for Duskglass's bio-quality squad for three years, and the board she keeps is the kind other teams point to. Committed points and delivered points land on top of each other, sprint after sprint. Ask her for a ship date in October and she'll give you a day, not a month, and she'll be right.

For thirteen sprints, that reputation held without her having to think about it. The backlog was onboarding screens, settings pages, a referral flow, a save-for-later button. She'd estimate a ticket, watch it come in close, and move on. When a ticket ran a little long, another ran a little short, and the two canceled out the way they always had. She stopped double-checking individual misses around sprint nine, because for months the misses were noise, never a pattern.

Sprint fourteen, a ticket came in from the roadmap: "Make the playful-tone bios sound less generic." Loveday sized it the way she sized everything, against the last comparable thing on the board. A save-for-later button, also five points, had shipped exactly on schedule in sprint eleven. Five felt right. She wrote it in the box.

Hand sketched icon list titled What we ruled out before blaming the AI tickets. Four numbered rows. One, no one left the team, headcount held steady. Two, no new process or tool slowed anyone down. Three, meeting load stayed flat, sprint over sprint. Four, this row in rose, only the mix of ticket types changed, from sprint 14 on.
None of the ordinary explanations moved. That's what made the real one worth digging for.

It didn't close. It came back in the sprint fifteen retro, half finished, and again in sixteen. Loveday wasn't alarmed. Tickets slipped sometimes; she'd seen it before, always for a reason that resolved itself in a sprint or two. This one resolved in three, finally reaching "good enough" in sprint seventeen. By then a second AI ticket, "stop inventing hobbies the user never typed in," was already open, and it wasn't closing cleanly either, dropping from 6 percent to 2 percent and then getting reopened to chase 0.4, with nobody quite agreeing on whether 0.4 was the finish line or just the next stop.

Loveday didn't connect it yet. A slipped ticket here, a reopened one there, nothing that looked, sprint to sprint, like anything other than ordinary variance. By sprint twenty, roughly a third of the board's points sat in tickets like these, and commit-hit-rate had drifted from 91 percent down to 54, a little at a time, never with a single sprint bad enough on its own to demand an explanation.

A blended average never lies exactly. It just stops telling you the thing you actually need to know.

Jorunn Sennwright, who owns the roadmap, had already told marketing a relaunch date for Openline's tone-matching feature, built off the velocity number Loveday's board had always delivered on. The date came and went. It shipped 24 days later, into a campaign that had already run.

The moment that actually cracked it open wasn't the missed date. It was smaller than that. Tsering Oakburn, three weeks into the job, sat in sprint twenty one planning and asked a plain question nobody senior had thought to ask in months: why was the "stop inventing hobbies" ticket, still open after three sprints, sized the same as a button that had shipped in one. Loveday started to answer with the usual line, estimates are just estimates, and stopped halfway through the sentence, because she didn't actually have a better answer than that.

She pulled the last six sprints of ticket data that night. Split by type: CRUD and UI tickets, burn ratio 1.04, tight, barely moving. AI-tagged tickets, burn ratio 2.3 on average, but ranging from 0.9 to 4.1, a spread wide enough that no single number could ever have represented it honestly.

Hand sketched quadrant diagram titled Why one ruler can't size both. X axis, how solvable the effort is, from clear path to unknown if possible. Y axis, how defined done is, from subjective and moving to binary and checkable. Save-for-later button and referral code flow plotted top left, clear path and binary. Tone less generic and stop inventing hobbies plotted bottom right, unknown if possible and subjective.
The CRUD tickets and the AI tickets were never scattered randomly. They sat in two different corners the whole time.

She thought back to the kickoff meeting, over a year earlier, where the point scale had been set. Someone had asked whether AI-behavior work should be estimated differently from day one. The answer, reasonable at the time, was that there wasn't enough of it yet to bother, one shared scale was simpler, and they'd revisit if it ever became a real share of the board. Nobody ever circled back, because nothing forced the question until a third of the board was AI-tagged and the board had already stopped telling the truth.

Hand sketched full page metaphor scene titled One ruler, two different jobs. Left panel, a scale icon labeled CRUD WORK, caption a tape measure against a wall, same number every time. Right panel, a question mark icon labeled AI-BEHAVIOR WORK, caption a tape measure against fog, the target keeps moving.
The whole answer, in one picture. A number can hold still against a wall. It can't hold still against fog.

What Loveday would tell herself, back at that kickoff: skipping a separate estimation practice for AI work wasn't careless. It was the sensible call when AI-behavior tickets were one a quarter. Nobody ever agreed to revisit it once they became one in three, and by then the board had been quietly lying for months, in a language that looked exactly like ordinary variance right up until it didn't.

TRACE, and why a ticket's own shape decides whether a point means anything

Not a way to spot a bad estimate. TRACE is what you run when the estimates all looked reasonable individually, because that's exactly how this kind of drift hides.

Hand sketched labeled parts diagram titled TRACE, on a velocity chart that stopped predicting anything. A gauge icon at the center labeled Why velocity broke, with five callouts arranged around it: timeline when it started, recut by ticket type, assume nothing rule out capacity, cause candidates 3 real reasons, evidence test burn ratio by type.
Five checks, and every one of them assumes the velocity number in front of you still looks basically fine.
TTimeline. When it shipped, and when it actually started drifting.
Thirteen clean sprints, 91 percent commit-hit-rate. Sprint fourteen, the first AI ticket enters, sized like a button. It doesn't clear until sprint seventeen. The pattern isn't named until sprint twenty one, seven sprints after it began.
Puts a real clock on the gap between "the number started sliding" and "someone asked why."
RRecut. Slice by ticket type, not by sprint or by person.
CRUD and UI tickets: burn ratio 1.04, tight and consistent across six sprints. AI-tagged tickets: 2.3 on average, ranging from 0.9 to 4.1.
The strongest move here: a blended 1.5-ish average would have looked survivable and hidden the exact same story just as well as 91 percent once did.
AAssume nothing. Rule out capacity and process before blaming the work itself.
Headcount held steady. No new process or tool entered the picture. Meeting load stayed flat, sprint over sprint. Only the mix of ticket types moved, starting sprint fourteen.
This is where most candidates skip straight to blame. Ruling out the boring explanations first is what makes the real one credible.
CCause candidates. Three real reasons, not one vague complaint.
One: a "build the feature" ticket like the tone one isn't unsure how long it takes, it's unsure whether the model can hit the bar at all. Two: an eval-style ticket like the hobby one has no natural finish line the way a button does, it can always go a little further. Three: a prompt or model change touches every bio at once, so the work can't be split into independent sub-tasks the way a UI ticket can.
Three specific, checkable habits beat one line about "AI work being unpredictable."
EEvidence test. The one check that confirms it.
Pull the last six sprints. Compute burn ratio, actual effort divided by estimated points, split by whether the ticket is AI-tagged or not. If CRUD stays near 1.0 and AI-tagged tickets spread wide, that's the confirmation, no argument needed.
Turns "the estimates feel off lately" into a number a planning meeting can actually act on.

The recap, one line per letter: thirteen clean sprints, then a ticket type nobody re-scoped the ruler for. Recut by type, not sprint, and the average splits cleanly into two very different stories. Capacity and process never moved. Three habits, not one mystery, explain the gap. One burn-ratio check, run on data the team already had, would have shown it seven sprints earlier.

Three things worth stating directly, since this is where the real judgment sits. The alternative Loveday's team considered, and rejected, was multiplying every AI-tagged ticket's estimate by a fixed factor, say three times its CRUD-equivalent size, instead of pulling AI tickets out of blended velocity entirely. It lost, because the burn ratio itself ranged from 0.9x to 4.1x across six sprints; a fixed multiplier would just have been a different wrong number wearing a fix's clothes. The AI-specific failure worth naming by name is hallucination: Openline sometimes filled a bio with a hobby, an employer, or a detail the user never typed in, because a tidy, plausible story was closer to what its best-tested prompts had always produced. The guardrail is a grounding check: before any draft reaches a user, flag any concrete noun phrase in the bio that doesn't trace back to something the user actually entered. And the trade-off is real, and accepted on purpose: timeboxing an AI-behavior ticket to two sprints instead of a point estimate trades away a perfect, undefined bar for a schedule the team can actually promise to marketing, and Duskglass accepts shipping "good enough, measured" bios over chasing a target that was never going to hold still.

And if you want to be sure it really works, try it somewhere else

Same five letters, a legal-translation tool instead of a dating app, and this time the ticket with no natural finish line isn't about tone. It's about whether a contract clause survives translation with its meaning intact.

Faithword, built by Nettlebridge Linguistics, drafts and reviews translations of contracts and filings for translation agencies, flagging risky or ambiguous renderings for a human reviewer. Zephira Castelic leads engineering there, and Faithword's backlog broke the same way Openline's did, on a different kind of sentence.

Hand sketched decision tree titled Faithword's new rule for sizing a ticket. Root box reads A new ticket enters the backlog, branching into three outcomes. Clear CRUD or UI shape leads to size it normally, trust the points. Quality bar unclear, no fixed finish line leads to timebox it, track burn ratio apart. Not sure which it is leads to ask, is there a binary finish line.
Same shape of fix as Duskglass's, built as a rule the whole team can apply to the next ticket, not just this one.

CRUD tickets there, an upload flow, a reviewer assignment queue, a billing export, held a burn ratio near 1.06 across six sprints, steady the whole way. One AI-tagged ticket, "reduce mistranslation of standard indemnity clause boilerplate," was sized at five points, matched against a CRUD ticket of the same size that had shipped on schedule. It took two extra sprints to reach a reviewer-approved bar, and the AI-tagged slice across six sprints averaged a burn ratio of 2.1, ranging from 0.8 to 3.6.

The decision Nettlebridge would take back Faithword launched with the same shape of point scale Openline started with: one shared ruler, calibrated against CRUD work, with no separate track for eval-style tickets. It made sense for a small pilot backlog with a handful of translation-quality tickets a year. It stopped making sense once indemnity, liability, and warranty clause tickets became a real share of every sprint.

Mapped straight onto TRACE: the timeline is a five-point ticket that entered looking ordinary and took two extra sprints before anyone flagged the pattern. The recut is CRUD tickets near 1.06 against AI-tagged tickets averaging 2.1, spread from 0.8 to 3.6. Assume nothing rules out headcount and process, both held steady while the mix of ticket types shifted. The cause candidates are the same three habits in new clothes: whether the model can preserve legal meaning across two languages' idioms is a feasibility unknown, not an effort estimate; "close enough" for a legal rendering is a judgment call with no binary check; and a fine-tune pass touches every clause type at once, so the work can't be split across engineers the way a queue screen can. The evidence test is identical in shape: burn ratio by ticket type, over six sprints, no argument needed once it's on the table.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: recut the backlog by ticket type, and track burn ratio for AI-tagged tickets on its own, full stop.
Cost: no time this sprint to build new tracking. Add a single dropdown to the ticket tool you already use, AI-tagged or not, that's nearly free, then compute the split from data you already have.
The model got better, for real: say Openline's hallucination rate drops close to zero after a base-model upgrade. Keep timeboxing AI-behavior tickets anyway, because "no natural finish line" doesn't go away just because today's bar got easier to clear. The next quality target will have the same shape.

Where people run it wrong.
They add a fixed story-point multiplier for AI work instead of tracking the real, unstable ratio.
They blame the individual engineer's estimate instead of the ticket type.
They keep chasing a subjective quality bar past the timebox because stopping feels like giving up, instead of shipping "good enough, measured" and moving on.

How to use it live. When an interviewer says "your team's velocity has gotten unreliable, what do you do," ask one thing back before answering: "is the backlog still one kind of work, or did a second kind quietly get mixed into the same points scale?" That question alone is usually the exact distinction being tested.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question about why a team's velocity metric stopped predicting delivery?
Tap to flip
ANSWER
TRACE: lay out the timeline, recut by ticket type, assume nothing about capacity or process, name the real causes, then run the one evidence test.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Loveday Hallanby, who runs sprint planning for Openline's bio-quality squad; Jorunn Sennwright, the product manager who owns the roadmap; and Tsering Oakburn, the new engineer whose plain question started the dig.
3 · THE TIMELINE
What shipped, and when did the real problem actually surface?
Tap to flip
ANSWER
Thirteen clean sprints of CRUD and UI work, 91 percent commit-hit-rate. Sprint fourteen, the first AI-behavior ticket enters, sized like a button. It doesn't clear until sprint seventeen. Nobody names the pattern until sprint twenty one.
4 · THE RECUT
What's the real difference the recut shows?
Tap to flip
ANSWER
Not ticket count, ticket type. CRUD and UI tickets held a burn ratio near 1.04x across six sprints. AI-tagged tickets averaged 2.3x, swinging from 0.9x to 4.1x, wildly inconsistent, not just biased high.
5 · THE OLD DECISION
What decision would this answer take back?
Tap to flip
ANSWER
Estimating every ticket on one shared point scale, calibrated entirely against CRUD and UI work, with no separate practice for AI-behavior tickets, until AI-behavior work became a third of the board.
6 · THE NUMBER
Fill in the blank: commit-hit-rate fell from ___ percent to ___ percent between sprint 13 and sprint 20, and the public relaunch date slipped ___ days past what marketing had promised.
Tap to flip
ANSWER
91 percent to 54 percent. 24 days.
7 · THE EVIDENCE TEST
What's the one check that confirms the diagnosis?
Tap to flip
ANSWER
Pull six sprints of tickets and compute burn ratio, actual effort over estimated points, split by ticket type. If CRUD tickets stay near 1x and AI-tagged tickets spread wide, that confirms it's the ticket type, not the team.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent ticket that broke velocity there?
Tap to flip
ANSWER
Faithword, Nettlebridge Linguistics' legal-translation review tool. The equivalent ticket reduces mistranslation of standard indemnity clause boilerplate, an eval-style ticket with the same feasibility-unknown, no-natural-finish-line shape.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Openline's blended velocity number stay looking healthy for months even while the AI-tagged tickets were already failing underneath it?
  • A. CRUD and UI tickets made up most of the committed points, so their accurate estimates masked the AI slice's cratering ratio inside the average.
  • B. The team stopped reporting AI ticket results honestly.
  • C. AI tickets were tracked on a separate board nobody looked at.
  • D. Story points are never accurate for any kind of work.
Show hint
Look at the recut chart, and how far apart the CRUD and AI-tagged burn ratios actually are.
Show answer
A. A blended average is a weighted mix. A small, unstable slice can already be broken while the majority slice keeps the overall number looking fine.
Fill in the blank
2. The "make the tone less generic" ticket was estimated at ___ points, the same size as a save-for-later button that had shipped on schedule. It didn't clear until sprint ___, ___ sprints later than planned.
Show hint
Check "Let's learn," right after the knowledge spark on burn ratio.
Show answer
5 points, sprint 17, 3 sprints. A ticket sized to fit inside one sprint took four, and the estimate never signaled that in advance.
True or false
3. True or false: the real fix here was for engineers to be more careful and give bigger point estimates to AI-behavior tickets.
  • True
  • False
Show hint
Look at the burn ratio range for AI-tagged tickets, 0.9x to 4.1x, not just the average.
Show answer
False. A bigger fixed number is still one fixed number for a ratio that swings from 0.9x to 4.1x. The fix is tracking and timeboxing AI-behavior tickets separately, not estimating harder on the same scale.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The decision I would take back," in Let's learn.
Show answer
Model answer: Estimating every ticket on one shared point scale, calibrated entirely against CRUD and UI work. It made sense when Openline's roadmap was 95 percent forms, flows, and screens, before AI-behavior tickets became a real share of the backlog.
Short answer, apply it yourself
5. Think of a team or backlog you know that mixes clearly-scoped work with fuzzy, subjective-quality work. Name one ticket there that would likely show the same wide, unstable burn ratio if anyone measured it.
Show hint
Look for a ticket with no fixed, checkable finish line, the way "less generic" or "less repetitive" never quite is.
Show answer
Model answer: A content team's backlog mixing "add a comment count to posts" (a clean CRUD ticket) with "make the homepage recommendations feel less repetitive" (a tuning ticket with no fixed finish line and an unclear feasibility bar). The second would show the same wide spread if anyone timed it against its point estimate.
Fill in the blank, work the number
6. If AI-tagged tickets grew from 35 percent of sprint points to 60 percent, would you expect the blended commit-hit-rate to rise, fall, or stay the same, and roughly why?
Show hint
Think of the blended rate as a weighted average of the CRUD rate and the AI-tagged rate.
Show answer
Fall further. The blended rate is a weighted mix of each type's own rate. Shifting more points toward the unpredictable AI-tagged slice pulls the average down more, even if neither type's own ratio changed at all.
Before you close the answer
Why this works
Tests whether you can tell the difference between "the team got worse at estimating" and "one kind of work in the backlog doesn't fit the tool being used to measure it." Most candidates blame the team, or reach for more process, instead of recutting the data by type.
Follow-up traps
"Couldn't you just add a bigger buffer to every AI ticket's estimate instead of doing all this?" Response: a fixed buffer is still one number for a ratio that ranged 0.9x to 4.1x across six sprints. A buffer big enough for the worst case wastes most sprints, and one too small still blows the forecast.

"Isn't this just proof AI tickets should never share a board with everything else?" Response: no, they stay on the same board. Only the estimation and tracking method splits: CRUD keeps its points, AI-behavior tickets get a timebox and their own burn-ratio tracking.
If pressed
The burn ratio wasn't self-reported. It came from actual engineer hours logged against each ticket, divided by the sprint's own points-to-hours conversion, roughly 0.8 engineer days per point at Duskglass's historical average. The 2.3x figure for AI-tagged tickets was measured against logged time across all six sprints, not a team's gut sense that things felt slower.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more