CalculationAdvancedQuality, Cost & Token Economics / Cost modeling and unit economics / #6

What is the cost impact of moving from a single call to a five-step agent?

Reiko budgeted the new research agent the way she had budgeted three smaller tools before it: multiply the single-call price by the step count. It did not hold, because two of the five steps grow with whatever the step before them found, and the deep report turned out good enough that people asked for it far more often too.

The direct answer
Never price a chained agent as the single-call cost times the step count. Each downstream step re-reads what earlier steps produced, search results, extracted notes, a draft, so a five-step agent is five different sized jobs added together, not one job repeated five times. For Rivalgraph, that honest build-up landed near twelve times the single-call cost per report, with a real range from about nine times to eighteen depending on how much raw text later steps carry forward. And because the deep version was trustworthy enough to actually use, requests for it climbed nearly eighteen times too, so the real monthly bill came in near two hundred and forty seven times the old one, not five.
Do this, in order
  1. Price each of the five steps by its own token count, never the single-call price times five.Why: the plan assumed a five times multiplier and the honest build-up came out near twelve times, because later steps read a growing pile of context, not a fixed one.
  2. Anchor the range to whether the team compacts context between steps, not to a single point number.Why: carrying raw search results all the way to the final draft instead of compact notes swings the report cost by about eight cents, the single biggest lever in the whole estimate.
  3. Forecast the request volume too, not only the per-report cost.Why: once the deep report earned enough trust to be requested before every client meeting, usage climbed almost eighteen times, and that alone moved the real bill more than the per-report multiplier did.
  4. Sanity check the total against what a person would have charged to do it by hand.Why: even the top of the range stays hundreds of times cheaper than a two and a half hour analyst brief, so the number is fine to spend, it just is not fine to plan for wrong.
  5. Add a verification pass that checks the draft's claims against the sources it actually names.Why: a step that has already thrown away the raw source text cannot tell a real claim from one it made up.
  6. Re-forecast the moment average tokens per report drift more than about fifteen percent from the last plan.Why: the real gap was only caught after three months of invoices, purely because nobody was watching the step by step numbers month to month.

How to answer this, stage by stage

Nobody is grading whether you can say the word "agent." They are grading whether you will price five unequal jobs honestly, or reach for a multiplier because it is the easy shortcut.

1
Scope it to one product before estimating anything in the abstract
Say it like this
"Let's ground this in one tool. Rivalgraph is a research agent Harrowfield Partners built for its own strategy team. It reads the open web and pulls together a competitor brief in minutes instead of an afternoon. Reiko Fujimori owns the spend forecast for it, and Marten Lepik is the engineer who built the five-step version."
Why this works
An abstract "what happens if you add steps" question turns into a hand wave fast. One product turns it into a real arithmetic problem.
2
Say your structure out loud before touching a single number
Say it like this
"I'm going to break the five steps into their own costs, own where every number comes from, give a range instead of one point figure, sanity check the total against what a person would charge, then say which assumption would move it most."
Why this works
Tells the interviewer you have a method before you have said a single dollar figure.
3
Break the equation down, and name the trap in the naive version
Say it like this
"The tempting shortcut is single-call cost times five steps. That's wrong, because steps two through five don't each do the same size job. Step two fetches raw pages off the web. Step five re-reads a chunk of that same raw text to check its own claims. The real equation is five separate costs, added, each built from its own token count."
Why this works
Naming the naive shortcut and why it fails is what shows you understand chained agents, not just the word "agent."
4
Own every number and where it came from
Say it like this
"I'll assume a single call costs about two cents, twelve hundred tokens in, nine hundred out. The five-step version runs about six tenths of a cent to plan, seven cents to search six sources, nine cents to extract facts from them, three and a half cents to compare across competitors, and five cents to draft the report and check two of its shakiest claims against the original page."
Why this works
A number nobody can trace back to a source is a guess wearing a decimal point.
5
Give the range, not one blended figure
Say it like this
"Add it up and a well built version runs about twenty four and a half cents a report, just over twelve times the single call. But it's a range. Compact the context hard and it can run as low as nineteen cents. Let the self check step retry, or let raw text pile up uncompacted across every step, and it climbs past thirty six cents, closer to eighteen times."
Why this works
A single point estimate hides exactly the design choice the interviewer is testing you on.
6
Run the sanity check and name the biggest lever together
Say it like this
"Compare that to what a strategy analyst would charge to do the same brief by hand, about two and a half hours, call it two hundred forty dollars loaded. Even the high end of our range is hundreds of times cheaper than that. And the assumption that moves the number most isn't retries or how many sources we fetch, it's whether later steps read compact notes or raw pages. That one choice is worth about eight cents on its own."
Why this works
A big number with nothing to compare it to is a guess with more decimal places. Naming the shakiest assumption is what a good estimator does that a bad one skips.
7
Close on the decision, and the volume it hides
Say it like this
"So: price the five steps separately, give a range anchored to context compaction, and don't stop at the per-report number. Once this thing was actually good, people asked for it constantly, and that's most of why the real bill was closer to two hundred fifty times the old one, not twelve."
Why this works
Ending on the decision and the volume effect, not the last cent computed, is what makes this sound like judgment instead of a spreadsheet read aloud.

Let's learn

Before Rivalgraph existed at all, a strategy analyst at Harrowfield Partners spent about two and a half hours building a three-competitor snapshot by hand, reading pricing pages and pulling notes together into a slide. Call it two hundred forty dollars of loaded time, once a quarter, when someone remembered to ask for it.

The first version of Rivalgraph made that nearly free to try. Type in three competitor names, get a one paragraph brief back in about three seconds, built entirely from whatever the model already knew from training. It cost about two cents a report, and nobody watched that line, because two cents times the handful of times a month someone remembered to use it was never worth its own row in a spending sheet.

Marten Lepik built a second version this spring: Rivalgraph Deep. Five steps instead of one. It plans what to check, searches and reads real pages, pulls out facts with the page they came from, compares across competitors, then drafts a report and double checks its two shakiest claims against the original text before it ships. Sourced, current, actually good enough that a strategy lead could hand it to a client without wincing.

Knowledge spark: what's context, in a chained agent? Everything the model gets shown before it answers: the original question, plus anything earlier steps produced. A single call has a small, fixed context every time. A chained agent's context can grow with every step, because each step needs to see what came before it.

The turn: Reiko's team had budgeted Rivalgraph Deep like any bigger version of a smaller tool: five steps, so five times the cost of the old single call, about twelve cents a report with a small buffer. That was never really the problem. The problem was that two of the five steps do not cost a fixed amount at all. They cost whatever the step before them handed them to read.

The five-step version was never five times one small job. It was five different sized jobs, and two of them got bigger every time the report ran.

Here's the arithmetic behind the two numbers that actually happened.

Hand sketched timeline titled one call versus five, dollar by dollar. A single call sits at the start at zero point zero two dollars baseline. Five milestones follow along the line: plan at zero point zero zero six dollars, search at zero point zero six nine dollars, extract at zero point zero eight seven dollars, compare at zero point zero three four dollars, draft and check at zero point zero four nine dollars. The line ends on a circled rust orange milestone marked five step total, zero point two four five dollars.
The five-step version is not one price repeated five times. It is five different sized jobs added along the line, and the total lands near twelve times the single call, not five.

The single call cost about two cents, twelve hundred tokens in and nine hundred out, priced at three dollars a million input tokens and fifteen dollars a million output, a fair stand in for a mid tier model that is good enough for a short answer.

The five-step version breaks into five separate jobs. Step one, plan what to check, costs about six tenths of a cent. Step two, search and read six real pages, is the first place the cost jumps, about seven cents, because the model has to actually read everything it pulled off the web. Step three, pull structured facts out of those same six pages with a citation for each one, costs about nine cents, the single most expensive step in the chain. Step four, compare the three competitors using just the compact facts, drops back down to about three and a half cents. Step five, draft the final report and check its two shakiest claims against the original page text, costs about five cents.

The build-up: what one deep report actually costs, step by step
$0.25 $0.125 0 1. Plan $0.006 2. Search $0.069 3. Extract $0.087 4. Compare $0.034 5. Draft+check $0.049 Total $0.245
PlanSearchExtractCompareDraft and check
Search and extract, the two steps that have to read real web pages, make up about two thirds of the whole report's cost. Plan and compare, the two steps that only read compact notes, barely register.
The choice that mattered Reiko's original forecasting rule, multiply the single-call price by the step count, was accurate for the three earlier tools Harrowfield had shipped before Rivalgraph Deep, because each of those tools' steps did the same small, fixed size job every time. Rivalgraph Deep was the first one where a step's job size depended on what the step before it found on the open web. The rule was never wrong on purpose. It just met its first exception.

At its worst, this is not really about twenty four cents being expensive. It is cheap, in absolute terms. At its worst, an agent that quietly runs three or four times its planned cost, unnoticed, month after month, is worse than a product honestly forecast to cost more from day one, because nobody is watching the number that is actually moving.

One thing worth flagging on its own, since a person's behavior is doing real damage here, not the token math. A report that used to cost basically nothing and took ten seconds got used a handful of times a month. A report that is actually good enough to trust in front of a client gets asked for constantly. Cheaper and better does not always mean the bill gets smaller. Sometimes it means people finally use the thing enough that a bill shows up for the first time.

Before Rivalgraph Deep, the whole strategy team asked for a competitor brief about thirty five times a month, seventy cents total, small enough that finance never gave it a line of its own. Once Deep shipped, and word got around that the sourced version was actually right, requests climbed to about six hundred forty a month within a quarter, strategy leads pulling one before every client kickoff and every quarterly review. At a blended twenty seven cents a report, once you count the runs where the check step has to retry, that is about a hundred seventy three dollars a month. Twelve times the per-report cost, times almost eighteen times the volume, and the real bill landed near two hundred forty seven times where it started.

What moves the report cost most, if the assumption is wrong
Raw pages carried forward, vs compact notes ~$0.08 Sources fetched per report, six vs four ~$0.05 Self check retry rate ~$0.02
Biggest swingMedium swingSmaller swing
Estimated dollars a report moved if each assumption breaks the wrong way. Whether later steps read raw pages or compact notes swings the report cost about four times as much as the next biggest lever.

What I'd leave alone: Rivalgraph still runs a small, single call lookup called RivalPulse for one thing only, has a competitor's public funding round or headcount changed since last week. It never touches the open web beyond one clean API call, its context never grows step to step, and the old five-times rule is still exactly right for it. Giving RivalPulse the same rebuild Deep needed would spend engineering time solving a problem it does not have.

The lesson: a forecasting rule that was right three times running is not proof it is a rule. It is proof it has not met the exception yet. The exception here was not a bigger step count. It was the first step whose job size depended on the size of whatever page it happened to find.

Now here is the same thing as a story

Read the long version below when you want to feel why a rule that worked three times running still broke on the fourth, not just be told that it did.

Reiko Fujimori has forecast what Harrowfield's internal tools cost to run for two years now. She built her first spend model the week the firm's first small AI helper shipped, a single call that reformatted a client's messy spreadsheet into a clean slide, and she has kept the same one page method ever since: price one call, multiply by however many calls or steps a tool needs, done.

It worked. Three tools later, on time every time, her forecasts landed within a few dollars of the real monthly invoice. The finance partners upstairs stopped asking her to double check her own numbers, which, for Reiko, was the whole point of the exercise.

Marten Lepik pitched Rivalgraph Deep in March. The old single-call version had always been fine for a rough first look, but strategy leads had started quietly not trusting its answers, because it only ever knew what the model had been trained on, sometimes a year out of date, sometimes flatly wrong about a competitor's current pricing. Marten's version would actually go read the web. Five steps: plan, search, extract, compare, draft and check.

Reiko ran her usual method. Single-call cost, about two cents, times five steps, call it ten cents a report, add a buffer, round to twelve. She sent the forecast up, it got approved without much conversation, because her numbers were always fine, and Deep shipped in April.

For the first few weeks it looked fine too. A handful of strategy leads tried it, the invoice line for that month barely moved, and Reiko had three other budgets to watch.

There was no one day this went wrong. Week by week, more of Harrowfield's strategy leads heard from a colleague that the new version actually cited its sources and was worth trusting before a client call, so more of them started asking for it before every kickoff instead of only when they remembered. Nobody announced a new habit. The invoice for the "AI tools" line just kept being a little bigger than the month before, and a little bigger than that, for three months running, and Reiko kept telling herself it would flatten out once the novelty wore off.

It did not flatten out. At the third month's close, going through the vendor invoice line by line instead of just checking the total against last month's, the way she usually did, she found the number for Rivalgraph alone: a hundred seventy three dollars, against a forecast of well under twenty.

The invoice was never wrong. The five-times rule was.

Reiko's first instinct was the sensible one: check for a billing error, a duplicate charge, a price change from the model vendor. There was none of that. So she pulled the actual per-step token logs Marten's team kept, something she had never once needed to open for the earlier, simpler tools, and did the arithmetic she should have done back in March. Five steps, five different sizes, not one size five times. Step three, pulling facts out of six real web pages with a citation for each, alone cost about nine cents, nearly as much as her entire twelve-cent forecast for the whole report.

It was never really about whether five times was a reasonable guess back in March. It was about nobody having the one job whose whole purpose was noticing the moment a rule that had worked for three simple tools met a fourth tool built differently underneath.

The decision that opened the door went back to that same one page method Reiko built the week the first tool shipped. It made complete sense then. Every tool Harrowfield had built up to that point ran through steps that each did a fixed amount of work: reformat a spreadsheet, summarize one document, answer one question. Multiplying was the honest shortcut for a chain of identically sized links. Nobody in the March meeting asked whether Rivalgraph Deep's links were actually the same size as each other, because every tool before it had made that a fair question not worth asking.

Run that third month's close again with one change: Reiko's team now tracks the average tokens per report for every multi-step tool, not just the total invoice, and re-forecasts the moment that average drifts more than about fifteen percent from the last plan. Rivalgraph Deep's average drifts past that mark in its second week, not its twelfth. The gap gets caught at forty dollars, not a hundred seventy, and the fix, a real compaction step so later stages read notes instead of raw pages, ships before the third invoice ever lands.

One design let a rule built for identical links keep pricing a chain whose links were not identical, for three full months, because nobody was watching the one number that would have said so. The other design watches that number every month, on every tool, whether or not the total invoice looks fine.

What I'd tell myself, back in that March meeting: a shortcut that has been right three times is not a rule yet. It is a streak. And the fourth tool was always going to be the one where a step's job depended on the size of whatever it found lying on the open web.

BOUND: pricing five unequal steps instead of one link repeated

Not a story question wearing a framework's clothes. This is an estimation problem, and BOUND is what keeps a comforting five-times shortcut from hiding two very different steps underneath it.

BBreak it down. What's the actual equation?
Cost per report equals the sum of five step costs, plan, search, extract, compare, draft and check, each its own input tokens times input price plus output tokens times output price. Not one call's price times five.
Say the equation before touching a number, or the shortcut that fits three tools quietly gets applied to a fourth one it does not fit.
OOwn the numbers. Where did each one come from?
Single call baseline: 1,200 tokens in, 900 out, at $3 and $15 per million. Five-step build: plan 600 in / 250 out; search six pages 17,650 in / 1,100 out; extract 18,600 in / 2,100 out; compare 3,350 in / 1,600 out; draft and check 9,300 in / 1,400 out.
This is also where the rejected alternative sits, see below: a hard token ceiling instead of a real compaction step.
UUse a range, not one number.
A well compacted build runs about $0.245 a report, just over twelve times the single call. Skip compaction and let raw pages ride all the way to the draft step, and it climbs past $0.33. Add the roughly one time in five the self check step has to retry, and the honest high end sits near $0.36, about eighteen times.
A single point figure this precise, from a chain this new, is exactly what let Reiko's forecast sound more solid than it was.
NNail the sanity check. Does the number survive being compared to something real?
Compare it to what Harrowfield would have paid a person: about two and a half hours of an analyst's time, call it $240 loaded, for the same three-competitor brief. Even the top of the range, $0.36, is about six hundred sixty times cheaper than that. And compare step three's cost alone, nine cents, against Reiko's entire twelve-cent forecast for the whole report, that gap alone should have been the tell in March.
The hardest step, and the one most answers skip. A number with nothing real to measure it against is a guess with more decimal places.
DDirection. Which assumption would move the answer most?
Not the retry rate, and not how many sources get fetched. Whether steps four and five read compact extracted notes or the full raw pages moves the report cost by about eight cents on its own, more than the other two assumptions combined.
Naming the assumption you trust least, out loud, is what a good estimator does that a bad one skips.

Three things worth stating directly, since this is where the real judgment sits. The alternative Marten's team tried first, and dropped, was a hard token ceiling on each step instead of a real compaction step, capping how much raw page text any step could read no matter what. It lost because a blunt ceiling cut real facts along with the excess, and in testing the agent missed a competitor's price change about eleven percent of the time simply because the fact it needed had fallen past the cutoff. The AI specific failure worth naming by name is a chain quietly compounding its own context, each step re-reading more than it needs because nobody built the step whose only job is throwing away what the next step does not need, and it costs real money without ever throwing an error. The guardrail is the extraction step itself, plus one hard rule: any claim in the final report has to point at the specific source line it came from, checked in the last step against the compact notes, not trusted on the model's word alone. That guardrail is not free, it is most of step three's nine cents, small next to the two hundred forty dollars the same brief would have cost by hand. And the bar Rivalgraph Deep holds itself to was never a fixed cost per report, no chain this new earns a fixed number yet. It is a range, rechecked whenever the average tokens per report drift more than fifteen percent, not one comforting decimal point standing in for five steps of very different size.

And if you want to be sure it really works, try it somewhere else

Same five letters, a customs desk at a shipping terminal instead of a strategy floor, and nothing about competitors anywhere in sight.

Quaybrief is an AI agent Wrenmoor Terminals built to classify shipping manifests, matching each line item to the right customs tariff code before a human broker signs off. Solange Beaumont runs digital operations for the terminal.

The build-up: a single-call classification, guessing the tariff code straight from the line item's description, costs about a cent and a half. The five-step version, parse the manifest, pull the current tariff schedule plus any past ruling for a similar item, match candidate codes, check the result against a restricted and sanctions list, then draft the classification with a confidence flag for the broker, runs about fourteen cents. A little over nine times, not the five times Wrenmoor's finance team assumed when they copied the multiplier straight off a vendor's slide.

The decision Solange would take back Budgeting Quaybrief's rollout with the same flat five-times rule a vendor's slide deck used for a much simpler pipeline, without asking how many past-ruling lookups a complicated, many-line manifest actually triggers.

That assumption held fine for a simple manifest, three or four line items, one obvious code each. It broke on a complex one, forty line items, several needing a lookup against a prior ruling because the item did not match any tariff code cleanly. Complex manifests were rare at first, then became most of what came through the terminal's new machinery-import contracts, and the average cost per manifest crept from fourteen cents toward twenty two without anyone changing a setting.

Same method, different lever: for Rivalgraph, the lever that moved the estimate most was whether later steps carried raw text or compact notes. For Quaybrief, it is not context compaction at all, every step there is already compact. It is how many prior-ruling lookups one manifest triggers, which depends on how unusual its cargo is, not on how the software was written.

A customs broker reviewing one complex manifest by hand takes about forty minutes, call it $47 loaded. Even at twenty two cents, Quaybrief runs close to two hundred times cheaper than that, the same shape of answer as Rivalgraph: cheap in absolute terms, just not the multiplier anyone had actually written down.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: price the steps separately, anchor the range to the assumption that actually drives cost, then sanity check the total against a person doing the same job.
Cost: there's no budget this quarter for both a bigger review sample and a compaction rebuild. The compaction rebuild wins, it's the one that stops the bill from drifting every month on its own, a bigger sample just measures the drift more precisely.
The model got better, for real: say Rivalgraph's underlying model gets a price cut next quarter. That's not proof the report gets cheaper by the same percentage, the search and extract steps are priced by how much of the web they have to read, not by the model's own price, so a cheaper model shrinks the bill less than it looks like it should.

Where people run it wrong.
They price a chained agent by multiplying the single-call cost by the step count, and never ask whether every step actually does the same size job.
They forecast the per-report cost and stop there, missing that a genuinely better tool gets asked for far more often, which moves the real bill more than the per-report number does.
They let raw material pile up in context step after step because it's easier to build than a real compaction step, and only notice the cost once the invoice is already three times the plan.

How to use it live. Say the real question out loud before quoting a number: "before I give you a multiplier, are the steps in this chain actually the same size job, or does one of them grow with what an earlier step found?" That buys a beat to think instead of repeating a shortcut that worked on the last three tools.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: show the arithmetic, own the assumptions. Built for estimation and sizing questions like this one, not a story about a person's habit.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Reiko Fujimori, the finance partner who forecasts what Harrowfield Partners' internal AI tools cost to run. Her one page pricing method had been accurate for two years before Rivalgraph Deep.
3 · THE OLD RULE
What rule did Reiko stop checking, because it had worked three times running?
Tap to flip
ANSWER
Multiply the single-call price by the step count. It was exactly right for three earlier tools whose steps all did the same fixed size job, and nobody re-tested it on a tool whose steps did not.
4 · THE EQUATION
What five costs make up Rivalgraph Deep's real per-report price?
Tap to flip
ANSWER
Plan, search and read six pages, extract facts with citations, compare across competitors, draft and check. Five different sized jobs added together, not one job repeated five times.
5 · THE OLD DECISION
What decision would Reiko take back?
Tap to flip
ANSWER
Approving Rivalgraph Deep's launch forecast using the same single-call-times-steps shortcut that had worked for three simpler tools, without asking whether this chain's steps were actually the same size as each other.
6 · THE NUMBER
Fill in the blank: the honest per-report cost ran from about $0.19 up to $___, with a best estimate near $___, against a single call cost of two cents.
Tap to flip
ANSWER
$0.36, and $0.245. Roughly nine to eighteen times the single call cost, not the five times the original forecast assumed.
7 · THE REPLAY
Same third month's close, new design, what changes?
Tap to flip
ANSWER
The team tracks average tokens per report monthly and re-forecasts the moment it drifts past fifteen percent. The real gap gets caught in week two at about $40, not month three at $170, and the compaction fix ships before the next invoice lands.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the different lever there?
Tap to flip
ANSWER
Quaybrief, a customs manifest classifier at Wrenmoor Terminals. There the multiplier is driven by how many prior-ruling lookups a complex manifest triggers, not by how much raw text later steps carry forward.

Check yourself Score: 0 / 0

Fill in the blank
1. The single-call version of Rivalgraph cost about $___ a report. The five-step version's best estimate came out to about $___ a report.
Show hint
Look at the build-up numbers in the timeline diagram and the chart in Section 1.
Show answer
$0.02, then $0.245. A little over twelve times the single call cost, not the five times the original forecast used.
Multiple choice
2. Why did Reiko's five-times forecast undercount Rivalgraph Deep's real cost so badly?
  • A. The model vendor raised prices partway through the quarter.
  • B. It assumed all five steps cost the same as one call, when two of them actually grow with how much raw web content earlier steps pulled in.
  • C. Harrowfield's strategy team stopped using the old single-call version entirely.
  • D. The finance system double counted every invoice that quarter.
Show hint
Look at which two steps in the build-up chart cost far more than the other three.
Show answer
B. Search and extract are the two expensive steps, because each has to read real web pages instead of doing a small, fixed size job like the other three steps do.
True or false
3. True or false: because the five-step agent cost about twelve times as much per report as the single call, Harrowfield's total monthly spend on Rivalgraph also grew by about twelve times.
  • True
  • False
Show hint
Check the paragraph in Section 1 that gives both the old and new monthly totals, not just the per-report cost.
Show answer
False. The per-report cost grew about twelve times, but requests for the report grew almost eighteen times on top of that, once the deep version was good enough to trust before every client meeting. The two multiply together, so the real monthly bill grew close to two hundred forty seven times, not twelve.
Short answer, name the rejected alternative
4. What alternative did Marten's team try first to control Rivalgraph Deep's cost, and why did it lose?
Show hint
Look at the O step in the framework recap, in the paragraph right after the five letters.
Show answer
Model answer: A hard token ceiling on each step, capping how much raw page text any step could read no matter what. It lost because the ceiling cut real facts along with the excess, and the agent missed a competitor's price change about eleven percent of the time in testing, simply because the fact it needed fell past the cutoff.
Short answer, apply it yourself
5. Pick a multi-step AI tool you use or have heard of. Name one step in it that probably costs more than the others because it has to read something the earlier steps produced, and how you would check.
Show hint
Think about which step in that tool has to read something long, a document, a page, a transcript, rather than just the original question.
Show answer
Model answer: A meeting-notes app that records, transcribes, then summarizes probably has one expensive step, the summarizer, because it has to read the entire transcript, not just a short prompt. I'd check by asking for the input token count on that specific step, not just the tool's total monthly bill, since a summarizer's cost scales with how long the meeting ran.
Multiple choice
6. If Rivalgraph Deep let raw page text ride uncompacted all the way to the final draft step, instead of extracted notes, what happens to the report cost?
  • A. It falls, because there is less work for the extraction step to do.
  • B. It stays about the same, since the model reads the same web pages either way.
  • C. It rises by roughly eight cents, since steps four and five each have to read far more tokens to find the same facts.
  • D. It becomes impossible to estimate without switching frameworks entirely.
Show hint
Look at the sensitivity chart in Section 1. Context compaction is the longest bar for a reason.
Show answer
C. Reading compact notes versus full raw pages is the single biggest lever in the whole estimate, worth about eight cents a report on its own, more than the retry rate and the source count combined.
Before you close the answer
Why this works
Tests whether you will price a chained agent step by step or default to a comforting multiplier, and whether you will forecast usage growth alongside the unit cost instead of stopping at the per-report number.
Follow-up traps
"Twenty four cents a report still sounds cheap, why does the multiplier even matter?" Response: it matters because the multiplier is what the forecast is built on. A twelve-times reality budgeted as five times means every volume projection built on top of it is wrong too, and volume is what actually moved this bill.

"Couldn't you just cap total tokens per report to control cost?" Response: tried and dropped, a flat ceiling cut real facts along with the excess and cost the agent about eleven percent of its accuracy on price changes in testing. The fix has to be a real compaction step, not a blunt cutoff.
If pressed
The retry rate itself is not fixed either. It is driven by how many claims the self check step flags as unsupported, which climbs when a competitor's website structure changes and the extraction step's usual pattern stops matching, worth watching as its own number, not folded into the general range.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more