What is the cost impact of moving from a single call to a five-step agent?
Reiko budgeted the new research agent the way she had budgeted three smaller tools before it: multiply the single-call price by the step count. It did not hold, because two of the five steps grow with whatever the step before them found, and the deep report turned out good enough that people asked for it far more often too.
- Price each of the five steps by its own token count, never the single-call price times five.Why: the plan assumed a five times multiplier and the honest build-up came out near twelve times, because later steps read a growing pile of context, not a fixed one.
- Anchor the range to whether the team compacts context between steps, not to a single point number.Why: carrying raw search results all the way to the final draft instead of compact notes swings the report cost by about eight cents, the single biggest lever in the whole estimate.
- Forecast the request volume too, not only the per-report cost.Why: once the deep report earned enough trust to be requested before every client meeting, usage climbed almost eighteen times, and that alone moved the real bill more than the per-report multiplier did.
- Sanity check the total against what a person would have charged to do it by hand.Why: even the top of the range stays hundreds of times cheaper than a two and a half hour analyst brief, so the number is fine to spend, it just is not fine to plan for wrong.
- Add a verification pass that checks the draft's claims against the sources it actually names.Why: a step that has already thrown away the raw source text cannot tell a real claim from one it made up.
- Re-forecast the moment average tokens per report drift more than about fifteen percent from the last plan.Why: the real gap was only caught after three months of invoices, purely because nobody was watching the step by step numbers month to month.
How to answer this, stage by stage
Nobody is grading whether you can say the word "agent." They are grading whether you will price five unequal jobs honestly, or reach for a multiplier because it is the easy shortcut.
Let's learn
Before Rivalgraph existed at all, a strategy analyst at Harrowfield Partners spent about two and a half hours building a three-competitor snapshot by hand, reading pricing pages and pulling notes together into a slide. Call it two hundred forty dollars of loaded time, once a quarter, when someone remembered to ask for it.
The first version of Rivalgraph made that nearly free to try. Type in three competitor names, get a one paragraph brief back in about three seconds, built entirely from whatever the model already knew from training. It cost about two cents a report, and nobody watched that line, because two cents times the handful of times a month someone remembered to use it was never worth its own row in a spending sheet.
Marten Lepik built a second version this spring: Rivalgraph Deep. Five steps instead of one. It plans what to check, searches and reads real pages, pulls out facts with the page they came from, compares across competitors, then drafts a report and double checks its two shakiest claims against the original text before it ships. Sourced, current, actually good enough that a strategy lead could hand it to a client without wincing.
The turn: Reiko's team had budgeted Rivalgraph Deep like any bigger version of a smaller tool: five steps, so five times the cost of the old single call, about twelve cents a report with a small buffer. That was never really the problem. The problem was that two of the five steps do not cost a fixed amount at all. They cost whatever the step before them handed them to read.
Here's the arithmetic behind the two numbers that actually happened.
The single call cost about two cents, twelve hundred tokens in and nine hundred out, priced at three dollars a million input tokens and fifteen dollars a million output, a fair stand in for a mid tier model that is good enough for a short answer.
The five-step version breaks into five separate jobs. Step one, plan what to check, costs about six tenths of a cent. Step two, search and read six real pages, is the first place the cost jumps, about seven cents, because the model has to actually read everything it pulled off the web. Step three, pull structured facts out of those same six pages with a citation for each one, costs about nine cents, the single most expensive step in the chain. Step four, compare the three competitors using just the compact facts, drops back down to about three and a half cents. Step five, draft the final report and check its two shakiest claims against the original page text, costs about five cents.
At its worst, this is not really about twenty four cents being expensive. It is cheap, in absolute terms. At its worst, an agent that quietly runs three or four times its planned cost, unnoticed, month after month, is worse than a product honestly forecast to cost more from day one, because nobody is watching the number that is actually moving.
One thing worth flagging on its own, since a person's behavior is doing real damage here, not the token math. A report that used to cost basically nothing and took ten seconds got used a handful of times a month. A report that is actually good enough to trust in front of a client gets asked for constantly. Cheaper and better does not always mean the bill gets smaller. Sometimes it means people finally use the thing enough that a bill shows up for the first time.
Before Rivalgraph Deep, the whole strategy team asked for a competitor brief about thirty five times a month, seventy cents total, small enough that finance never gave it a line of its own. Once Deep shipped, and word got around that the sourced version was actually right, requests climbed to about six hundred forty a month within a quarter, strategy leads pulling one before every client kickoff and every quarterly review. At a blended twenty seven cents a report, once you count the runs where the check step has to retry, that is about a hundred seventy three dollars a month. Twelve times the per-report cost, times almost eighteen times the volume, and the real bill landed near two hundred forty seven times where it started.
What I'd leave alone: Rivalgraph still runs a small, single call lookup called RivalPulse for one thing only, has a competitor's public funding round or headcount changed since last week. It never touches the open web beyond one clean API call, its context never grows step to step, and the old five-times rule is still exactly right for it. Giving RivalPulse the same rebuild Deep needed would spend engineering time solving a problem it does not have.
The lesson: a forecasting rule that was right three times running is not proof it is a rule. It is proof it has not met the exception yet. The exception here was not a bigger step count. It was the first step whose job size depended on the size of whatever page it happened to find.
Now here is the same thing as a story
Read the long version below when you want to feel why a rule that worked three times running still broke on the fourth, not just be told that it did.
Reiko Fujimori has forecast what Harrowfield's internal tools cost to run for two years now. She built her first spend model the week the firm's first small AI helper shipped, a single call that reformatted a client's messy spreadsheet into a clean slide, and she has kept the same one page method ever since: price one call, multiply by however many calls or steps a tool needs, done.
It worked. Three tools later, on time every time, her forecasts landed within a few dollars of the real monthly invoice. The finance partners upstairs stopped asking her to double check her own numbers, which, for Reiko, was the whole point of the exercise.
Marten Lepik pitched Rivalgraph Deep in March. The old single-call version had always been fine for a rough first look, but strategy leads had started quietly not trusting its answers, because it only ever knew what the model had been trained on, sometimes a year out of date, sometimes flatly wrong about a competitor's current pricing. Marten's version would actually go read the web. Five steps: plan, search, extract, compare, draft and check.
Reiko ran her usual method. Single-call cost, about two cents, times five steps, call it ten cents a report, add a buffer, round to twelve. She sent the forecast up, it got approved without much conversation, because her numbers were always fine, and Deep shipped in April.
For the first few weeks it looked fine too. A handful of strategy leads tried it, the invoice line for that month barely moved, and Reiko had three other budgets to watch.
There was no one day this went wrong. Week by week, more of Harrowfield's strategy leads heard from a colleague that the new version actually cited its sources and was worth trusting before a client call, so more of them started asking for it before every kickoff instead of only when they remembered. Nobody announced a new habit. The invoice for the "AI tools" line just kept being a little bigger than the month before, and a little bigger than that, for three months running, and Reiko kept telling herself it would flatten out once the novelty wore off.
It did not flatten out. At the third month's close, going through the vendor invoice line by line instead of just checking the total against last month's, the way she usually did, she found the number for Rivalgraph alone: a hundred seventy three dollars, against a forecast of well under twenty.
Reiko's first instinct was the sensible one: check for a billing error, a duplicate charge, a price change from the model vendor. There was none of that. So she pulled the actual per-step token logs Marten's team kept, something she had never once needed to open for the earlier, simpler tools, and did the arithmetic she should have done back in March. Five steps, five different sizes, not one size five times. Step three, pulling facts out of six real web pages with a citation for each, alone cost about nine cents, nearly as much as her entire twelve-cent forecast for the whole report.
It was never really about whether five times was a reasonable guess back in March. It was about nobody having the one job whose whole purpose was noticing the moment a rule that had worked for three simple tools met a fourth tool built differently underneath.
The decision that opened the door went back to that same one page method Reiko built the week the first tool shipped. It made complete sense then. Every tool Harrowfield had built up to that point ran through steps that each did a fixed amount of work: reformat a spreadsheet, summarize one document, answer one question. Multiplying was the honest shortcut for a chain of identically sized links. Nobody in the March meeting asked whether Rivalgraph Deep's links were actually the same size as each other, because every tool before it had made that a fair question not worth asking.
Run that third month's close again with one change: Reiko's team now tracks the average tokens per report for every multi-step tool, not just the total invoice, and re-forecasts the moment that average drifts more than about fifteen percent from the last plan. Rivalgraph Deep's average drifts past that mark in its second week, not its twelfth. The gap gets caught at forty dollars, not a hundred seventy, and the fix, a real compaction step so later stages read notes instead of raw pages, ships before the third invoice ever lands.
One design let a rule built for identical links keep pricing a chain whose links were not identical, for three full months, because nobody was watching the one number that would have said so. The other design watches that number every month, on every tool, whether or not the total invoice looks fine.
What I'd tell myself, back in that March meeting: a shortcut that has been right three times is not a rule yet. It is a streak. And the fourth tool was always going to be the one where a step's job depended on the size of whatever it found lying on the open web.
BOUND: pricing five unequal steps instead of one link repeated
Not a story question wearing a framework's clothes. This is an estimation problem, and BOUND is what keeps a comforting five-times shortcut from hiding two very different steps underneath it.
Three things worth stating directly, since this is where the real judgment sits. The alternative Marten's team tried first, and dropped, was a hard token ceiling on each step instead of a real compaction step, capping how much raw page text any step could read no matter what. It lost because a blunt ceiling cut real facts along with the excess, and in testing the agent missed a competitor's price change about eleven percent of the time simply because the fact it needed had fallen past the cutoff. The AI specific failure worth naming by name is a chain quietly compounding its own context, each step re-reading more than it needs because nobody built the step whose only job is throwing away what the next step does not need, and it costs real money without ever throwing an error. The guardrail is the extraction step itself, plus one hard rule: any claim in the final report has to point at the specific source line it came from, checked in the last step against the compact notes, not trusted on the model's word alone. That guardrail is not free, it is most of step three's nine cents, small next to the two hundred forty dollars the same brief would have cost by hand. And the bar Rivalgraph Deep holds itself to was never a fixed cost per report, no chain this new earns a fixed number yet. It is a range, rechecked whenever the average tokens per report drift more than fifteen percent, not one comforting decimal point standing in for five steps of very different size.
And if you want to be sure it really works, try it somewhere else
Same five letters, a customs desk at a shipping terminal instead of a strategy floor, and nothing about competitors anywhere in sight.
Quaybrief is an AI agent Wrenmoor Terminals built to classify shipping manifests, matching each line item to the right customs tariff code before a human broker signs off. Solange Beaumont runs digital operations for the terminal.
The build-up: a single-call classification, guessing the tariff code straight from the line item's description, costs about a cent and a half. The five-step version, parse the manifest, pull the current tariff schedule plus any past ruling for a similar item, match candidate codes, check the result against a restricted and sanctions list, then draft the classification with a confidence flag for the broker, runs about fourteen cents. A little over nine times, not the five times Wrenmoor's finance team assumed when they copied the multiplier straight off a vendor's slide.
That assumption held fine for a simple manifest, three or four line items, one obvious code each. It broke on a complex one, forty line items, several needing a lookup against a prior ruling because the item did not match any tariff code cleanly. Complex manifests were rare at first, then became most of what came through the terminal's new machinery-import contracts, and the average cost per manifest crept from fourteen cents toward twenty two without anyone changing a setting.
Same method, different lever: for Rivalgraph, the lever that moved the estimate most was whether later steps carried raw text or compact notes. For Quaybrief, it is not context compaction at all, every step there is already compact. It is how many prior-ruling lookups one manifest triggers, which depends on how unusual its cargo is, not on how the software was written.
A customs broker reviewing one complex manifest by hand takes about forty minutes, call it $47 loaded. Even at twenty two cents, Quaybrief runs close to two hundred times cheaper than that, the same shape of answer as Rivalgraph: cheap in absolute terms, just not the multiplier anyone had actually written down.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: price the steps separately, anchor the range to the assumption that actually drives cost, then sanity check the total against a person doing the same job.
Cost: there's no budget this quarter for both a bigger review sample and a compaction rebuild. The compaction rebuild wins, it's the one that stops the bill from drifting every month on its own, a bigger sample just measures the drift more precisely.
The model got better, for real: say Rivalgraph's underlying model gets a price cut next quarter. That's not proof the report gets cheaper by the same percentage, the search and extract steps are priced by how much of the web they have to read, not by the model's own price, so a cheaper model shrinks the bill less than it looks like it should.
Where people run it wrong.
They price a chained agent by multiplying the single-call cost by the step count, and never ask whether every step actually does the same size job.
They forecast the per-report cost and stop there, missing that a genuinely better tool gets asked for far more often, which moves the real bill more than the per-report number does.
They let raw material pile up in context step after step because it's easier to build than a real compaction step, and only notice the cost once the invoice is already three times the plan.
How to use it live. Say the real question out loud before quoting a number: "before I give you a multiplier, are the steps in this chain actually the same size job, or does one of them grow with what an earlier step found?" That buys a beat to think instead of repeating a shortcut that worked on the last three tools.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't you just cap total tokens per report to control cost?" Response: tried and dropped, a flat ceiling cut real facts along with the excess and cost the agent about eleven percent of its accuracy on price changes in testing. The fix has to be a real compaction step, not a blunt cutoff.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Cost modeling and unit economics
- #1 Build the cost-per-interaction model for a feature with a 2,000-token prompt and a 500-token response.
- #2 What cost drivers exist for an AI feature beyond model tokens?
- #3 Explain how a RAG pipeline's cost structure differs from a single model call.
- #4 How does prompt caching change your unit economics, and when does it not help?
- #5 Model the monthly cost of a feature used by 50,000 users averaging 12 interactions each.
- #7 Describe how you would find the most expensive one percent of your traffic.