ConceptAdvancedAI Opportunity & Model Strategy / Opportunity identification for AI / #11
How do you find opportunities where AI changes the unit economics rather than just the experience?
LEAD · why a flat satisfaction score can hide a real cost collapse, tested on Vintermark's pricing engine Tallyrook
Tallyrook is Vintermark's pricing engine. It watches what competitors charge across the web and hands an online seller a new price back, all day, without anyone opening a spreadsheet. Yrsa Aabakken owns its roadmap. Oxcombe Outdoor Supply, a fourteen thousand SKU camping and gear retailer, is one of its biggest clients. For most of a year, the number everyone watched on Tallyrook was how satisfied Oxcombe's sellers felt. It barely moved. Underneath that flat number, something much bigger already had.
The direct answer
Map the real workflow behind an AI pricing idea and price every step in it: minutes, times a loaded dollar rate, times how often it repeats. Chase the opportunity that collapses the single most expensive step on that map, not the one that only makes a screen feel nicer. If the cost of one transaction does not move once you run that math, you found an experience fix, not a unit economics one, no matter how good the demo looks.
Do this, in order
Map the workflow and price every step in minutes and dollars before picking an idea.Why: this is the direct answer, laid out as an action.
Go after the single most expensive manual step first.Why: that is the one place a real cost collapse can actually live.
Check cost per transaction before and after, every time.Why: a feature that only trims a click usually leaves that number exactly where it was.
Track cost per unit weekly, on its own, not folded into a satisfaction score.Why: satisfaction can sit still for months while the real economics story already happened.
Call out a fake unit economics pitch the moment it reuses an old cost number instead of a fresh one.Why: the label gets borrowed long after the real win was already spent.
Once a step collapses, point the next roadmap slot at the next expensive manual step, not a nicer version of the one you just fixed.Why: the win keeps compounding only if you keep chasing cost, not comfort.
How to answer this, stage by stage
Nobody is grading whether you found a bigger number. They are grading whether you would have gone looking for it at all.
1
Ground it in one real pricing budget
Say it like this
"Let's ground this. Vintermark builds Tallyrook, a pricing engine e-commerce sellers plug into their catalog. Yrsa Aabakken owns the roadmap. One client, Oxcombe Outdoor Supply, runs fourteen thousand SKUs. I'll answer with their pricing team, not with AI opportunities as an idea floating in the air."
Why this works
Naming a real budget and a real headcount stops the answer from staying a slogan about "finding value."
2
Name the method in one breath
Say it like this
"I'll run this as LEAD. Link ties 'changes the economics' to a real cost per transaction, not a feeling. Early signal is the number that would have shown this before anyone noticed. Abuse is how a team fakes an economics win. Decision is what actually changes on the roadmap because of it."
Why this works
Two seconds of structure tells the interviewer this is a method, not a lucky story picked after the fact.
3
Reframe what "changes the economics" actually means
Say it like this
"Sellers get two different kinds of pricing feature. One kind makes the tool feel less annoying and does not touch what it costs to run. The other kind takes a step a person used to do by hand and lets a model do it for pennies. Only the second kind is a unit economics opportunity, and the question is really asking how you tell those two apart before you build either one."
Why this works
This is the reframe. Without it, the rest of the answer is just a list of features with no test attached.
4
Give the one decision, before any story
Say it like this
"Short version: map the workflow, price each step in minutes and dollars, find the one step costing the most, and go after that with the model. If cost per transaction does not move, you found an experience fix, not an economics one."
Why this works
The interviewer never has to wait for the story to find out what you would actually do.
5
Prove it with the near miss, compressed
Say it like this
"Here's what almost got missed at Oxcombe. For a year we watched satisfaction on the price suggestion screen. It sat at four point one out of five and never moved. Then someone ran the actual cost per repricing check, manual against automated, and found Oxcombe's whole pricing team had gone from a hundred eighty thousand dollars a year down to about sixty one thousand, while covering seventeen times more products. That story had already been true for weeks. The satisfaction score never once said so."
Why this works
This is the real content of the early signal step, the actual comparison the interviewer is testing for.
6
Catch the way this gets faked
Say it like this
"Watch for this: someone takes a feature that trims a few seconds off a click, like a one tap price approval, and calls it a unit economics story because it sounds like the last one. Check the actual cost per transaction. If it was four cents before and it is still four cents after, nothing about the economics moved. Only the click count did."
Why this works
Naming the fake version out loud separates a candidate who understands the test from one who just liked the last answer's shape.
7
Say what to chase next, and what to leave alone
Say it like this
"Once one step collapses, look for the next step still done by hand, not a nicer version of the one you just fixed. At Oxcombe that is about six hundred bundle SKUs still priced manually because the model cannot match a kit's parts yet. That is next. The one tap approval screen is fine to ship. It is just not the win."
Why this works
This is the discipline check: does the candidate know when a good idea is not the point.
8
Close on the decision, restated in one breath
Say it like this
"So here's what I'd actually do. Before any roadmap idea gets built, I price the workflow it touches. I chase the step that costs the most, not the one that would feel the nicest. And I track cost per transaction every week, because a satisfaction score can sit still for months while the real story already changed underneath it."
Why this works
This is the actual decision, said once, plainly, tying the whole answer back to deliverable zero.
Let's learn
Here's what happens when a feature makes everyone a little happier and the actual bill barely changes size.
Say a company builds Tallyrook, a tool that checks what competitors charge and hands a seller a new price back, without anyone opening a spreadsheet.
Four steps in one seller's pricing workflow. Only one of them was ever expensive.
Before Tallyrook, Oxcombe Outdoor Supply ran this whole workflow by hand. Three pricing analysts, a hundred eighty thousand dollars a year in loaded pay, checking competitor prices at roughly three dollars a check: six minutes of work at about thirty dollars an hour. That budget could only cover about eight hundred of the fourteen thousand SKUs on a real weekly cycle, the ones that sold enough to be worth an analyst's morning. The other thirteen thousand two hundred sat on prices nobody had rechecked in a quarter.
Same catalog. One version of it only ever got checked for one SKU in seventeen.
Now, an automated check costs about four cents: a scrape, a match against the right competitor listing, and a price recommendation, with no analyst involved unless the model flags a shaky match. Every one of the fourteen thousand SKUs gets checked, most of them daily. Oxcombe's whole pricing operation, one analyst left plus the automation bill, now runs about sixty one thousand dollars a year. Full coverage, at a third of the old cost.
Cost per SKU reprice check, Oxcombe, weekly, as automated coverage rolled out
The line was already flat near four cents by week twelve. Nobody checked it against the budget until finance needed a number for a renewal slide.
What is unit economics?
The real cost, or the real profit, on one single unit: one SKU repriced, one order shipped, one claim processed. Not the whole company's revenue. Just what one transaction actually costs to do, tracked on its own.
Satisfaction sat at four point one for a year and told us nothing was happening. Cost per repricing check had already fallen ninety eight percent underneath it.
What it costs at its worst: if nobody had run that math, Vintermark would have kept pitching Tallyrook on "sellers like it a bit more," a soft story that is easy to cut from a roadmap. The real fact, a ninety eight percent drop in the cost of running a whole catalog, would have sat undiscovered, and it was the actual reason a client would sign a longer contract.
The choice I would take back
Setting customer satisfaction as Tallyrook's headline success metric at launch, because it was the number every product dashboard already tracked, and it was fine when the product's whole job was feeling less annoying. It stopped being fine once automation started collapsing real costs behind that same screen, because a satisfaction score was never built to catch a cost line falling that far.
What I would leave alone: the price suggestion screen itself, the part sellers actually rate, is fine exactly as it is. Nobody would notice if it never got another design pass, because it was never carrying the real value. Leave the screen alone. The win never lived there.
The lesson: a number can hold perfectly still and still be hiding something enormous underneath it. A satisfaction score measures how a screen feels. It was never built to catch a cost line falling ninety eight percent somewhere nobody's dashboard was pointed at. If a metric only watches the interface, it will miss the moment the real economics change, every time, and nobody will know until someone happens to run the math.
Now here is the same thing as a story
Use this version when you have the time. The short version above is what you would actually say out loud in the room. This one is for feeling why it matters.
Every Monday, Yrsa Aabakken pulls the same three numbers before she opens her inbox, out of habit more than need. She has run Tallyrook's roadmap at Vintermark for four years, since before the model could tell a hiking boot from a hydration pack, and she built the very first competitor matching pass herself, on a spreadsheet, checking forty SKUs by hand just to prove the idea would work at all.
Tallyrook launched properly eighteen months ago, and for the first stretch, the number Yrsa watched was the one everyone watches: how satisfied a seller felt with the price suggestions Tallyrook handed them. It opened at three point six out of five. Sellers said the screen was nice. Nobody said much more than that.
Oxcombe Outdoor Supply came on as a client that spring. Fourteen thousand SKUs, three pricing analysts, and a habit going back years of manually checking whatever competitor prices there was time for. Within two months of turning Tallyrook on, satisfaction crept up to four point one. It sat there. Week after week, four point one, sometimes four point zero, once briefly four point two before sliding back. Steady. Unremarkable. The kind of number a roadmap meeting glances at and moves past.
Nobody was hiding anything. The number just never moved enough to make anyone stop and ask what was actually happening underneath it.
Then, in October, Vintermark's VP of Growth, Bjarke Hemsley, brought that same flat four point one into a board prep deck, arguing the pricing engine had plateaued and the roadmap should shift toward something friendlier to use, starting with a one tap price approval button. His deck had eleven months of satisfaction data behind it. It looked responsible. It looked like evidence.
One of these was already true. Nobody had asked it the right question yet.
Yrsa almost let it go. The number really had sat still. Nothing in her own dashboard gave her a reason to ask a different question.
What stopped it was smaller than a meeting. A finance partner, prepping Oxcombe's renewal call that same week, asked Yrsa for one line: how much had Tallyrook actually saved this client, in dollars, this year. Not a satisfaction score. A dollar figure, for a slide.
Yrsa didn't have one ready. She built it in an afternoon, pulling Oxcombe's own pricing operations line from the shared finance sheet, before Tallyrook and after.
Before: three analysts, a hundred eighty thousand dollars a year, and real coverage on maybe eight hundred of the fourteen thousand SKUs, the ones that sold enough to be worth checking. The other thirteen thousand two hundred sat on prices nobody had rechecked in a quarter.
After: one analyst left, handling only the cases the model flagged as uncertain, plus the automation bill itself. Sixty one thousand dollars a year, total, covering every single SKU, most of them rechecked daily.
We did not make Oxcombe's sellers a little happier. We cut what it costs them to price their whole catalog by two thirds and gave them coverage they had never had, and the satisfaction score never once said that story was even happening.
Yrsa remembers the exact meeting where satisfaction became the metric on Tallyrook's own dashboard. It made sense at the time. Tallyrook was young, mostly a nicer screen wrapped around a rough model, and satisfaction was the number every other product review already used, sitting right there on the same dashboard as five other features. Nobody sat down and asked whether a screen feeling number could ever catch a cost line collapsing behind it. It just came with the template.
She brought the real number to Bjarke before the board deck shipped. Cost per repricing check, three dollars down to four cents. Ninety eight percent, not a rounding error. Oxcombe's own renewal, which had been shaky, closed for three years within the month, on the strength of that one dollar figure, not the satisfaction score that had sat flat the whole time.
The one tap approval button still shipped, a few months later. It is a fine button. It just is not the reason anyone renewed.
What I'd tell myself, back when the dashboard template asked for a satisfaction number and nothing else: the screen was never lying. It just was never built to see the ledger underneath it. I would have gone looking for the dollar number a year earlier, not because the satisfaction score was wrong, but because I never should have trusted it to tell the whole story on its own.
LEAD, so a cost story doesn't hide behind a satisfaction score
Not a way to prove a feature is popular. LEAD forces you to name the real cost line before anyone gets to call something a unit economics win.
LLink. What "changes the economics" actually has to connect to.
Not a satisfaction score, and not how a feature demos. One real number: cost per transaction, or the margin sitting on top of it. A feature earns the label only if that number actually moves.
Every pricing feature Tallyrook shipped could score four out of five in a survey. What actually told anyone whether the economics moved was whether one repricing check got cheaper to run.
Four questions, multiplied together. That is the whole tool.
EEarly signal. The number that moved before anyone noticed.
Cost per repricing check, tracked weekly, on its own. It fell from three dollars to four cents over twelve weeks while satisfaction sat between four point zero and four point two the entire time. The leading number was already telling the real story. Nobody had a dashboard pointed at it.
By week eight, cost per check was already down to twenty two cents. The satisfaction score did not tell anyone anything new for another month after that, and even then, barely.
Oxcombe's yearly pricing operations spend, before and after Tallyrook
BeforeAfter
The bar dropped by two thirds while coverage went from 800 SKUs to all 14,000. The leading number, cost per check, had already predicted this weeks before anyone ran the total.
The cost line finished falling weeks before the meeting that finally looked at it.
AAbuse. How "changed the economics" gets claimed loosely.
The one tap approval button is the clean example. Old flow: three screens, about forty seconds, to approve one suggested price change. New: one tap, about twenty two seconds. Eighteen seconds saved, and cost per repricing check does not move at all, still four cents, because the model was already generating the recommendation before either version shipped. Calling that a unit economics story borrows the label from a real win it never earned. The team also considered making price suggestion acceptance rate the headline metric instead of cost per check, since it is cheap to pull from existing logs. It was rejected: a team can lift acceptance just by suggesting cautious, easy to approve prices, without the underlying cost ever moving, which is exactly the kind of gaming a real economics metric has to resist.
Bjarke's version, said out loud: "sellers like the new button, so this counts too." Liking a button was never the test.
Two features can sit at opposite corners of this chart and still show up in the same roadmap deck.
DDecision. What actually changes on the roadmap.
Every new idea gets a cost map before a design review: minutes per step, dollar rate, how often it repeats, who still does it by hand. Ideas that attack the priciest step get built. Ideas that only polish a screen get parked. Full daily coverage of all fourteen thousand SKUs would cost more in scraping calls than it is worth on slow moving tail items, so checks run tiered: high velocity SKUs daily, mid tier every three days, long tail weekly, trading a little price freshness on slow SKUs for a lot less compute spend. And because the matching model can grab the wrong competitor listing, a low confidence match never reprices on its own; it gets routed to a person first, the same way a threshold decides who checks a shaky output anywhere else in the product.
The next roadmap slot goes to Oxcombe's six hundred bundle SKUs, still priced by hand because the model cannot yet match a kit's own parts. Not to a nicer version of the button that already shipped.
Three questions, asked before a single design screen gets drawn.
The recap, one line per letter: link "changes the economics" to a real cost per transaction, not a satisfaction score. The early signal is cost per unit, tracked weekly, and it moved weeks before anyone's dashboard caught it. Name the abuse plainly: a feature that trims a few seconds off an already cheap step, wearing the language of a real cost collapse, and a metric like acceptance rate that can be hit without the underlying cost ever moving. And the decision is what makes any of this real: price the workflow first, build against the priciest step, and route anything the model is unsure about to a person rather than trust a round confidence number.
And if you want to be sure it really works, try it somewhere else
Same four letters, a freight brokerage instead of a webstore, and this time the metric that almost fooled everyone was speed, not satisfaction.
Ingimar Kestholm runs pricing operations for Redcastle Freight Partners, a mid size freight brokerage. Loadline, their AI quoting tool, used to get judged the same way Tallyrook almost did.
Before Loadline, a broker manually checked four or five competing carrier rates to quote one shipping lane: about twelve minutes, at a loaded rate near forty five dollars an hour, roughly nine dollars a quote. Only the big repeat customers' lanes got that real check. Smaller, one off shipments got a flat rate guess, because nobody had time to check them properly, and that guess quietly bled margin on lanes nobody was watching.
Loadline's automated lane check costs about fifteen cents. Every quote now gets a real competitive check in seconds instead of a guess. The team's headline metric was quote turnaround time, which fell from two hours to about four minutes, and sales loved it. That number almost became the whole story.
The turnaround number was true. It just was not the number that mattered.
The decision Ingimar would take back
Quote turnaround time sat on Loadline's own dashboard as the main success number, because it was the easiest thing to measure from a timestamp and it was genuinely impressive. It never once pointed at the long tail lanes that used to get a flat rate guess. A finance partner ran a lane by lane margin check instead and found about six hundred and twenty thousand dollars a year in margin that had been quietly leaking, and was now being recovered, a story turnaround time never suggested on its own.
Mapped onto LEAD, the shape holds. The link is a real cost per quote and the margin sitting behind it, not how fast the quote appears on screen. The early signal is the same lesson Yrsa learned: a leading cost number can already be telling the real story while the metric everyone is watching stays flat, or in this case, looks great for the wrong reason. The abuse Ingimar watched for was the same too: an "instant quote PDF download" button got pitched as an economics win, but cost per quote never moved, since the quote was already computer generated; only the delivery screen changed. And the decision matched Yrsa's: price the workflow, chase the step that actually costs the most, and keep the win moving to the next expensive manual step instead of resting on it.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: map the workflow's cost, chase the priciest step, check whether the transaction cost actually moved.
Cost: no budget for a dedicated finance review. Whoever owns the model spot checks a random sample of quotes or reprices by hand every week instead of waiting on a quarterly report. Slower to build, but it closes the same gap.
The model got better, for real: say the matching model's accuracy genuinely improves. Keep chasing cost per transaction anyway, because the point was never how right the model looked. It was always what the manual step used to cost, and whether that cost actually fell.
Where people run it wrong.
They let a feature borrow the language of the last real economics win because it feels similar, without checking the actual cost number.
They watch a satisfaction or turnaround metric and assume flat, or good, means nothing else is worth checking.
They chase the most visible manual step instead of the most expensive one, and those are not always the same step.
How to use it live. When an interviewer asks you to find an AI opportunity, buy yourself a second by asking out loud what the workflow actually costs today, per unit, before naming a single feature. That question does most of the answer already.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to find where AI changes unit economics instead of just the experience?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It forces "changes the economics" to mean a real cost per transaction number, not a feeling about the screen.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Yrsa Aabakken, who owns the roadmap for Tallyrook, Vintermark's AI pricing engine, and Bjarke Hemsley, Vintermark's VP of Growth, who nearly shipped the wrong roadmap call off a flat satisfaction score.
3 · THE HABIT
What number was the team watching instead of the real cost story?
Tap to flip
ANSWER
Customer satisfaction on the price suggestion screen. It sat between 4.0 and 4.2 for most of a year and never moved, while the real cost story was already happening underneath it.
4 · THE REAL DIFFERENCE
What separates a real unit economics opportunity from a feature that only feels nicer?
Tap to flip
ANSWER
Whether cost per transaction actually moves. Automated competitor matching collapsed the cost of a step from $3.00 to $0.04. The one tap approval button saved eighteen seconds of clicking and left the $0.04 cost exactly where it was.
5 · THE OLD DECISION
What decision would Yrsa take back?
Tap to flip
ANSWER
Setting customer satisfaction as Tallyrook's headline success metric at launch. It made sense when the product was mostly a nicer screen. It stopped working once automation started collapsing real costs behind that screen.
6 · THE NUMBER
Fill in the blank: Oxcombe's cost per repricing check fell from $___ manual to $___ automated, while total pricing operations spend dropped from $180,000 to about $___ a year.
Tap to flip
ANSWER
$3.00 to $0.04, about a 98 percent drop. Total spend fell to about $61,000 a year, while coverage went from 800 SKUs to all 14,000.
7 · THE REPLAY
Same flat satisfaction score, new lens. What changes?
Tap to flip
ANSWER
Instead of reading 4.1 as "nothing happened," Yrsa pulls the actual cost per check number. It shows the 98 percent drop weeks before a quarterly review would have caught it, and it is what closes Oxcombe's three year renewal.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the parallel find?
Tap to flip
ANSWER
Loadline, Redcastle Freight Partners' AI quoting tool, run by Ingimar Kestholm. There, quote turnaround time looked like the win, two hours down to four minutes, but the real story was about $620,000 a year in margin recovered on long tail lanes that used to get a flat rate guess.
Check yourself Score: 0 / 0
Fill in the blank
1. Before Tallyrook, Oxcombe's pricing team really covered only ___ of its 14,000 SKUs on a weekly cycle. After, ___ of them got checked, most of them daily.
Show hint
Check the coverage numbers in "Let's learn," right after the workflow diagram.
Show answer
800 SKUs before; all 14,000 after. That's about seventeen times more coverage, on a smaller total budget.
True or false
2. True or false: since the one tap approval button never moved cost per transaction, Vintermark should never have built it.
True
False
Show hint
Check the Decision step. Shipping an experience fix is fine, it just is not the same thing as an economics win.
Show answer
False. The button is a fine experience improvement, and it was fine to ship it. The mistake was only ever calling it a unit economics story when the real cost per transaction never moved.
Multiple choice
3. Why did automating competitor price research count as a real unit economics opportunity at Oxcombe, while the one tap approval button did not?
A. The one tap button used an older version of the pricing model.
B. Automating price research collapsed the cost of the step itself, while the approval button only trimmed a few seconds off a step that was already cheap.
C. Sellers rated the one tap button lower in surveys.
D. The competitor price research feature shipped first, so it earned the credit.
Show hint
Check the Link and Abuse steps in the framework recap.
Show answer
B. Cost per repricing check fell from $3.00 to $0.04. The approval button left that same $0.04 number exactly where it was, no matter how popular it got.
Short answer, name the reversal
4. What old decision would Yrsa take back, and why did it make sense the first time nobody thought to change it?
Show hint
Look at the block labeled "The choice I would take back" in Let's learn.
Show answer
Model answer: Setting customer satisfaction as Tallyrook's headline metric at launch. It made sense when the product was mostly a nicer screen with a rough model behind it, and satisfaction was the number every other product review already tracked. It stopped making sense once automation started collapsing real costs that satisfaction was never built to see.
Short answer, apply it yourself
5. Pick a workflow a tool you use today actually touches. What is the single most expensive manual step behind it, and would automating that step change what it costs, or just how it feels?
Show hint
Think about who used to do the step by hand, how long it took, and whether a model could do it for pennies instead.
Show answer
Model answer: A photo storage app's manual customer support review of a "wrongly flagged" account is far more expensive per case than its automated tagging of vacation photos. Automating the support review would change what it costs to run the service. A nicer photo grid would only change how it feels to scroll through it.
Short answer, work the number
6. If Oxcombe's catalog doubled to 28,000 SKUs, would cost per repricing check likely double too? Why or why not?
Show hint
Think about what the four cent automated cost is actually made of, and what the three dollar manual cost was actually made of.
Show answer
Not necessarily. The manual cost scaled with analyst headcount, so doubling SKUs really did mean hiring more people. The automated cost is mostly compute and data calls, which can scale with volume without needing more people, so cost per check likely stays close to four cents, with the real new constraint being compute budget and how many low confidence matches need a person, not linear analyst hours.
Before you close the answer
Why this works
Tests whether you go looking for a real cost number instead of resting on a satisfaction score, and whether you can tell a genuine unit economics change from an experience feature wearing the same label.
Follow-up traps
"Isn't ignoring satisfaction reckless? Doesn't it matter too?" Response: satisfaction isn't wrong, it's just a lagging, low resolution signal for this specific question. Track it alongside cost per transaction, never instead of it.
"Couldn't a seller be unhappy with an objectively cheaper system?" Response: possible, and worth checking, but that's a different problem, a rollout or trust issue, not whether the economics moved. Those get answered with two different numbers, not one.
If pressed
The matching confidence cutoff that decides which SKUs skip straight to auto reprice, versus getting flagged for a person, sits at 92 percent, tuned against a hand checked sample of 500 SKU matches, not picked as a round number. A wrong competitor match, like pricing against a discontinued tent model, is the kind of mistake that cutoff exists to catch before a price ever changes.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.