ConceptAdvancedQuality, Cost & Token Economics / Measuring ROI and business impact / #23

What is the ROI of doing nothing, and when is it the right answer?

BOUND · sizing whether to speed up a phishing model's retraining cadence for a hospital network's security team

Driftnet reads and scores every email Corscaden Security's customers receive before it reaches a real inbox. Amadis Ecclestone owns its model roadmap. Crowmarsh Regional Health, a three-campus hospital network and one of Driftnet's nine highest-risk customers, had a near miss last month: an email written to sound exactly like its own CFO nearly moved $184,000 to a fake vendor. Tovah Sedgeworth, Corscaden's VP of Engineering, wants a real number before she funds a quarter of engineering time to fix it.

The direct answer
Build the continuous retraining pipeline now. Even the low end of the expected cost of doing nothing, about $194,000 a year in likely wire-fraud losses across Driftnet's nine highest-risk customers, clears the roughly $123,500 a year the build and its upkeep actually cost, with real room to spare. Doing nothing is only the right call if the true chance of a hit is under about 6 percent, meaning the one near miss that already happened turns out to be a fluke and not a sign of a real gap. That is the one number worth checking again before spending another quarter on anything else.
Do this, in order
  1. Build the continuous retraining pipeline this quarter.Why: even the low end of the incident-risk estimate clears the full cost of the build and its upkeep, with real margin.
  2. Do not try to fix this by lowering Driftnet's score threshold instead.Why: the near-miss email scored 12 out of 100, far under even the review line at 40; no reasonable cutoff change would have caught it.
  3. Track the false-negative rate on generative-style phishing as its own eval slice, separate from overall accuracy.Why: overall accuracy still looks fine; a slice like this one can rot for a year before anyone official notices.
  4. Put the churn risk on the same page as the incident risk, labeled as the softer number.Why: it is real money, but it is the one figure in the whole model without a receipt behind it yet.
  5. Re-check the 12 to 20 percent hit-probability assumption within two quarters.Why: it is the one number the entire call swings on; every other number in the model is well pinned down.
  6. Leave Driftnet's quarterly cadence alone for the lower-risk customers.Why: their transaction sizes and exposure do not clear the same bar, so spending engineering time there first solves the wrong company's problem.

How to answer this, stage by stage

Nobody is grading whether you can say "it depends." They are grading whether you can turn a mood, doing nothing feels risky, into an actual number, and tell them exactly what would have to be true for that number to flip.

1
Ground it in one real case, one real near miss
Say it like this
"Let's ground this in one case. Driftnet is Corscaden Security's phishing and spam detector, sold to enterprise security teams. Amadis Ecclestone owns its model roadmap. One of Driftnet's customers, Crowmarsh Regional Health, had a near miss last month that's forcing the question: is it worth speeding up how often the model gets retrained, or is the current setup good enough for now?"
Why this works
An abstract "how would you think about the ROI of doing nothing" answer turns into a slogan fast. One real near miss keeps every number after this checkable.
2
Say the method out loud before touching a number
Say it like this
"I'm going to run BOUND here, since this is really a sizing question, not a story question. Break the real comparison into its parts, own every number and say where it came from, give a range instead of one confident guess, sanity-check it against something real, and say which assumption would flip the answer."
Why this works
Naming the method up front tells the interviewer you have a structure, not five numbers arriving in whatever order they occur to you.
3
B: break the comparison into its real parts
Say it like this
"The real comparison is this. Cost of building a faster retraining pipeline for Driftnet, engineering time plus ongoing compute, against the cost of leaving the quarterly cadence alone: the chance of a real loss from a phishing email the model can't catch, plus a manual review cost that's already growing, plus the risk that a competitor closes this gap first and some customers leave."
Why this works
Saying the equation out loud, before any figure, is what stops the answer from turning into one scary anecdote dressed up as a number.
4
O: own each number, say where it came from
Say it like this
"Here's what I'd plug in, and why. Nine of Driftnet's thirty-four customers sit in the higher-risk segment, healthcare and financial services, where a wire-fraud attempt is worth real money. I'd assume a 12 to 20 percent chance one of them has a real, successful hit in the next year if nothing changes, based on the fact that one already surfaced as a near miss in six months, purely because an accounts-payable lead happened to be paying attention. I'd cost each hit at 180 to 260 thousand dollars, based on the real 184 thousand dollar transfer that almost went out. Building the faster pipeline runs about 127 thousand dollars once, and about 60 thousand dollars a year after that to keep running."
Why this works
Every number gets a reason attached to it, which is what lets someone push back on the assumption instead of the arithmetic.
5
U: turn it into a range, not a single confident figure
Say it like this
"So the expected cost of doing nothing lands somewhere between about 194 thousand and 468 thousand dollars a year, across those nine accounts. Building and running the pipeline costs about 123 thousand dollars a year, once you spread the one-time build over two years. Even the low end of that range clears the cost of building it."
Why this works
A single number here would claim a precision nobody actually has. The range is what an honest estimate sounds like out loud.
6
N: sanity-check it against something already real
Say it like this
"Does that pass a smell test? The low end, about 194 thousand dollars a year across nine accounts, is close to just one more event like the one that already almost happened. I'm not asking anyone to believe something bigger than the real thing that nearly went out the door last month."
Why this works
A sanity check against a real, already-proven number is what separates an estimate from a guess with a confident voice.
7
D: name the one assumption that would flip it
Say it like this
"If I'm wrong about one thing here, it's the hit probability, not the cost per incident and not the build cost, both of those are pinned down pretty well. If the real chance of a hit is closer to 3 to 5 percent instead of 12 to 20, meaning the near miss was a fluke rather than a sign of a real gap, the expected cost drops under what the build costs to run, and doing nothing is actually the right call, at least until a second real data point shows up."
Why this works
Naming the swing assumption is what a strong estimator does that a weak one skips, and it's exactly what "when is doing nothing the right answer" is asking for.
8
Close on the call, and its own kill condition
Say it like this
"So: build the faster retraining pipeline now, because even the conservative case clears its own cost with real room to spare. The one thing I'd want to re-check before committing further is that 12 to 20 percent, because if it turns out closer to 5, doing nothing was the right call all along, and I'd rather find that out in a quarter than after we've spent the budget."
Why this works
Closing on the decision and its own kill condition, in one breath, is what actually answers "when is doing nothing the right answer," not just "what should we build."

Let's learn

Driftnet reads and scores every email that reaches a customer's inbox, in under a second, and decides whether it is safe, worth a second look, or dangerous enough to block outright.

Hand sketched left to right flow diagram titled How Driftnet scores one email. Four rounded boxes connected by wobbly arrows: Rules layer, ML classifier, Score 0 to 100, this box outlined in rust to mark the step everything else depends on, Block review or deliver.
A rules layer catches known bad senders. A learned model reads the rest for language and context, and turns it into one score.
Knowledge spark: what makes an email "generative-style" phishing? One written by a language model instead of a template. No typos, no generic urgency, no reused kit. It can be built from a target's own public writing, an earnings call, a LinkedIn post, so it reads like them. The old tells a model learns to spot just are not there anymore.

Two years ago, before Driftnet, Crowmarsh's own security team hand-reviewed anything a basic spam filter flagged as unsure, which meant slower delivery for real mail and a small team triaging a growing pile of maybes by feel. Driftnet now auto-blocks about 9,400 emails a month across Crowmarsh's 260,000 external emails, and routes the ones it is not sure about to a security-team review queue. A year ago that queue held about 480 emails a month. It holds about 640 now.

Emails a month landing in Driftnet's review queue at Crowmarsh, last 12 months
700 350 0 near miss caught here 480 640 -12mo -9mo -6mo -3mo now
Emails a month in the review queueThe month of the near miss
480 grew to 640 a month, about a third more, before the near miss ever happened. The climb was sitting in the data for months. Nobody called it a trend until it nearly cost real money.

Here is the turn. The extra emails in the review queue were never the real problem, since a person still reads every one of those. The real danger is the email that never reaches the queue at all, because Driftnet scored it low enough to just deliver.

The extra work in the queue was never the real cost. The real cost was the email that scored low enough to walk straight past the queue and into someone's inbox.
Hand sketched gauge diagram titled The gap a threshold change can't close. A dial in the center labeled Driftnet's score, needle pointing low. Four labeled callouts around it: classic phishing scores 70 to 95, auto block line at 80, SOC review line at 40, the near miss email scored 12.
A classic phishing email scores 70 to 95. The near-miss email, written to sound like Crowmarsh's own CFO, scored 12, well under even the review line at 40.

At its worst, this costs a real transfer. The email impersonating Crowmarsh's CFO, Adaora Thornhollow, asked Crowmarsh's accounts-payable lead to rush $184,000 to a "new vendor" ahead of a system migration. It was written using Adaora's own earnings-call transcript and a LinkedIn post, so it read exactly like her. Driftnet's dashboard never flagged it. Nobody at Corscaden even knew the model was blind to this pattern until a person happened to notice, about an hour before the payment would have gone out.

The decision that mattered Driftnet's retraining cadence was set to quarterly at launch, a default nobody has revisited since. That was fine when phishing kits and templates barely changed for months at a stretch. It stopped being fine once attackers started writing with a language model instead of a kit, since that changes what "normal phishing" looks like far faster than a calendar can keep up with.

What I would leave alone: Driftnet's twenty-five lower-risk customers, small businesses with no large wire-transfer workflow. Even a fully successful attack there is unlikely to cost enough to justify moving engineering time away from the higher-risk segment first.

The lesson: a schedule that was right when the world it was watching held still stops being a schedule and starts being a blind spot the moment that world starts moving on its own.

Now here is the same thing as a story

Read the short version above when you are in the room. Read this one when you want to feel why an hour, not a headline, was the whole margin Crowmarsh actually had.

Philippa Duskin has run accounts payable at Crowmarsh Regional Health for eleven years. She knows the rhythm of a real vendor request the way a nurse knows a normal heartbeat. A real one has a purchase order number attached. A real one never asks her to skip the second signature. She has caught three fake invoices by feel in the last two years alone, before Driftnet ever flagged a single one of them.

Hand sketched timeline titled The AP lead's three mornings. Three labeled points on a line: Before Driftnet, checks by hand. The routine, trusts the review band. The near miss, this point circled in rust, score 12, her gut said no.
Before Driftnet, she checked everything by hand. Once it arrived, she leaned on the review queue to catch what needed a second look. The near miss never reached the queue at all.

When Driftnet arrived, it took the worst part of her week off her desk. Instead of squinting at every external request, she trusted the ones Driftnet delivered clean and gave a proper second look to the handful it routed to review. For a year and a half, that trade held. The routine got smaller. She stopped assuming every unusual request was hiding something, because the model had been right often enough to earn that.

Then, on a Tuesday afternoon, a request landed from "Adaora Thornhollow," Crowmarsh's CFO. It asked her to push $184,000 to a new vendor before a system migration locked the payment queue for the week. No typos. No generic urgency. It even referenced a real detail from the migration timeline. Driftnet had already delivered it clean, score of 12, nothing to review.

Philippa read it twice. Something about the rhythm was wrong, not the words themselves, more like a sentence Adaora would never actually build that way. She called Adaora's office directly instead of replying to the email. Adaora had sent nothing. The payment queue closed forty minutes later. The transfer never went anywhere.

Hand sketched two panel comparison titled Two people one spreadsheet. Left panel, a person icon labeled Amadis, caption has one real near miss no clean number yet. Right panel, a gauge icon labeled Tovah, caption needs a real number before funding a quarter of engineering.
Amadis had a real story and no clean number. Tovah needed a number before she would hand over a quarter of engineering time.

The story reached Amadis Ecclestone two days later, secondhand, through Crowmarsh's account team. Amadis brought it straight to Tovah Sedgeworth, Corscaden's VP of Engineering, expecting the near miss alone to be enough. It was not. Tovah had heard near misses before. What she wanted was a number she could actually defend to her own leadership, not a story about one very alert accounts-payable lead.

It was never really about the twelve out of a hundred. It was about the fact that nobody at Corscaden knew the model could be that wrong until someone got lucky.

The decision that opened this gap traced back to a planning meeting almost two years earlier, the week Driftnet first shipped. Someone asked how often the model should retrain. Quarterly was the answer, because back then phishing kits barely changed for months at a time, and quarterly retraining kept the model current without burning engineering time nobody had spare. Nobody wrote down a condition for revisiting that choice. It was just the setting, and settings that work quietly stop getting looked at.

Run that same near miss again, with the faster pipeline already live. This time Driftnet compares the email's writing style against ninety days of Adaora's real messages, not just its surface content, and the mismatch alongside an unusual payment request pushes the score to 61, well past the review line at 40. It never reaches Philippa's inbox marked "clean." Crowmarsh's on-call security analyst opens it, checks it, and kills it in twelve minutes. Not an hour of one person's gut standing between Crowmarsh and $184,000. Twelve minutes, by design.

What Amadis would tell that planning-meeting version of themselves: quarterly retraining was not a careless call. It was the right one for a world where attackers were still reusing kits. It was never built to survive attackers switching from a kit to a writing tool, and nobody ever asked it to.

BOUND, or the price tag on doing nothing

Not a way to dress up "I have a bad feeling about this" in five letters. BOUND is what forces a real number out of a feeling, and forces you to say exactly what would make that number wrong.

BBreak it down.What's the real equation?
Cost of building Driftnet's faster retraining pipeline, engineering time plus ongoing compute, against the cost of leaving the quarterly cadence alone: the chance of a real loss like Crowmarsh's near miss, a manual review cost that keeps climbing, and the risk that Mistport, a rival vendor already marketing writing-style detection, wins the next renewal on the strength of it.
Say the comparison out loud before a single figure appears, or the whole answer collapses into one scary anecdote.
OOwn the numbers.Where did each one come from?
Nine of Driftnet's thirty-four customers sit in the higher-risk segment. A 12 to 20 percent chance one has a real hit this year, based on one near miss surfacing in six months purely by luck. $180,000 to $260,000 per hit, based on the real $184,000 transfer. $127,000 to build the pipeline, $60,000 a year to run it, both priced off real engineering rates.
Every figure needs a reason attached, or there is nothing for anyone to actually push back on.
UUse a range.What's honest, low to high?
$194,400 to $468,000 a year in expected incident cost, against $123,500 a year to build and run the fix. This is the step a rushed answer collapses into one confident-sounding figure, and the one that actually reveals how much is genuinely known versus assumed.
A single number here claims a precision nobody has. The range says exactly how sure you are, out loud, instead of hiding it.
Hand sketched number line titled The range and what it has to clear. Four points left to right: Breakeven at 123,500 dollars a year, this point circled in rust, Near miss at 184,000 dollars, Low estimate at 194,400 dollars a year, High estimate at 468,000 dollars a year.
Even the low estimate clears the breakeven line. The near miss, the one real number in the whole model, sits close to it too.
NNail the sanity check.Does this survive a smell test?
The low end of the range, $194,400 a year, is close to the cost of just one more event like the $184,000 near miss that already almost happened. Nobody is being asked to believe something bigger than the real thing that nearly went out the door last month.
A number checked against a real event beats a number that just sounds plausible in the room.
DDirection.Which assumption swings it most?
The 12 to 20 percent hit probability, not the cost per incident and not the build cost, both of which are well pinned down. If the real rate is closer to 5 percent, expected cost drops to roughly $99,000 a year, under the $123,500 breakeven, and doing nothing becomes the right call, at least until a second real event shows up.
If the ranking of what to build would look the same regardless of this number, it was decided by gut and the number was written afterward to match.
What doing nothing costs against what building it costs, per year
$500k $250k $0 $123,500/yr to build and run it $479k high case $205k low case Doing nothing $123.5k/yr Building it
Manual review, already growingIncident risk, low caseIncident risk, extra to high caseOngoing costBuild, amortized over 2 years
Even the low case for doing nothing sits above the dashed line, what building and running the fix costs. The gap only closes if the true hit rate is far below what one near miss already suggests.

Three things worth stating directly, since this is where the real judgement sits. The alternative Corscaden considered, and rejected, was cheaper and faster: just lower Driftnet's score threshold so more borderline mail lands in review. It does not work, because the near-miss email scored 12, far under even the 40-point review line; tightening a cutoff catches classic phishing harder, it does nothing for an email that never looked like classic phishing in the first place, and pushed aggressively enough it would delay real mail, referrals, lab results, vendor invoices, into a queue nobody asked to grow. The AI-specific failure worth naming by name is a form of drift: the model's learned features, typos, generic urgency, reused phrasing, no longer discriminate once attackers write with a language model instead of a kit, so accuracy looks fine overall while a whole slice of real risk goes quietly unscored. The guardrail is a held-out eval set built only from generative-style phishing, tracked separately from overall accuracy, with a standing rule that a false-negative rate crossing 25 percent on that slice forces an out-of-cycle retrain, calendar or no calendar. And the trade-off being accepted plainly: the new pipeline will flag more legitimate mail into review while it tunes, a real cost in a hospital's inbox, where a late referral is not nothing. Corscaden is accepting that friction, and the $60,000 a year it costs to run, in exchange for closing a gap a threshold change was never going to close.

And if you want to be sure it really works, try it somewhere else

Same five letters, a grain co-op instead of a hospital network, and this time the thing evolving mid-cycle wasn't an attacker with a language model, it was a fungus.

Blightscope is Tarnworth Cooperative's own crop-disease detector, built in-house by its agronomy-data team to flag early blight and mildew across 2,200 acres and 46 grower fields from drone and scout photos. It retrains once a season, a default set three years ago when the strains it watches for barely changed within a five-month growing window.

The decision Tarnworth would take back Blightscope's once-a-season retraining made sense when a strain that showed up in April still looked the same in August. This year, a wetter spring pushed a faster-mutating fungal strain into the fields, and its early symptoms looked different enough from what Blightscope learned at season start that the model's confidence on it drifted down over about ten weeks, quietly, with nobody watching that slice on its own.

A field scout caught a spreading patch in Field 14 by eye, days before Blightscope's own score would have flagged it at its current trend. The catch saved an estimated $46,000 in crop value in that field alone. Coenraad Farrowfield, who owns Blightscope, rebuilds the case the same way Amadis did: somewhere between one and three patches a season are likely going undetected past the point where they can still be saved, at $35,000 to $50,000 each, putting $35,000 to $150,000 a season on the table. Moving to biweekly retraining during the growing season, using scout-confirmed photos, costs about $18,000 once and $9,000 a season to run. Even the low end clears that with room to spare, and the one real save already on the books, $46,000 in a single field, is bigger than the entire season's running cost on its own.

Hand sketched two panel comparison titled Blightscope retrained once a season vs mid season. Left panel, a gauge icon labeled Season start only, caption trained once blind to a strain that shows up 10 weeks in. Right panel, a document icon labeled Mid season retraining, caption updates every two weeks from scout confirmed photos.
Same shape of problem as Driftnet's. A model trained once against a target that keeps changing shape underneath it.

Same rank, different swing assumption: validating against a real save comes first here too, but the number that would flip the call is different. At Crowmarsh, it was a hit probability across nine accounts. At Tarnworth, Isambeau Vosswinkel, the co-op's general manager, would only reverse the call if it turns out the new strain never spreads past Field 14, since a single contained patch would not have justified the spend, and the co-op will not actually know that until the season is already over, which is itself the argument for not waiting to find out.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the range: the low case for doing nothing already clears the cost of building, doing nothing is only right if the hit rate is under about 6 percent.
Cost: there's no budget this quarter for the full build. Instrument the eval slice first, cheap and immediate, and treat the probability estimate as provisional until a second real data point lands.
The model got better, for real: say Driftnet's classifier already halves its false-negative rate on generative-style phishing on its own. The case still holds, since the breakeven math does not change, only the assumed probability shrinks, and it is still worth checking whether the new number lands above or below 6 percent.

Where people run it wrong.
They treat "no incident yet" as proof nothing is wrong, instead of checking whether the model's confidence has quietly drifted underneath.
They try to fix a feature problem by turning a threshold dial, and either still miss the same emails or drown real mail in false positives.
They build the range once and never re-check the one assumption it actually swings on, so an estimate calcifies into a fact nobody re-tests.

How to use it live. Ask yourself, out loud if you have to: "which single number in this estimate, if wrong, changes the answer, and how would I find out?" Naming the swing assumption is what shows real judgement instead of a rehearsed calculation.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: break a comparison into its real parts, own each number, give a range, sanity-check it, and name what would flip it. Built for estimation and sizing questions, including the ROI of doing nothing.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Amadis Ecclestone, who owns Driftnet's model roadmap at Corscaden Security, a phishing and spam detector sold to enterprise security teams.
3 · THE EQUATION
What's the real comparison being sized here?
Tap to flip
ANSWER
Cost of building a faster retraining pipeline for Driftnet against the cost of leaving the quarterly cadence alone: incident risk, a growing manual review cost, and the risk of losing customers to a rival vendor.
4 · THE GAP
What did Driftnet's score say about the near-miss email, and what should it have said?
Tap to flip
ANSWER
It scored 12 out of 100, low risk. A classic phishing email with the same intent usually scores 70 to 95. The gap is why a threshold change alone could not have fixed this.
5 · THE OLD DECISION
What decision would Amadis take back?
Tap to flip
ANSWER
Setting Driftnet's retraining cadence to quarterly by default at launch. It made sense when phishing kits and templates changed slowly, not once attackers started writing with a language tool instead of a kit.
6 · THE NUMBER
Fill in the blank: the expected cost of doing nothing lands between $___ and $___ a year, against a breakeven of about ___ percent.
Tap to flip
ANSWER
$194,400 to $468,000 a year, against a breakeven of about 6 percent.
7 · THE REPLAY
Same near miss, the faster pipeline already running, what changes?
Tap to flip
ANSWER
The email's style mismatch against Adaora's real writing pushes its score to 61, over the review line. Crowmarsh's on-call analyst catches it by design in about 12 minutes, not by one person's gut an hour before the payment would have gone out.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which product, and what's the different swing assumption?
Tap to flip
ANSWER
Blightscope, Tarnworth Cooperative's in-house crop-disease detector. The swing assumption there is how far the new fungal strain actually spreads before season's end, not a hit probability.

Check yourself Score: 0 / 0

Multiple choice
1. Which single assumption would flip whether building the retraining pipeline is worth it?
  • A. The cost per incident, $180,000 to $260,000.
  • B. The build cost, about $127,000.
  • C. The chance a higher-risk customer has a real hit this year.
  • D. The manual review cost, about $44,500 a year.
Show hint
Look at the D step, direction, in the framework recap.
Show answer
C. The cost per incident and the build cost are both well pinned down. The hit probability, 12 to 20 percent, is the number the whole answer swings on.
True or false
2. True or false: lowering Driftnet's auto-block threshold would have caught the near-miss email.
  • True
  • False
Show hint
Check the paragraph after the framework recap about the rejected alternative.
Show answer
False. The email scored 12 out of 100, far under even the 40-point review line. A stricter threshold catches classic phishing harder, it does nothing for an email that never looked like classic phishing.
Fill in the blank
3. The near-miss email asked Crowmarsh's accounts-payable lead to rush $___ to a new vendor, and Driftnet scored it ___ out of 100.
Show hint
Look at the labeled gauge diagram in "Let's learn," or flashcard 4.
Show answer
$184,000, scored 12 out of 100. That $184,000 is also the one real number the whole sanity check leans on.
Short answer, name the old decision
4. What old decision would Amadis take back, and why did it make sense when it was made?
Show hint
Look at the key point box titled "The decision that mattered," in "Let's learn."
Show answer
Model answer: Setting Driftnet's retraining cadence to quarterly at launch. It made sense at the time because phishing kits and templates barely changed for months, so quarterly retraining kept the model current without spending engineering time nobody had spare.
Short answer, apply it yourself
5. Think of a tool you use that runs on a fixed update schedule, weekly, monthly, quarterly. What would have to change in the world around it before that fixed schedule became the wrong default?
Show hint
Think about what the schedule was quietly assuming would stay stable.
Show answer
Model answer: A store's monthly reorder list assumes demand for each item moves slowly. If a product suddenly goes viral on social media, monthly reordering means running out for weeks before the schedule catches up, the same kind of gap Driftnet had against a faster-moving attacker.
Short answer, work the arithmetic
6. If the true hit probability turned out to be 5 percent instead of 12 to 20 percent, would building the pipeline still be worth it? Show the arithmetic.
Show hint
Use 9 accounts, $220,000 average cost per hit, and the $123,500 a year breakeven.
Show answer
No, not on the numbers alone. 9 accounts times 5 percent times $220,000 is about $99,000 a year, under the $123,500 a year the build and its upkeep cost. At that rate, doing nothing would be the right call, at least until more evidence shows up.
Before you close the answer
Why this works
Tests whether you can turn "doing nothing feels risky" into an actual number, and whether you understand that an AI feature's status-quo cost is a moving target, not a fixed one, because the thing it is failing to catch keeps changing shape too.
Follow-up traps
"Isn't the 12 to 20 percent probability just made up?" Response: it is anchored to the one real near miss that already surfaced within six months across a nine-account segment, not free-floating. The honest move is to treat it as provisional and re-check it, not to pretend a range beats a guess automatically.

"Why not just buy Mistport's tool instead of building your own?" Response: worth asking, but Mistport's writing-style detection is still unproven in the market, and Driftnet already has two years of SOC-confirmed phishing labels a bought tool would start without. The churn-risk number is what keeps that option honest instead of dismissed outright.
If pressed
The new feature compares an email's writing style against a rolling 90-day baseline of that sender's real messages, not a fixed rulebook, so it adapts as a person's own writing naturally shifts. It only fires alongside a request pattern, a payment, credentials, an urgent approval, never on style mismatch alone, to keep the false-positive rate from spiking just because someone writes differently on a Monday.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more