ConceptIntermediateResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #3
How do you measure adoption of an internal AI tool?
LEAD the metric: the City of Hollow Ridge's Building and Permits Office, and FormGate, its permit pre-check assistant
Interviewer's question: "How do you measure adoption of an internal AI tool?" The City of Hollow Ridge issues building permits out of one municipal office. Teodora Salas has worked as a permit technician there for ten years.
The direct answer
Measure the voluntary reopen rate: how often a technician opens FormGate on a new application without anyone telling them to, in the weeks after any training or mandate has worn off. That number moves weeks before permit cycle time does, in either direction. Total logins, the number most teams reach for first, can look perfectly healthy for months while real use is already collapsing underneath it, because one mandated login counts exactly the same as a hundred genuine ones.
Do this, in order
Track voluntary reopens, not total logins, as the real adoption number.Why: a mandated first login and a genuine habit look identical in a login count, and only one of them means anything.
Watch it weekly starting at rollout, not after the first quarterly report.Why: the drop happens quietly for weeks before a lagging metric like backlog size ever moves.
Name how the number gets gamed before someone games it by accident.Why: a training mandate that requires one login per technician will always keep the raw login count looking fine.
Set a real threshold and a real action at each one.Why: a metric nobody acts on is a chart nobody needed to build.
Separate "no one complained" from "everyone's using it."Why: quiet abandonment never generates a support ticket, so silence is not proof the tool is working.
How to answer this, stage by stage
Nobody is grading whether you can name a metric. They're grading whether you can name the one that would have warned you before the dashboard did.
Stage 1
Scope it to one office, one tool
Say it like this
"I'll answer this for Hollow Ridge's permits office, and FormGate, the tool that pre-checks applications before a technician like Teodora reviews them."
Why this works
Keeps "measure adoption" from turning into a generic list of dashboard metrics.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the real outcome. Early signal, what moves first. Abuse, how it gets gamed. Decision, what I'd do at each threshold."
Why this works
Signals a method before naming a single number.
Stage 3
Name the real outcome
Say it like this
"The thing that actually matters is permit cycle time, how long from submission to issuance. Everything else is downstream of whether people genuinely use the tool."
Why this works
Anchors the whole answer in a business outcome, not the model's own score.
Stage 4
Give the early signal
Say it like this
"Not total logins. Voluntary reopens, a technician using FormGate on a new application weeks after any training mandate, with nobody telling them to."
Why this works
This is the direct answer: the metric that moves before the lagging outcome does.
Stage 5
Name how it gets gamed
Say it like this
"If a supervisor requires one login for training, the raw login count stays high forever, whether or not anyone actually keeps using it afterward."
Why this works
Shows you understand every metric has a way to be satisfied without the real behavior happening.
Stage 6
Give the thresholds, and close
Say it like this
"If voluntary reopen rate falls under thirty percent, I'm interviewing technicians that week, not waiting for the backlog to prove it. Above seventy, I'm expanding to the next office."
Why this works
Turns the metric into an action plan instead of a number on a slide.
Let's learn
FormGate is a tool that reads a building permit application and flags missing documents or zoning citations before a technician's own review.
Before it, Teodora and her colleagues flip through a zoning code binder and a paper checklist for every application, about twenty-five minutes each. With FormGate, that drops to about ten, since the flags are already sitting there when she opens the file.
Voluntary reopen rate, by week since FormGate's rollout
The fall started in week three, right after the training mandate ended. Nobody was watching this number, only total logins, which stayed flat the entire time.
The turn: the login count wasn't lying, exactly. It was just measuring something that had stopped meaning anything the moment the mandate ended and one login per technician became the whole story it could tell.
Same desk, same technician. What changed is how easy it became to stop noticing who was still really using it.
The decision that mattered
Hollow Ridge's launch dashboard tracked total FormGate logins as the adoption metric. That looked reasonable when the tool was new and everyone was genuinely curious. It stopped meaning anything the day a mandatory training login counted exactly the same as a technician choosing to open it on their own, six weeks later.
At its worst: the backlog of unreviewed permits grows for four straight months while the login dashboard says adoption is at ninety percent the whole time, because ninety percent of technicians did, in fact, log in once, back in week one.
All three of these were true at Hollow Ridge, at the same time, for months.
What I would leave alone: the mandatory training login itself is fine. Making sure every technician tries the tool once is a reasonable, low-cost requirement. The mistake isn't requiring it, it's counting it as proof of anything beyond that one day.
The lesson: a metric that can be satisfied by a policy instead of a behavior will always look calm right up until the thing it was supposed to warn you about finally shows up somewhere else.
Now here is the same thing as a story
The short version above is what you'd say defending this metric to Hollow Ridge's city manager. Read this one for how the drop actually got noticed.
Teodora Salas has processed building permits for ten years. She can tell a straightforward renovation permit from one that's going to need three follow-up calls before she's finished reading the cover letter.
When FormGate rolled out, every technician in the office was required to log in once during a half-day training. Teodora kept using it after that, on her own, because it genuinely caught things she used to miss on a rushed Friday. She assumed everyone else did too.
Knowledge spark: what makes a metric "leading" instead of "lagging"?
A leading metric moves before the thing you actually care about does. A lagging one only moves after the damage is already done. Backlog size is lagging. Voluntary reopen rate is leading, because people stop returning to a tool weeks before a backlog visibly grows.
The remark that started it came from a colleague, offhand, in the break room: "Have you noticed nobody actually reopens FormGate anymore? I think people just did the training and went back to the binder." Teodora hadn't noticed, because nobody's dashboard tracked that. It tracked logins, and logins looked fine.
The fourth step is invisible on a login count. It's the only one that actually mattered.
She started asking around, informally. Out of eleven technicians, three still opened FormGate regularly. The rest had gone back to the binder within a few weeks of training, mostly because the first time it flagged something they thought was wrong, nobody had told them what to do about a disagreement, so they just quietly stopped trusting it and went back to what they knew.
The login dashboard never lied. It just kept answering a question nobody was actually asking anymore: did everyone try it once. Nobody was asking who kept coming back.
Here's the decision I'd take back: building the adoption dashboard around total logins in the first place, because it was the easiest number to pull from the system. It made sense at launch, when a login really did mean someone was engaging with the tool. It stopped making sense the moment a mandate could produce the exact same number with none of the same meaning behind it.
Replayed with voluntary reopen rate tracked from day one: the same three-week dip after training shows up immediately, not four months later. The office's operations lead sees it drop from seventy to fifty-five percent in week three, pulls two technicians aside that same week, learns about the disagreement-handling gap, fixes it, and reopen rate climbs back to sixty-eight percent by week six instead of bottoming out at twelve.
I built the first dashboard around logins because it was the number the system already tracked, and asking for anything else felt like extra engineering for a "nice to have." It took a break-room remark and a four-month backlog to see that the extra engineering was the entire point of measuring adoption at all.
LEAD, mapped onto one login countNot a dashboard checklist. LEAD is what tells you which number would have warned you before the backlog did.
L
Link. The real business outcome.
Permit cycle time, how long from submission to issuance, not any score the model gives itself.
Anchors the whole metric question in something the city actually cares about.
E
Early signal. What moves first.
Voluntary reopen rate: a technician using FormGate on a new application, weeks after any mandate, with nobody telling them to.
The hardest step and the direct answer: the number that moves before cycle time ever does.
A
Abuse. How it gets gamed.
A mandatory training login keeps total logins looking healthy indefinitely, whether or not anyone keeps using the tool.
Names the exact way the obvious metric fails without anyone cheating on purpose.
D
Decision. What you'd do at each threshold.
Below thirty percent voluntary reopen rate, interview technicians that week. Above seventy, expand to the next office.
Turns the metric into an action, not a number nobody responds to.
Four conditions have to be true for a reopen to actually mean something. A mandated login meets none of them.
Three months separate the real drop from anyone noticing it. That gap is the whole cost of watching the wrong number.
The recap, one line per letter: link is permit cycle time, early signal is voluntary reopen rate, abuse is a training mandate inflating raw logins, and decision is a real threshold with a real action behind it.
Average permit cycle time, three scenarios
Watching logins alone still delivers some improvement. It leaves most of the value on the table, quietly, for months.
And if you want to be sure it really works, try it somewhere elseSame four letters, a county library's cataloging desk instead of a permits counter. Nothing else about the two jobs is alike.
Larchwood County Library System uses an assistant that suggests subject headings and call numbers for new titles before a cataloger finalizes the record. Idris Whitlock has catalogued books there for twelve years.
Mapped onto LEAD: link is how quickly new titles reach the shelves ready to browse, not the model's own suggestion-accuracy score. Early signal is the same shape as Hollow Ridge's: catalogers voluntarily accepting or lightly editing the tool's suggestion on a new title, weeks after the branch-wide training session ended, versus quietly re-typing the whole record from scratch. Abuse is structurally identical too: a mandated one-time training session keeps the raw usage count high forever, whatever catalogers actually do afterward. Decision: below a set voluntary-use threshold, a supervisor sits with catalogers to find out what specifically stopped working for them, rather than assuming the tool failed across the board.
A different shape of picture than Section 2 used: not a timeline, a sort. The mandated login sits exactly in the middle, telling you nothing by itself.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "measure voluntary reopens, not total logins, since a mandate can fake the second and never the first," and stop.
Cost: if tracking "voluntary" reopens is hard to instrument precisely, a rough proxy works too: reopens more than two weeks after any training event, tagged separately from week-one logins.
The model gets better, for real: even if FormGate's suggestions get more accurate over time, a rising voluntary reopen rate is still the number that tells you people trust it, not just that it's technically correct.
Where people run it wrong.
They report total usage as adoption, without separating a mandated first touch from a genuine habit.
They wait for a lagging outcome, like a backlog or a complaint, to confirm a problem that a leading metric would have shown weeks earlier.
They fix the tool the moment usage dips, without first asking technicians what specifically made them stop trusting it.
How to use it live. When someone asks how to measure adoption, ask yourself first: could this number stay high for months while real use has already collapsed? If yes, that's not your metric yet, keep looking for the one a policy can't fake.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "how do you measure adoption of an internal AI tool"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. The early signal here is voluntary reopen rate, not total logins.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Teodora Salas, a ten-year permit technician at Hollow Ridge, who noticed the adoption problem only after an offhand break-room remark.
3 · THE HABIT
What did most technicians stop doing a few weeks after training?
Tap to flip
ANSWER
They stopped voluntarily reopening FormGate on new applications and quietly went back to the paper zoning binder.
4 · THE FLIP
What's the two-setting switch in this story?
Tap to flip
ANSWER
Opening FormGate on their own, unprompted, versus going back to the binder for good. Eight of eleven technicians flipped within a few weeks, with no complaint filed.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building the adoption dashboard around total logins, the easiest number already sitting in the system, instead of a number a mandate couldn't fake.
6 · THE NUMBER
Fill in the blank: voluntary reopen rate fell from 70 percent in week 1 to ___ percent by week 12.
Tap to flip
ANSWER
12 percent. The drop had already started by week 3, three months before the backlog made the problem visible.
7 · THE REPLAY
Same three-week dip, new metric tracked from day one. What changes?
Tap to flip
ANSWER
The drop to 55 percent gets caught in week 3. Two technicians get interviewed that same week, the disagreement-handling gap gets fixed, and reopen rate climbs back to 68 percent by week 6.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the early signal there?
Tap to flip
ANSWER
Larchwood County Library System's cataloging assistant. Same early signal: voluntary acceptance of suggestions weeks after training, not the one-time training login count.
Check yourself Score: 0 / 0
Multiple choice
1. Why does total login count fail as an adoption metric for FormGate?
A. Logins are too hard to measure accurately.
B. A mandatory one-time training login counts exactly the same as a genuine, repeated habit.
C. FormGate doesn't actually log who uses it.
D. Technicians share logins with each other.
Show hint
Look at the "abuse" step.
Show answer
B. The metric fails because it can't tell a mandated, one-time touch from a genuine, repeated one. Both look like the same login.
True or false
2. True or false: the drop in real FormGate usage first became visible in the permit backlog, in month four.
True
False
Show hint
Look at the line chart and the timeline diagram.
Show answer
False. The drop actually started in week 3. It just wasn't visible in the backlog, the lagging metric, until month four.
Fill in the blank
3. Fill in the blank: with the mandate masking real abandonment, permit cycle time only fell to ___ days, versus 9 days with real adoption tracked and fixed.
Show hint
Look at the three-bar chart.
Show answer
15 days. Watching only the login count still delivered some improvement over the original 18 days, but it left most of the value on the table.
Short answer, where it wouldn't matter
4. Name a kind of permit task where FormGate's adoption barely matters either way.
Show hint
Look at the quadrant diagram's bottom-right corner.
Show answer
Model answer: A very routine, low-ambiguity renewal permit that a technician could process correctly either way, with or without FormGate's flags.
Short answer, apply it yourself
5. Pick a workplace tool you've seen adopted. What would tell you real adoption apart from everyone just completing a mandatory training?
Show hint
Think about whether people keep using it once nobody's checking anymore.
Show answer
Model answer: Whether people use the tool on a task where nobody's watching, weeks after the training ended, not whether they clicked through it once during onboarding.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Building the dashboard around total logins, the easiest number the system already tracked. It made sense before any mandate existed to distort it, and stopped making sense the moment one did.
Before you close the answer
Why this works
Tests whether you'll settle for the easiest available number or go looking for one a policy can't fake. Most candidates stop at "track usage."
Follow-up traps
"Isn't voluntary reopen rate just usage with extra steps?" Response: the extra step is exactly the point, it filters out the one thing that makes raw usage lie, a mandate that produces the same login without the same meaning.
"What if some technicians just don't need FormGate because they're already fast?" Response: fair, which is why the threshold triggers a conversation, not an automatic verdict; a low reopen rate is a question to ask, not a score to punish.
If pressed
Hollow Ridge's actual instrumentation tags a reopen as "voluntary" only if it happens more than fourteen days after that technician's last mandatory training event, which is a rough proxy, not a perfect one, but it's enough to separate a real habit from a policy-driven blip.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.