ConceptIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #7

What signals in user research suggest an AI solution rather than a better interface?

TRACE · a hard task mistaken for a hard-to-find button, tested on Kernelwell's recipe planner

Kernelwell recommends recipes and plans a family's week of meals around their allergies, their budget, and their macros. Ionara Vradenburg owns the part of Kernelwell that decides what to recommend and why. This is the quarter she learned that one of her loudest complaints was never about the screen at all.

The direct answer
Two signals matter: people describing the task itself as hard, not just hard to find, and people who already built their own spreadsheet to do that hard part by hand. Test it by asking whether the hard part would still be there if navigation were perfect. If yes, that's a signal for AI, not a redesign.
Do this, in order
  1. Test every "hard to use" note against one question: would it survive a perfect navigation fix?Why: this is the actual signal test. Everything else here just feeds it good evidence.
  2. Read the real words behind each complaint, not the tag a queue gave it.Why: one tag, "hard to use," can hide two totally different problems wearing the same label.
  3. Treat "the task itself is hard" as a judgment signal, not a UI signal.Why: checking three lists by hand is a different problem from not finding a button, and no menu redesign touches it.
  4. Treat a household's own spreadsheet or workaround as proof, not just a complaint.Why: they've already written the spec for the AI feature themselves, by hand, for free.
  5. Fix the real navigation cases with a plain interface change, nothing fancier.Why: building a model for a findability problem wastes effort and the button is still hard to find.
  6. Ground any AI recommendation in a real ingredient database before it ships, never the model's own claim alone.Why: a hallucinated "dairy free" label is worse than the manual check it was supposed to replace.

How to answer this, stage by stage

Nobody is grading whether you know recipe apps. They're grading whether you can tell a hard-to-find button apart from a hard-to-do task, using what people actually said.

1
Scope it to one product, one household
Say it like this
"Let me ground this in one real case. Kernelwell recommends recipes and plans a family's week of meals around their allergies, their budget, and their macros. I want to talk about one household that kept telling us the app was hard to use, and what we almost built for them instead of what they actually needed."
Why this works
One real household in one real kitchen keeps the answer from turning into a lecture about user research in general.
2
Say your structure out loud
Say it like this
"I'd run this as TRACE. Timeline: when the complaints actually started, and what we shipped that looked like a fix. Recut: slice the complaints by what people actually said, not the ticket tag. Assume nothing: don't call every hard thing a navigation problem, and don't call every hard thing an AI problem either. Cause candidates: three honest guesses. Evidence test: the one question that tells the two apart."
Why this works
Two seconds of structure tells the interviewer this is a method, not a guess dressed up as an answer.
3
Reframe what's actually being asked
Say it like this
"The real question isn't whether people are unhappy. It's whether 'this is hard' means 'I can't find it' or 'I can't do it.' Those read identical in a support queue, and they need completely different fixes."
Why this works
Naming the reframe up front shows the interviewer you won't jump straight to a redesign, or straight to a model, before checking which one it actually is.
4
Give the decision, committed
Say it like this
"So here's what I'd actually do. I'd pull the raw words behind every 'hard to use' tag this quarter, not the tag itself, and sort them into two piles: people describing a control they couldn't find, and people describing a judgment they had to make by hand. Whatever's genuinely in the second pile is where I'd look for an AI feature. The first pile gets a plain interface fix, nothing cleverer."
Why this works
This is the direct answer, said plainly, before a single number shows up.
5
Prove it with the evidence
Say it like this
"Here's what we actually found. A hundred and eighteen pieces of feedback got tagged 'hard to use' this quarter. Fifty one really were about a control people couldn't find. Sixty seven were about a task that was still hard once you'd found it, like checking a recipe against two kids' allergy lists and your own macros by hand. Twenty three households described that exact check unprompted, and nine of them had already built their own spreadsheet to do it."
Why this works
Real numbers pulled from real quotes are what make "this is a judgment task" checkable instead of a hunch.
6
Say what you'd leave alone, then close
Say it like this
"One thing worth saying plainly. I would not touch the fifty one navigation cases with a model. Someone who can't find the allergy filter needs a clearer menu, not software guessing at what they meant. So: read the actual words, split hard to find from hard to do, and only the second one is a candidate for AI."
Why this works
Naming what you'd leave alone is what proves this is judgment, not a reflex to bolt AI onto everything.

Let's learn

What does it actually mean when someone tells you a task is hard?

Kernelwell recommends recipes and plans a family's week of meals, matched to their allergies, their budget, and their macros.

For four straight quarters before this one, every single piece of feedback tagged "hard to use" went straight into the interface team's backlog. No exceptions. Nobody read the actual words behind the tag, only the tag itself.

Hand sketched icon list titled Every slice we checked. Five rows, each an icon and a line of text: which screen or device the note came from, which tag the ticket got on intake, how many people are in the household, whether the household has an allergy or a macro goal, and what the person's own words actually say, this last row highlighted in a different color.
Four of the five slices came back flat. The fifth one, what people actually said, split the pile almost exactly in half.

This quarter, a new researcher read every one of them first. Over thirteen weeks, a hundred and eighteen pieces of feedback got the same tag, "hard to use." This time somebody actually reread the words behind each one before it got routed anywhere.

What 118 "hard to use" notes were actually about
70 35 0 51 Navigation 23 Allergy + macro check 19 Substitution reasoning 25 Other judgment work
A real interface fixOzma's exact complaint
51 and 67 sit almost dead even. One tag, "hard to use," had been hiding two different problems in roughly equal amounts for four straight quarters.

Fifty one of those notes really were about a control someone couldn't find, a buried filter, a setting three taps deep. Sixty seven were not. They were about a task that was still hard once you'd found the right screen: checking a recipe against two allergy lists and a set of macro targets, by hand, every week.

We didn't take her spreadsheet away. We built a filter chip next to it and called the job done.
Knowledge spark: what is a hallucinated substitution? It's when a model swaps an ingredient with total confidence and gets it wrong. It might call a cashew cream sauce "dairy free" and never mention it still contains a tree nut. It reads calm and certain either way, right or wrong.

What it costs at its worst: leadership sees "hard to use" and ships another interface pass, more filters, better labels, and the sixty seven judgment cases don't move an inch, because the screen was never their problem. Or worse, someone builds a quick recipe assistant to look responsive, lets it answer "is this safe" in free text with no real check behind it, and it tells a parent a sauce is peanut free when it isn't. A slow, honest "maybe check this one" is annoying. A fast, wrong "yes, safe" is dangerous.

The choice I would take back Kernelwell's intake process merged tagging and routing into one step. Every "hard to use" note got auto-tagged and sent straight to the interface backlog, with nobody reading the actual words in between. That was fine for two years, back when the app was simple enough that "hard to use" really did just mean the screen. It stopped being fine the day half of what people meant by hard had nothing to do with the screen at all.

What I would leave alone: the fifty one real navigation cases. A missing filter chip, a buried setting, a label that doesn't say what it means, all get a plain interface fix. Building a model to guess what someone meant by a button they couldn't find would be solving a problem the button already solves, just badly placed.

The lesson: "hard" is not one word with one meaning. It hides two completely different problems, and a tag on a support ticket cannot tell them apart. Only the actual words can.

Now here is the same thing as a story

Say the short version out loud in an interview. Read this one when you want to feel exactly how a filter chip that fixed nothing sat unnoticed for eight weeks.

Ozma Bakare can tell in about four seconds whether a recipe is safe for her son, just from the ingredient list.

Hand sketched labeled parts diagram titled Ozma, most Sunday nights. Center icon a person labeled Ozma, four labeled parts around her: checks her son's peanut list, checks her daughter's dairy list, checks her own marathon macros, does all three by hand every week.
She's been running this same check in her head, every week, since long before Kernelwell existed.

She's managed her son's peanut allergy since he was two, her daughter's dairy intolerance since last year, and she's six weeks from a half marathon, chasing a protein target her coach set for her. Every recipe that lands on her table gets checked against all three, in her head, before anyone eats.

Kernelwell arrived in her house in February, recommended by a friend from her running club. For the first few weeks it was genuinely useful. It found her fifty new dinner ideas she'd never have searched for herself, tagged high protein, dairy optional, nut free wherever the recipe's headline ingredients allowed it.

But "wherever the headline ingredients allowed it" turned out to be doing a lot of quiet work in that sentence. Kernelwell tagged recipes by their main ingredients, not every line of the method. A "nut free" curry might still list a nut oil three steps down, in the marinade nobody scrolls to. So Ozma kept doing exactly what she'd always done: reading every recipe's full ingredient list herself, checking it against her son's list, her daughter's list, and her own macro target, by hand, every single week. The tool changed what she found. It never touched what she still had to check.

By March, in every research interview Kernelwell ran with her, she said some version of the same thing: "this app is hard to use." The research team wrote it down, tagged it "hard to use," and moved on to the next interview.

In week 3, the interface team shipped what looked like the obvious fix: allergy filter chips, so you could filter the whole recipe list down to "nut free" and "dairy free" in one tap instead of hunting through menus.

Hand sketched comparison diagram titled Two very different kinds of hard. Left panel, a question mark icon, labeled Hard to find, caption the allergy filter is buried three taps deep. Right panel, a scale icon, labeled Hard to do, caption cross checking three lists by hand every single week.
The filter chips answered the left panel. Ozma's actual complaint had always been the right one.

It didn't help. Ozma kept saying the same thing in April, and in May. The filter chips solved a problem she'd never actually had, finding recipes. They left completely untouched the one she did have, checking them.

Hand sketched timeline titled What we shipped, and when we actually looked. Four milestones left to right: filter chips ship in week 3, looks like a fix. Feedback still says hard in week 6. New hire rereads notes in week 9, the real trigger. Recut finds the split in week 11, this last one emphasized.
Eight weeks sit between the fix that looked obvious and the first person who actually checked whether it worked.

Nobody noticed the gap for six more weeks, because "hard to use" kept getting logged as one tag, and one tag looked like one problem.

In week 9, Tobit Sarafyan, a researcher who'd joined Kernelwell a month earlier, sat down to prep a quarterly report and did something nobody had done in four straight quarters: he reread the actual words behind every single "hard to use" ticket, instead of trusting the tag. Twenty minutes in, he found Ozma's line from March, word for word: "I have to check every recipe against my son's list, my daughter's list, and my own macros by hand, every week, and the app doesn't help with any of it."

He brought it to Ionara with one question: "why are we treating 'I can't find the filter' and 'I have to check three lists by hand' as the same ticket?"

That one question was the whole quarter, sitting inside a single sentence nobody had reread until week nine.

Ionara ran the recut herself. She pulled every "hard to use" note from the last thirteen weeks, a hundred and eighteen of them, and this time she sorted by what the words actually said, not the tag.

Hand sketched quadrant diagram titled Sorting what people actually said. X axis how specific the words are, from vague to exact. Y axis about doing versus about finding. Four points plotted: general grumble, low on both axes. Filter request, medium specific, low on the doing axis. Typo report, medium on both. Ozma's verbatim, high on both axes, in the top right corner.
The most specific complaint came from the household doing the most by hand. That's not a coincidence, it's the whole diagnosis.

Twenty three households, unprompted, described the same cross check Ozma had: allergies plus macros, done by hand, every week. Nine of those twenty three had already built their own workaround, a shared spreadsheet, a running notes app list, to do it themselves.

Households describing their own manual workaround, week by week
10 5 0 Wk 1 filter chips ship Wk 3 Wk 5 Wk 7 Tobit rereads notes Wk 9 Wk 11 Wk 13
Households with a manual workaroundFilter chips shipVerbatims reread
The line never bends at week three. Whatever the filter chips fixed, it wasn't the thing driving people to build their own spreadsheet.

Before Ionara acted on any of it, she wrote out three honest guesses, because a real pattern and a convenient one can look identical from across a support queue. One, the "hard to use" count was just noise, unlucky households, no shared cause. Two, it really was navigation, and the March filter chips just hadn't reached everyone's app version yet. Three, it was a genuine task, checking allergies and macros across a recipe, that no filter or menu could ever remove, because the checking itself was the hard part.

Only the third guess survived contact with the actual numbers. The filter chips had reached every user by week 4. The households still describing the cross check by hand, in week 11, all had the update. They just weren't being helped by it, because it was never their problem.

Hand sketched decision tree titled The evidence test. Root question, if navigation were perfect, is the hard part gone. Two branches: it was about finding, leads to gone, real UI fix. It was about judging, leads to still there, AI candidate.
One question, run against all 118 notes, is what actually separated the two piles.

The test that told the two apart was simple to say and took real work to run: if navigation were fixed perfectly, would the hard part still be there? For the fifty one, no. Find the filter, and the complaint disappears. For the twenty three, yes. Ozma could have every allergy chip Kernelwell ever built, tap them all correctly on the first try, and she would still be sitting at her counter on Sunday night, reading every ingredient line by hand, because Kernelwell never actually checked the recipe against her son's allergy list. It only ever checked whether the recipe's headline said "nut free."

Ionara's team had one faster option on the table before this: a single disclaimer on every recipe, "please check ingredients carefully," and call it done. She rejected it. It would have covered Kernelwell, not Ozma, and it still would have left her doing the exact same manual check every week, just with an extra sentence to read first.

The AI-specific risk sat right underneath the fix Ionara actually wanted to build: a model that reads a recipe and a household's allergy and macro profile, and flags a real conflict, has to be right, not just fluent. A model asked "is this dairy free" will answer in the same calm, certain voice whether it's correct or hallucinating a swap that still contains a dairy derivative. Before it could ship, that check had to clear a set pass rate on a labeled set of recipes with known allergens, not just look right on the handful the team happened to try. The guardrail Ionara built in: every flag gets checked against a real, structured ingredient and allergen database before it ever reaches a screen, never just the model's own free text claim. That check adds real time, about two to three seconds per recipe, against a near instant unchecked answer. She took that trade on purpose, because a wrong "safe" is worse than a slow one.

TRACE, so a hard task doesn't get diagnosed as a hard-to-find button

Not a way to decide whether Ozma was right to be frustrated. TRACE is what stops "just add a filter" and "build us a model" from getting the same confident answer, when only one check actually tells them apart.

TTimeline. Lay out what shipped, including things that looked like fixes.
Kernelwell shipped allergy filter chips in week 3, the fix that looked obvious. "Hard to use" feedback kept coming at the same rate through week 6. Tobit reread the actual verbatims in week 9. The recut that separated the two problems finished in week 11.
The gap that mattered wasn't week 11. It was the eight weeks between the filter chips shipping and anyone checking whether they'd actually worked.
RRecut. Slice the pile every real way you have.
Ionara sliced the hundred and eighteen notes by device, by feature tag, by household size, by whether the household had an allergy or macro goal, and by what the words themselves actually said. Four of those slices came back flat. The fifth split the pile almost exactly in half, fifty one and sixty seven, on a line no ticket tag had ever drawn.
A tag built for routing tickets fast is not built for telling two different problems apart. Only the actual words could do that.
AAssume nothing. Neither easy story gets the benefit of the doubt.
Ionara didn't assume every "hard to use" complaint was a navigation problem just because that's what four straight quarters of routing had assumed. She also didn't assume all sixty seven judgment cases were true AI opportunities the moment she saw the number, without checking whether a simpler fix might quietly cover some of them too.
Both easy guesses are cheap to reach for. Neither one gets to stand in for a real check of the actual words.
CCause candidates. Three honest guesses, checked against real evidence.
One, random noise, no shared cause. Two, genuine navigation friction that the March filter chips simply hadn't reached yet. Three, a real judgment task, checking allergies and macros across a recipe, that stays hard no matter how good the menu is.
Only the third guess matched what the filter chip rollout data and the raw verbatims both actually showed.
EEvidence test. The one question that tells the two apart.
Ask it of every remaining case: if navigation were fixed perfectly, would the hard part still be there? For the fifty one, the answer is no, so they get an interface fix. For the twenty three describing the cross check by hand, and the nine who'd already built a spreadsheet for it, the answer is yes, so that's where the AI feature belongs.
This is the strongest move in the whole framework. It's checkable against real behavior, not a guess about which explanation sounds better in a meeting.

Three things worth saying plainly, since interviewers push here. The team's first, faster idea was a blanket ingredient disclaimer on every recipe, rejected because it still left Ozma doing the same manual check every week, just with an extra line to read. The AI-specific risk worth naming by name: a model checking a recipe against an allergy list can sound exactly as certain when it's wrong as when it's right, so every flag it raises gets checked against a real structured ingredient database before it ships, never trusted as free text alone. And the trade-off, accepted on purpose: that check costs two to three seconds per recipe against an instant, unchecked answer, worth it because a wrong "safe" is worse than a slow one.

And if you want to be sure it really works, try it somewhere else

Same five letters, a loading dock at six in the morning instead of a kitchen counter. This time most of the "hard to use" pile really was navigation, and only a genuine minority was ever a candidate for AI.

GridWrench is a scheduling and dispatch app HVAC field technicians use to see their day's jobs, reroutes, and parts. Wynfield Kastor owns dispatch product there. Over an eight week stretch, forty pieces of feedback said some version of "I need smarter routing."

Wynfield ran the same recut. He read the actual words behind every one of the forty notes, not the tag. Thirty one were genuinely about not being able to find where a reassigned job had gone, buried three screens deep after a dispatcher moved it. Nine were about something else entirely: three jobs go red at once, and the technician has to decide which one to hit first, weighing drive time, parts on the van, and which customer's contract has the tightest service window. No amount of finding the reassigned job faster removes that decision.

Ticket friction, by how many screens it took to find a reassigned job
1 2 3 4 5 Screens to find the reassigned job Judgment cases: low screens, still high friction
Rises with screens, a real UI problemHigh friction regardless, a judgment problem
Thirty one tickets fall in a clean line: more screens, more friction. Nine sit off that line completely, low screens, still high friction, because the friction was never about the screens.

Wynfield's team floated one AI idea early: a model that predicts the "best" next job automatically and just tells the technician what to do. He rejected it for the thirty one navigation cases, because building a smart recommendation on top of a screen someone simply couldn't locate would treat a navigation problem as if it were a judgment problem. He shipped a persistent "reassigned today" list on the home screen instead, and the thirty one cases dropped out almost immediately.

The decision Wynfield would take back GridWrench buried a reassigned job wherever the dispatcher happened to leave it in the schedule tree, with no separate view for "moved since this morning." That was fine when reassignments were rare. It stopped being fine once a third of a shift's jobs could move in a single morning, and no screen ever collected them in one place.

The nine judgment cases stayed open, a real candidate for a future tool that weighs drive time against parts against contract terms. But nine out of forty wasn't worth a model before the thirty one cheap, honest wins had shipped.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it, don't assume every "hard to use" note is a UI problem or an AI problem, check what the words actually say first.
Cost: no time to reread all hundred and eighteen notes. Sample twenty, and ask the same one question of each: would this survive a perfect navigation fix?
The model got better, for real: say a future version of Kernelwell checks allergies with total accuracy. The evidence test still runs exactly the same way, it just moves where the honest line between UI and AI sits, not whether you still need to check.

Where people run it wrong.
They read the tag and skip the words, because the tag is faster and feels like enough.
They find one real judgment case and assume the whole pile is the same, building AI for complaints that were genuinely just about a missing button.
They find a real UI problem and build a clever model around it anyway, because "AI" sounds better in a roadmap review than "we moved a filter."

How to use it live. When an interviewer throws this at you cold, buy two seconds by asking one thing back: "do we know yet whether people mean they can't find it, or they can't do it?" That question alone is usually exactly what a question shaped like this one is listening for.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits telling a hard-to-find button apart from a hard task?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Built to find the real cause before deciding what to build.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ozma Bakare, a parent managing two kids' allergies and her own marathon macros, and Ionara Vradenburg, who owns Kernelwell's recipe recommendation product.
3 · THE TIMELINE
What shipped in week 3, and what actually got found in week 11?
Tap to flip
ANSWER
Allergy filter chips shipped in week 3, looking like a fix. In week 11 the recut showed 67 of 118 notes were really about a judgment task, not navigation, and the chips had never touched them.
4 · THE RECUT
What did slicing by the actual words reveal?
Tap to flip
ANSWER
118 notes tagged "hard to use" split into 51 about a control someone couldn't find, and 67 about a task that stayed hard even after they found the right screen.
5 · THE OLD DECISION
What decision would Ionara take back?
Tap to flip
ANSWER
Kernelwell's intake process merged tagging and routing into one step, so every "hard to use" note went straight to the interface backlog with nobody reading the actual words in between.
6 · THE NUMBER
Fill in the blank: ___ of 118 notes were about a task that stayed hard even once found. ___ households described the exact same cross check unprompted. ___ of those had already built a spreadsheet to do it.
Tap to flip
ANSWER
67. 23. 9. The spreadsheet number is the one that turns a complaint into proof.
7 · THE EVIDENCE TEST
What's the one question that tells "hard to find" apart from "hard to do"?
Tap to flip
ANSWER
If navigation were fixed perfectly, would the hard part still be there? Yes means it's a candidate for AI. No means it's a real interface fix.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what did the recut find that time?
Tap to flip
ANSWER
GridWrench, an HVAC dispatch app run by Wynfield Kastor. That time most of the pile, 31 of 40, really was a navigation problem. Only 9 were a genuine judgment task worth an AI feature later.

Check yourself Score: 0 / 0

Fill in the blank
1. Of the 118 notes tagged "hard to use" this quarter, ___ were genuinely about a control people couldn't find, and ___ were about a task that stayed hard even after they found the right screen.
Show hint
Look at the bar chart in Let's learn.
Show answer
51. 67. The recut split almost exactly down the middle, on a line the ticket tag itself never drew.
Multiple choice
2. Which of these is the strongest signal that a complaint is a candidate for AI, not a redesign?
  • A. The complaint uses the word "confusing."
  • B. The person describes checking several things by hand, and would still have to after a perfect redesign.
  • C. The complaint came in more than once.
  • D. The person asked for a new button.
Show hint
Check the evidence test in the framework recap.
Show answer
B. Frequency and words like "confusing" don't tell UI apart from judgment work. Surviving a perfect navigation fix does.
True or false
3. True or false: because the March filter chips didn't fix Ozma's complaint, that proves the filter chips were a bad idea.
  • True
  • False
Show hint
Look at what the bar chart says the filter chips were actually built for.
Show answer
False. The filter chips were a genuine fix for a genuine problem, the 51 real navigation cases. They just weren't built for Ozma's problem, which was never about finding a control.
Short answer, where it wouldn't matter
4. Name a part of Kernelwell's research pipeline that should NOT be rebuilt around AI, based on this quarter's recut. Why not?
Show hint
Look at "What I would leave alone" in Let's learn.
Show answer
Model answer: The 51 real navigation cases. A missing filter chip or a buried setting needs a plainer menu, not a model guessing at what someone meant by a button they couldn't find.
Short answer, apply it yourself
5. Think of a product you use where you've built your own workaround, a spreadsheet, a note, a habit, to do something the product itself doesn't do. What would the evidence test say about it?
Show hint
Ask whether a better menu in that product would actually remove the thing you're doing by hand.
Show answer
Model answer: If the workaround is standing in for real judgment work, comparing options, cross checking numbers, that a perfect redesign wouldn't touch, the evidence test says it's a real candidate for an AI feature, not just a UI complaint.
Short answer, the number question
6. If only 9 of 118 notes this quarter had described the allergy and macro cross check, instead of 23, would the evidence test still be worth running? Why or why not?
Show hint
Think about what the evidence test is actually checking for.
Show answer
Model answer: Yes. The evidence test isn't about hitting a minimum count, it's about whether the hard part survives a perfect fix. A smaller number might change how urgently you'd build it, but not whether it's a UI problem or a judgment problem.
Before you close the answer
Why this works
Tests whether you'll read the actual words behind a complaint before reaching for a fix, instead of pattern matching "user says hard" straight to either a redesign or a model. Most candidates pick one lens and apply it to everything.
Follow-up traps
"Isn't reading raw verbatims instead of trusting your own tags just extra work most teams won't do at scale?" Response: it only has to happen once per pattern, not per ticket. Ionara reread 118 notes one time and it changed how every future "hard to use" tag gets triaged.

"What if a complaint has both signals, someone can't find the filter and also does the cross check by hand?" Response: split it. Fix the finding half with the interface, and still run the evidence test on the doing half separately, since fixing one doesn't retire the other.
If pressed
The 23 households who named the cross check by hand were themselves a sample of the 67 broader judgment cases, so before committing engineering time to the AI feature, Ionara still needed a second check: whether the 44 remaining judgment cases, the ones that never explicitly mentioned a workaround, were the same problem in different words, or something else again. That second recut hadn't finished by the time this quarter closed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more