ConceptFoundationalModel Fluency & the AI PM Role / The AI literacy baseline every PM needs / #10
What does it mean for a model to be multimodal and what product surfaces does that unlock?
PICK · a laundromat flyer's smallest line nearly costs Suvarna Rathnayake the warehouse job Vellacourt was supposed to help her land
Vellacourt reads a job posting and rewrites a job seeker's resume and cover letter to fit it. Most postings get pasted in as plain text. Suvarna Rathnayake found hers taped to a laundromat bulletin board, a printed flyer for a warehouse job at Runcible Distribution, so she photographed it instead. Ozara Mainwaring runs product for Vellacourt, and for ten weeks nobody there knew that one small, creased line on flyers like Suvarna's kept vanishing on the way to the resume.
The direct answer
A multimodal model can take in, or make, more than one kind of thing: words plus a picture, sound, or video, not text alone. For Vellacourt that only opens a real new product surface twice: when the picture is the only copy of a posting that exists anywhere, and when a screenshot's layout, what's bold, what's boxed, what's crossed out, carries a requirement a plain paste would drop. Build reliably for those two cases. Treat every other "let people upload a photo" idea as a shortcut around the same text box, not a new capability, and never let a low-confidence read from a real photo vanish silently instead of getting flagged.
Do this, in order
Build photo upload only for postings with no text version, and screenshots whose layout genuinely carries a requirement.Why: that is the actual new capability multimodal buys; everything else is a camera icon on a text box.
Run the kill test on every new "let them upload a picture" idea before it's built.Why: does the picture carry something a paste couldn't. If not, it's a gimmick, not a feature.
Score every extracted field's confidence and show the low ones instead of guessing them.Why: a model that's unsure and a model that found nothing must never look the same on the screen.
Never auto-generate straight from an image with no editable preview step.Why: this is the exact catch point a messy real photo needs, and the one Vellacourt removed to make the demo feel instant.
Track extraction accuracy by source type, screenshot versus photographed flyer, not one blended number.Why: a blended average hid a category quietly cratering for nine weeks.
Leave the plain paste box exactly as it is.Why: most postings still arrive as text, and pasting them already works.
How to answer this, stage by stage
Nobody is grading whether you can define "multimodal" like a glossary entry. They're grading whether you can say, plainly, which pictures are worth building for and which ones are just a nicer looking text box.
1
Scope it to one photo, one flyer
Say it like this
"Let's make this real. Vellacourt tailors resumes and cover letters to a job posting. Say a job seeker named Suvarna finds a warehouse job posted on a laundromat bulletin board, no website, just a printed flyer with a phone number. That's the picture I want to answer this question about."
Why this works
A real photo stops "multimodal" from staying a vocabulary word nobody can actually build against.
2
Say the structure out loud
Say it like this
"I'll run this with PICK. Position, what multimodal actually means and what it unlocks here. Impact, who pays when we get the call wrong either way. Cost asymmetry, what's cheap to demo versus what's expensive to make real. Kill criteria, the one test that tells a real feature from a gimmick."
Why this works
Naming the structure first tells the interviewer you have a method, not just a definition memorized for the interview.
3
Define it plainly, then commit, in the same breath
Say it like this
"Multimodal just means a model that can take in, or make, more than one kind of thing, words and a picture, not words alone. For Vellacourt that's real in exactly two places: a flyer with no text version anywhere else, and a screenshot whose bold and color tell you which requirements are must-haves. Everywhere else, a picture is just a slower way to fill in the same box."
Why this works
Giving the definition and the commitment together is what makes this a direct answer instead of two separate thoughts.
4
Show the impact both ways, with real numbers
Say it like this
"Undersell it, and Suvarna still has to type a flyer out by hand, badly, on her phone, in a laundromat. Oversell it, and Vellacourt reads her flyer's forklift certification line at nine percent confidence and just drops it, no warning, and she submits a resume that never once mentions the one thing that employer actually asked for."
Why this works
Naming both losers, not just the flashy failure, proves you're weighing a real tradeoff and not just telling a scary story.
5
Point straight at the cost asymmetry
Say it like this
"A demo of 'take a photo, get a resume' is maybe two days of engineering with any vision model wired up. Making it hold up on a creased flyer under bad laundromat lighting, that's the expensive part, and it's not glamorous, it's confidence scoring and a review screen nobody claps for in a demo. Skip that part, and the failure doesn't show up on launch day. It shows up nine weeks later, in an audit."
Why this works
This is the hardest move in PICK: naming which mistake is loud and cheap, and which one hides.
6
Give the kill test itself
Say it like this
"Here's the test I'd run on any new 'let them upload a photo' idea: does the picture carry something the person genuinely could not have typed as well? If yes, build it properly, with confidence scoring and a review step. If the honest answer is no, nothing's lost by typing it, it's a gimmick, leave it behind a plain text field."
Why this works
A test you can run on the next idea is worth more than an opinion about this one.
7
Say what you'd leave alone
Say it like this
"The plain paste-the-text box needs nothing done to it. Most postings still live online as text. Pasting them is instant and complete, and no picture makes that faster or safer."
Why this works
Naming what's already fine stops the answer from sounding like blanket suspicion of every feature that touches a model.
8
Close on one line
Say it like this
"Multimodal is real here exactly twice: the flyer with no other copy, and the screenshot whose layout is the requirement. Build those two properly, with a confidence score and a review step. Leave the rest as a text box that happens to have a camera icon nobody needed."
Why this works
One sentence the interviewer can actually write down is the sentence that gets remembered.
Let's learn
Vellacourt reads a job posting and rewrites a job seeker's resume and cover letter to fit it. Type or paste the posting in, and out comes a resume and a cover letter built around that exact job, not a generic one.
Five steps. The middle one, turning a posting into fields Vellacourt can build with, is the step this whole answer turns on.
Vellacourt built the pasted-text path first. A job seeker copies a posting off a company site or a job board and pastes it in. About 22,000 resume and cover letter pairs a week come from that path, and extraction is close to perfect, because the text was already typed by a computer, clean, in order, nothing to misread.
Knowledge spark: what is a multimodal model?
A model that can take in, or make, more than one kind of thing. Words plus a photo, a sound clip, or a video, instead of words alone. Reading a picture and reading a sentence used to need two separate systems. A multimodal model does both with one.
Then Vellacourt added a second way in: take a photo or a screenshot instead of typing anything. Uploads grew fast, from about 400 a week to about 2,300 a week over ten weeks. Two very different kinds of pictures started arriving under the same button.
One picture is the only record of a job that exists. The other is the same words, in a picture instead of a paste box.
Some uploads were clean screenshots of a job board post, already digital text, just captured as an image instead of copied as words. Others were photos of real, physical flyers: taped to a wall, folded, lit by whatever light happened to be in the room. Vellacourt's own extraction accuracy split hard along that line.
Field extraction accuracy, clean screenshot versus photographed flyer
Clean screenshotPhotographed physical flyer
Screenshots barely moved the needle. Real, messy photos lost the most ground on exactly the fields that decide whether an application even gets seen.
The gap didn't show up as a wrong answer. It showed up as silence. Vellacourt's extraction gave every field a confidence score internally, but the product never surfaced it. A field read at 9 percent confidence and a field that was never on the page at all looked exactly the same to a job seeker: not there.
The messy photos were never the real problem. The real problem is a mistake nobody got told about.
What it costs at its worst: Suvarna's flyer had one line in small print, cut slightly by a fold crease, "must currently hold valid forklift certification." She has one, current. Vellacourt's photo-upload flow read the header and the duties confidently, wrote her a generic resume about punctuality and stamina, and never mentioned her certification, because that line scored 9 percent confidence and got dropped with no flag, no note, nothing. She heard nothing back for eight days, until a friend mentioned the flyer's forklift line by name. She went and reread it herself, rewrote her cover letter's opening by hand, and resubmitted. Eleven days after her first, generic try, she got a callback.
The choice I would take back
When Vellacourt built the paste-text path, it showed users an editable preview, "here's what we found," before writing anything, one screen, thirty seconds. When photo upload shipped, that screen got skipped for images, on purpose, to make "take a picture, get a resume" feel instant. That was fine while most photos were clean screenshots, extraction rarely missed anything worth catching. It stopped being fine the moment real, physical flyers started arriving, glare, creases, small print, exactly where extraction quietly failed, and the one screen built to catch it was the one just removed.
What I would leave alone: the plain paste-the-text box. Roughly 22,000 tailored packages a week still start there. Pasting a posting captures everything already, and no picture makes that path faster or safer. Nothing about this problem touches it.
The lesson: test every new "let's add a picture" idea against one question, is the picture the only copy of the thing, or does it show something words alone would drop. If neither is true, it's a slower typing box wearing a camera icon, and the real cost isn't the feature, it's the confidence someone builds in a product that quietly wasn't reading what it claimed to read.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for the ten weeks that built up to a question nobody could answer on the spot.
Ozara Mainwaring had run product at Vellacourt for three years, most of it spent making the paste-the-text path faster: fewer clicks, a cleaner preview, a resume that read like it was written by someone who'd actually looked at the posting. It worked. Job seekers pasted a listing in, checked the preview, hit generate, and got something they'd actually send.
Photo upload was supposed to be the same idea, just easier. Not everyone finds a job online. Some people still find one on a corkboard.
Nobody decided, on any single day, to let a real requirement disappear without a word. A screen built for one kind of photo just never got built for the other.
For the first few weeks, photo upload looked like a clean win. Most people who tried it uploaded a screenshot, a job board post captured as an image instead of copied as text, because that's who tries a brand new button first: people already comfortable online. Extraction on those was nearly as good as pasted text. The team watched one blended accuracy number, and it barely moved.
Knowledge spark: what is a confidence score?
A number the model gives itself, alongside its answer, for how sure it is. High means it's confident it read that line correctly. Low means it's guessing, or the text was too faint, blurred, or cut off to be sure. A product can show that number to a person. Vellacourt's didn't.
Word of mouth carried the feature past the early, screen-savvy users. By week six, more uploads were coming from people who'd found a job the old way, on a wall, a windshield flyer, a printed sheet handed across a counter, and photographed it because someone told them the app could handle a picture now. Those were the uploads the blended number was quietly hiding.
Suvarna's flyer arrived in week eight. She'd walked past the Runcible Distribution posting at her laundromat a dozen times before she photographed it, dim overhead light, one fold crease running straight through the bottom corner where the requirements sat in small print. Vellacourt's model read the bold header and the main duties without any trouble. The forklift certification line, the one requirement that actually mattered most to that employer, scored a 9 percent confidence internally and got left out. Nothing on her screen said a line had been found and doubted. Nothing said a line existed there at all.
Everything else on the flyer made it through fine. The one line an employer actually screened for is the one that never reached her resume.
She sent the generic version. Eight days of silence followed. Then a friend, who'd walked past the same corkboard, asked her, almost in passing, whether she'd mentioned that Runcible specifically wanted forklift-certified people. She hadn't seen a line about that in her own resume, because she'd never seen it in the tailored version Vellacourt built for her. She went back to the actual flyer, found the small print herself, and rewrote her cover letter's opening line by hand to lead with her certification. Eleven days after her first, generic submission, she got a callback. Three days after that, an interview.
It surfaced properly a week later, almost by accident. The newest member of Vellacourt's quality team, three weeks into the job, was doing a routine sample of support threads and found Suvarna's, a short, mild note thanking Vellacourt "eventually." Curious, she pulled the internal extraction log behind it and found the forklift line sitting at 9 percent confidence, silently dropped, not flagged anywhere, not even in the team's own internal review tools. She asked Ozara a question nobody could answer on the spot: "If we saw this line and weren't sure, why didn't we just say so?"
Vellacourt did not get the flyer wrong. It got the one line that mattered, and said nothing.
The audit that followed pulled 500 photo-uploaded postings from all ten weeks since launch and reran them with confidence logging turned on. The rate of flyer-sourced tailorings carrying at least one silently dropped requirement had been climbing the whole time, invisible inside a blended number that mostly reflected the early, clean screenshots.
Week 1: 12 percent of photographed flyers had a silent drop. Week 3: 19 percent. Week 5: 24 percent, past the 20 percent bar the team would have wanted, if anyone had set one. Week 7: 31 percent. Week 9: 39 percent, the week the question finally got asked.
Rate of at least one silently dropped requirement, photographed flyers, by week since launch
Nobody was watching this number on its own. It lived inside a blended average that stayed calm while a third of a fast growing category quietly went wrong.
By week nine, about 39 percent of that week's roughly 2,300 photo uploads carried at least one silently dropped requirement, close to 900 tailored packages in a single week, out into the world, missing the one line that might have mattered most.
The decision Ozara would take back sits in a planning meeting months earlier, when the photo-upload team debated whether to keep the editable preview screen for images. Someone argued it broke the magic: type nothing, get a resume, that was the whole pitch. Someone else pointed out the preview only ever caught small things on pasted text anyway. Both were true, for pasted text. Nobody in that room was thinking about a flyer with a fold crease through the one line that mattered.
Run the same ten weeks again, with the preview screen kept for every photo upload and a visible flag on anything scored under, say, 40 percent confidence. Suvarna's flyer still gets photographed in the same dim laundromat light. The forklift line still comes back at 9 percent. But this time it shows up on her screen with a small mark: "not sure about this line, worth checking." She glances at the real flyer, sees the requirement, adds it herself. Her first submission is the good one. No eight days. No friend's remark. No second try.
What Ozara would tell herself, back in that planning meeting: the magic was never the missing screen. It was the resume actually matching the job. Taking out the one thing built to guarantee that, to protect a feeling of speed, was never a fair trade.
The photo test: PICK, one letter at a time
Not a way to decide whether pictures are good or bad. PICK is what forces you to say, out loud, which pictures are actually worth the engineering, and which ones just look impressive in a demo.
PPosition. The claim, in one sentence, before any numbers.
Multimodal means a model that can take in, or make, more than one kind of thing, words plus a picture, not words alone. For Vellacourt, that unlocks a real new surface exactly twice: a posting with no text version anywhere else, and a screenshot whose layout, bold, color, a badge, carries a requirement a plain paste would drop.
Say the claim before the story. A position that only shows up after the evidence sounds reverse engineered from it.
IImpact. Who feels each kind of wrong.
Undersell multimodal, and a job seeker like Suvarna still has to retype a flyer by hand into a text box, slowly, on a phone, in a laundromat, badly. Oversell it, and Vellacourt reads a real, low-confidence line and drops it silently, so the one requirement an employer actually screens for never reaches the resume at all.
Naming both people who lose, not just the dramatic failure, is what keeps this from sounding like a one-sided horror story.
CCost asymmetry. The heart of it.
A demo of "take a photo, get a resume" is cheap: a couple of days wiring up any vision-capable model to the existing text pipeline, and it looks great in a pitch. Making that hold up on a creased, dimly lit, physical flyer is the expensive, unglamorous part: confidence scoring per field, a mandatory review screen, a way to tell a person "I'm not sure about this line" instead of quietly guessing. Vellacourt now holds every photo-sourced tailoring behind that review screen, on purpose, even for screenshots where it rarely changes anything. It costs 20 to 30 extra seconds a session. That's the trade accepted: a little slower, in exchange for never again generating a resume that silently drops the one thing that mattered.
This is the step that actually earns the position. Anyone can say "test it first." Naming which failure hides, and what it costs to stop hiding it, is what survives a follow-up question.
One cost is small and everyone can see it happening. The other one hides behind a resume that looks, on the surface, completely finished.
KKill criteria. The test that actually separates real value from novelty.
Does the picture carry something the person genuinely could not have typed as well? A flyer with no online listing passes, there's nothing to paste. A screenshot whose bold text marks the must-have requirements passes, typing loses the emphasis. A screenshot of a plain-text job post fails, typing it would carry exactly the same information. Any proposed multimodal feature that fails this test gets built as a convenience later, if at all, not treated as a new surface.
A position with no way to be proven wrong is just an opinion held tightly. A one-line test anyone can apply to the next idea is what makes this a method instead of a hunch.
Three things worth saying directly, since this is where the real judgment sits. Vellacourt considered a stricter fix instead of confidence scoring and a review screen: raise the extraction threshold so anything below a high bar gets left blank across the board, no guessing at all. It lost, because plenty of low-contrast lines on real photos were still read correctly, and a blanket threshold would have thrown those out too, trading one kind of missing information for a different kind, instead of just telling the user what the model wasn't sure about. The AI-specific failure worth naming by name is a silent low-confidence drop: a model that is unsure behaves exactly like a model that found nothing, and nothing in the product told anyone the difference. The guardrail is confidence-scored extraction paired with a mandatory review step for anything under the bar, not an engineering log only the team itself ever looks at. And the trade-off is real and accepted on purpose: the review screen adds 20 to 30 seconds to every photo-sourced session, a small, deliberate cost against ever again generating a resume that quietly leaves out the one line an employer was actually screening for.
And if you want to be sure it really works, try it somewhere else
Same four letters, an insurance claim instead of a job flyer, and this time the picture that matters is a dented bumper, not a printed page.
Castleton Mutual lets a policyholder file an auto claim by typing a description of the damage or uploading a photo of it. Osaretin Braithewaite runs AI product there, and faced a version of the same question Ozara did: which claims actually need the photo, and which ones are just asking for a picture because a picture feels more thorough?
Same shape of pipeline as the resume one. The step that decides everything is still the model, and still for the same reason, what it was actually shown.
Mapped onto PICK: Position, a photo genuinely unlocks a new surface when it shows something a typed description can't capture as precisely, dent depth, the exact point of impact, the color of paint transferred from whatever hit the car. A typed "small dent, rear bumper" collapses three different repair tiers into one vague sentence. Impact, undersell it and an adjuster keeps asking policyholders to type longer and longer damage descriptions that still miss the details that set the price; oversell it and Castleton starts asking for photos of paperwork, an insurance card, a VIN, that were already typed into the claim form, adding friction for zero new information. Cost asymmetry, a demo of "upload a photo, get an instant estimate" is quick to build; making it reliable across bad phone cameras, night photos, and rain-streaked bumpers is the expensive part, and skipping it means a wrong repair tier that only surfaces when a shop's own estimate doesn't match. Kill criteria, the same test travels unchanged: does the photo carry something the typed description genuinely could not. Damage, yes. A document that was already typed elsewhere, no.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the test: does the picture carry something the words alone would lose? If not, it's a gimmick, full stop.
Cost: no budget to build confidence scoring this quarter. Keep the mandatory preview screen for every photo upload instead, so a person catches what the model quietly missed, not nobody.
The model got better, for real: say the extraction model doubles its accuracy on messy photos next quarter. Keep the review screen anyway. Being better at reading blur is not the same as never being wrong about the one line someone actually needed.
Where people run it wrong.
They add a camera icon to a screen that already has a working text box, because a photo demo feels more impressive, whether or not the photo adds anything real.
They let "the model read it" and "the model actually knows what it read" collapse into one thing, so a low-confidence guess looks exactly like a correct answer on the screen.
They test the feature on their own clean, well-lit phone photos, then ship it for flyers taped to a wall in bad light, and call the gap a rare edge case instead of the whole real world.
How to use it live. Before answering, ask yourself one thing out loud: if I made the user type this instead of photographing it, would they lose anything real? If you can't name what they'd lose, you haven't found a real surface yet, you've found a nicer looking text box.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to define a model capability and commit to which product surfaces it actually unlocks?
Tap to flip
ANSWER
PICK: commit to a position, name who feels the impact each way, find the cost asymmetry between a cheap mistake and a hidden one, then give the test that would flip your pick.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ozara Mainwaring, who runs product for Vellacourt, and Suvarna Rathnayake, a job seeker who photographed a warehouse job flyer taped to a laundromat wall.
3 · THE POSITION
What's the actual claim, in one sentence?
Tap to flip
ANSWER
Multimodal only unlocks a real new surface twice for Vellacourt: a posting with no text version anywhere, and a screenshot whose layout carries a requirement a plain paste would drop. Everywhere else it's a slower text box.
4 · THE IMPACT
Name one real cost on each side of this tradeoff.
Tap to flip
ANSWER
Undersell it: a job seeker still has to retype a flyer by hand, badly. Oversell it: the model reads a low-confidence line and drops it silently, so a real requirement never reaches the resume.
5 · THE KILL CRITERIA
What's the one test for whether a multimodal feature is real value or a gimmick?
Tap to flip
ANSWER
Does the picture carry something the person genuinely could not have typed as well? If yes, build it properly. If no, nothing is lost by typing it, it's a gimmick, leave it behind a text field.
6 · THE NUMBER
Fill in the blank: Vellacourt read Suvarna's forklift line at ___ percent confidence and said ___ about it. By week nine, ___ percent of photographed flyer tailorings carried a silently dropped requirement.
Tap to flip
ANSWER
9 percent confidence. Nothing. 39 percent, close to 900 tailorings that single week.
7 · THE OLD DECISION
What decision would Ozara take back?
Tap to flip
ANSWER
Vellacourt skipped the editable preview screen for photo uploads, to keep the "take a picture, get a resume" flow feeling instant. That was fine while most photos were clean screenshots. It stopped being fine once real, messy paper flyers arrived, exactly where extraction quietly failed.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs PICK again for a different product. Which one, and what plays the role of the forklift line there?
Tap to flip
ANSWER
Castleton Mutual, an auto insurance claims tool. A photo of the actual dent, showing its depth and paint transfer color, carries information a typed damage description cannot capture as precisely.
Check yourself Score: 0 / 0
Fill in the blank
1. Vellacourt read the forklift certification line at ___ percent confidence, and the product ___ about it.
Show hint
Check the story, right after the new hire pulls the internal extraction log.
Show answer
9 percent. Said nothing at all. A low-confidence read and a field that was never found looked exactly the same on Suvarna's screen.
Multiple choice
2. Why does a photo of a laundromat flyer count as a real new product surface, while a screenshot of an already plain-text job listing usually does not?
A. Photos are always read more accurately than pasted text.
B. The flyer has no text version anywhere else; the screenshot's information could be typed just as well.
C. Screenshots take longer for Vellacourt to process than photos.
D. Vellacourt charges extra to process a screenshot.
Show hint
Check the kill test in stage 6 of the walkthrough.
Show answer
B. The test is whether the picture carries something typing genuinely could not. A flyer with no other copy passes. A screenshot of plain text does not, typing it loses nothing.
True or false
3. True or false: the mistake in this story is that Vellacourt's model misread the forklift certification line on the flyer.
True
False
Show hint
Check what the confidence score actually measured, and what the product did with it.
Show answer
False. The model didn't misread the line, it read it at low confidence and the product then treated that exactly like the line never existed, with no flag to anyone. The failure is the silence, not the read.
Short answer, name the rejected alternative
4. Vellacourt considered a stricter fix instead of confidence scoring and a review screen. What was it, and why did it lose?
Show hint
Look at the three-things-worth-saying paragraph after the K step.
Show answer
Model answer: Raise the extraction threshold so anything below a high confidence bar is left blank across the board, no guessing at all. It lost because plenty of low-contrast lines on real photos were still read correctly, and a blanket threshold would have thrown those out too, instead of just telling the user what the model wasn't sure about.
Short answer, apply it yourself
5. Think of an app you use that lets you upload a photo instead of typing something. Does the photo actually carry information typing couldn't, or could you type the same thing just as well?
Show hint
Ask whether anything real would be lost if the photo option didn't exist.
Show answer
Model answer: A banking app that lets you photograph a check to deposit it needs the photo, the actual handwriting and signature are the real record. A grocery app that lets you photograph a handwritten shopping list instead of typing it usually doesn't need to, typing the same items works just as well, the photo there is convenience, not new capability.
Short answer, work the number
6. If Vellacourt had set its review-flag bar at 30 percent confidence instead of having no bar at all, would Suvarna's forklift line, read at 9 percent, have been flagged for her to check?
Show hint
Compare 9 percent against 30 percent, then think about what the real gap was.
Show answer
Yes. Any reasonable bar set above 9 percent would have caught this specific line. The real gap wasn't the exact number chosen, it was that Vellacourt had no bar at all for photo uploads, so a 9 percent read and a 95 percent read went into the resume looking identical.
Before you close the answer
Why this works
Tests whether you can make a real call on your own product's capabilities instead of just defining a term, and whether the reasoning behind that call needs a model in the loop at all, not a generic "ship it carefully" instinct that would apply to any upload button. Most candidates stop at "multimodal means it can read pictures."
Follow-up traps
"Isn't dropping a low-confidence line better than showing the user something that might be wrong?" Response: no, dropping it silently is worse, because the user never learns there was anything to check. Flagging it, even uncertain, gives them the chance to look at the real flyer themselves, which is exactly what saved Suvarna's second try.
"Couldn't you just force every posting through the photo path, even clean screenshots, so the product feels consistent?" Response: that taxes every user with a slower flow for postings that were already perfectly typed. The extra care only earns its cost where a picture is doing real work, not everywhere.
If pressed
The confidence score itself comes straight from the extraction model's own token-level probability on that line, averaged across its words, not a separate classifier bolted on afterward. A blurry or partly cropped line naturally scores lower because the model is genuinely less sure what it's looking at, which is exactly why it works as a review-flag threshold instead of an arbitrary number picked to look responsible.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.