ConceptIntermediateAI Opportunity & Model Strategy / Model selection from a PM lens / #6
What does context window size actually buy you in product terms?
FLIPSfor eight months, the best part of Ines's week was a folder she never had to reread
For fourteen years, Ines Kovac has been the one person at Kestrel Rare Books Library who reads an entire donor correspondence folder before anyone else touches it. FolioMind is the tool the library built to draft a first-pass finding aid entry from a scanned folder, so junior catalogers could work through the backlog faster.
The direct answer
A bigger context window buys you the ability to hold an entire real document in the model's head at once, not just a fragment of it, so it can catch things that only show up when one part is compared against another: a name spelled three ways, a date that contradicts an earlier page, a footnote pointing back to something nobody re-checked. Below that size, the model isn't wrong about what it read. It just never actually read the whole thing at the same time, and it has no way to tell you that.
Do this, in order
Size the context window to your real document length, not your average one.Why: a model chunking a long document loses the ability to compare its own earlier and later parts.
Test specifically for cross-reference catches, not just per-chunk quality.Why: a chunked model can score well on every individual page and still miss everything that spans two of them.
Watch for people quietly delegating full trust to a tool that never saw the whole thing.Why: once a chunked summary looks fine often enough, the person who used to catch the gap stops checking for it.
Give the model a way to say when a document exceeded its window, instead of answering as if it saw it all.Why: silent truncation looks identical to a genuinely careful read until something contradicts itself.
Keep short-context handling for genuinely short documents.Why: a bigger window costs more per call, and most of your real documents may not need it.
How to answer this, stage by stage
Nobody is scoring whether you know what a token is. They're scoring whether you can say, in one sentence, what actually breaks when a document is bigger than the window.
Stage 1
Scope it to one real tool
Say it like this
"Let's ground this in something real. Kestrel Rare Books Library's FolioMind drafts a first-pass finding aid entry from a scanned donor correspondence folder. I'll walk through exactly what a bigger context window bought there, and what a small one quietly cost."
Why this works
Keeps the answer from turning into a definition of tokens with nothing real underneath it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as FLIPS. Find the person, whose folder is this. Locate the habit, what they stopped doing because the tool worked. Identify the flip, the verb that snaps. Pinpoint the old decision, which choice only made sense before. Show the replay, same folder, fixed design."
Why this works
Signals a repeatable way to reason about what a spec number costs a real person, not just a definition recited from memory.
Stage 3
Reframe: it isn't "how many tokens," it's "what can be compared at once"
Say it like this
"This isn't really about a number of tokens. It's about whether the model can hold the whole document in its head at the same moment, so it can notice when page two disagrees with page forty. A chunked model can't do that, no matter how good it is on any single chunk."
Why this works
This is where a strong answer separates from reciting "context window equals how much input it accepts."
Stage 4
Give the one decision
Say it like this
"Here's the direct answer: a bigger context window buys you cross-reference catches, a repeated name, a contradicting date, a footnote pointing back a chapter, things that only exist in the comparison between two parts of the same document, not in either part alone."
Why this works
Names the concrete capability, not a vague sense that "more context is better."
Stage 5
Prove it with the compressed evidence
Say it like this
"FolioMind's first version processed a folder in roughly eleven-page chunks, no memory between them. A new cataloger asked why a donor's name appeared three different ways across one folder, and why a letter dated 1932 seemed to reference an event the folder later placed in 1935. Nobody had caught either, because no single chunk ever saw both parts."
Why this works
Compresses the whole case into the one question a new hire asked that nobody in the room could answer.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't a generic tooling problem is that the model never signaled it had only seen a fragment. Each chunk summary read as confident and complete on its own. We accepted a real cost, a bigger model call, more expensive and a little slower per folder, to buy back the ability to catch what only shows up across the whole document."
Why this works
This is the load-bearing, AI-specific judgment. A normal document-splitting bug would at least fail loudly; a chunked model just answers fluently about the piece it saw.
Stage 7
Say what wouldn't change, then close
Say it like this
"For Kestrel's short acquisition memos, almost always under ten pages, the small context window was never the problem, and I wouldn't pay for the bigger one there. For folders over forty pages, the answer holds: the window has to cover the whole document, or the model's confidence is describing a fragment, not the folder."
Why this works
Closes with real judgment about where the fix doesn't matter, and restates the direct answer in one breath.
Let's learn
For eight months, the best part of Ines Kovac's week was a folder she never had to reread.
FolioMind reads a scanned archival folder and drafts a first-pass finding aid entry, so a junior cataloger can review and finalize it instead of writing one from scratch. Before FolioMind, Ines personally read a 60-page donor correspondence folder cover to cover, about three hours, to write one entry, catching every repeated name and contradicting date by holding the whole thing in her memory.
With FolioMind's first version, a junior cataloger got a full draft entry in about 25 minutes, since the tool processed the folder in roughly eleven-page chunks and summarized each one.
The five letters, held up as one page. Identify the flip is the one that takes the longest to find.
Here's the turn: the extra speed was never the real story. FolioMind's chunked reading meant no single call ever held the whole folder at once, so it could describe any one page confidently and still never notice that page two and page forty disagreed with each other.
The folder grew a little at a time. What the model could actually compare against itself did not move gently at all.
Ines's personal full-folder re-reads, hours per month, before and after the context-window upgrade
Nearly six workweeks a year, spent redoing a job the finding aid entry was supposed to have already done.
FolioMind was never wrong about page thirty. It just never held page two in the same hand while it read page thirty.
At its worst, a chunked read gives you a confident, well-written entry for a folder nobody can trust, and no way to tell which folders are the ones actually hiding a contradiction.
The choice I would take back
When FolioMind launched, its context window was scoped to 8,000 tokens per call, about eleven pages of scanned correspondence, to keep each call fast and cheap on the library's early digitization budget. That made sense when most folders ran under twenty pages. It stopped making sense once grant-funded digitization brought in donor folders regularly running forty to eighty pages.
What I would leave alone: for short acquisition memos, almost always under ten pages, the small window was never the actual problem, and the bigger, more expensive model call isn't needed there.
The lesson: a model can be completely honest about every page it reads and still mislead you about the folder, if the folder was never actually in front of it all at once.
Now here is the same thing as a story
The short version above is what you'd say in a budget review. Read this one for what it felt like the afternoon a new hire's question landed in a quiet reading room.
For eight months, the best part of Ines Kovac's week was a folder she never had to reread.
Ines inherited the cataloging backlog, and FolioMind, from a colleague who retired the year before. In the tool's good months, junior catalogers turned around folders in a fraction of the old time, and Ines spent her own hours on the library's hardest, most fragile manuscripts instead of routine donor correspondence.
The habit thinned in three beats, quietly, in what Ines herself did. At first, she still skimmed a junior cataloger's finished entry against the original folder before filing it. Within two months, with dozens of clean entries behind her, she stopped skimming and started trusting FolioMind's chunk summaries outright. By month six, she'd handed first-pass cataloging of every large folder to two junior staff, full stop, and moved entirely onto the manuscripts.
Ines assumed her own attention was a dial, turned down gradually. It had already become a switch, flipped months before.
A new hire, cataloging her third week, asked Ines a question over the folder in front of them both: why did a large donor correspondence folder list the same donor's name three different ways in different entries, and why did an early letter dated 1932 seem to reference an event the folder later described as happening in 1935? Ines had no answer ready.
Knowledge spark: why would a chunked model miss a contradiction like that?
A model reading in fixed-size chunks answers each chunk on its own, with no memory of what an earlier chunk said. It isn't guessing or hedging when it summarizes page forty; it's being completely accurate about page forty alone. It just has no way to notice that page forty disagrees with page two, because it was never shown page two at the same time.
Ines pulled 30 recently-cataloged large folders, all over forty pages, and personally reread each one cover to cover. Twelve of thirty, 40 percent, had at least one real cross-page inconsistency FolioMind's chunked process had missed entirely.
We didn't lose twelve entries. We lost the one thing a finding aid is actually supposed to promise: that someone checked the whole folder, not just its pieces.
The habit thinned two beats before anyone noticed. The question just happened to arrive on the right afternoon.
The real question was never whether FolioMind's chunk summaries were well written. They were. It was whether any single call had ever actually held the whole folder at once, long enough to notice it disagreeing with itself.
Four things that only exist in the comparison between two parts of a folder, never in either part alone.
When the 8,000-token window was set at launch, someone in the budget meeting said, "let's keep each call small and cheap while we're proving this out," and it sounded completely reasonable, since almost every folder on file that year ran under twenty pages.
Cross-reference catch rate, by folder length, chunked model vs long-context model
Both models look nearly identical at ten pages. The gap only shows up once a folder gets long enough to actually need comparing across itself.
Rerun the same 30 folders on FolioMind's upgraded 128,000-token window, wide enough to hold an 80-page folder whole: 29 of 30 now correctly flag the exact kind of cross-page issue the chunked version missed, at only five minutes more per folder than the chunked draft took.
What I'd tell myself, hearing that new hire's question land in a quiet reading room: FolioMind never lied to us. It just never once held the whole folder we were trusting it with.
The five steps, if you want to remember itNot a script for buying the biggest context window available. FLIPS is what tells you exactly how big yours actually needs to be.
F
Find the person. Whose folder is this?
Ines Kovac, senior archivist at Kestrel Rare Books Library, fourteen years in, who could once catch a contradiction across a folder from memory alone.
A specific person with a real, established skill makes the later loss mean something.
L
Locate the habit. What did she stop doing because it worked?
Skimming a junior cataloger's finished entry against the original folder before filing it. By month six, she'd stopped entirely and handed first-pass cataloging fully to junior staff.
The habit was rational; it kept confirming things were fine, right up until the folder that wasn't.
I
Identify the flip. What verb snaps?
Handing a task down to junior staff, then taking it back and doing it herself once quality dropped, a delegation flip with exactly two settings and no middle.
This is the hardest step. The flip isn't "the model got worse," it's what Ines did with the trust she'd extended.
P
Pinpoint the old decision. Which choice only made sense before?
Scoping FolioMind's context window to 8,000 tokens per call at launch, to keep each call fast and cheap, back when nearly every folder on file ran under twenty pages.
A repair-granularity decision: the only way to fix a missed cross-reference was to reread the entire folder by hand, since the chunked tool had no way to patch just the piece it missed.
S
Show the replay. Same folder, fixed design.
Same 30 folders, rerun on a 128,000-token window: 29 of 30 correctly flag the exact contradictions the chunked version missed, for five extra minutes per folder.
The recap, one line per letter: find is Ines, locate is the skimming she quietly dropped, identify is the delegation flip, pinpoint is the 8,000-token decision made for a smaller era, and show is the 29-of-30 replay.
And if you want to be sure it really works, try it somewhere elseSame five letters, a customs dock instead of a reading room. What "the whole document" means changes, the flip family doesn't.
Grace Tunbridge processes customs manifests for Vosberg Port Authority, which uses CargoLens, a model that reads a shipping manifest and flags line items needing a closer inspection. Mapped onto FLIPS: find the person is Grace, who used to catch a mismatched container weight across a long manifest by running her finger down every page. Locate the habit is cross-checking a flagged manifest against the ship's full bill of lading before clearing it, a habit she trusted CargoLens enough to drop within its first busy season. Identify the flip is a workaround flip: because CargoLens's short context window forced a 200-line manifest into five separate calls with no memory between them, Grace's team started keeping a private spreadsheet, copied by hand, to track a container's declared weight across all five chunks themselves, since the tool couldn't. Pinpoint the old decision is the same repair-granularity call as Ines's, scoping the window small to keep costs down when most manifests ran under 50 lines, before the port's new bulk-shipping contracts brought in manifests regularly topping 200. Show the replay is the same 200-line manifest, now read whole in one call: the declared-weight mismatch the private spreadsheet used to catch by hand gets flagged automatically, and the workaround spreadsheet, and the extra twenty minutes a manifest it cost, quietly stops being necessary.
The documents that need a bigger window aren't just long. They're long and full of parts that only make sense compared against each other.
Length alone isn't the trigger. It's length plus parts that need to be compared against each other.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a bigger window buys cross-reference catches, things that only exist in comparing two parts of the same document, and below your real document length, the model just never actually saw the whole thing," and stop.
Cost: no budget to widen the context window company-wide. Say so honestly, and commit to widening it only for the document types that are both long and cross-referential, not every call by default.
The model got better, for real: if a new model claims a much larger context window at the same price, that's still worth testing on your own longest real documents before trusting it, since "larger window" and "actually reasons across the whole thing well" aren't guaranteed to be the same claim.
Where people run it wrong.
They size the context window to their average document, not their longest real one.
They test chunk-by-chunk quality and never specifically test for a cross-reference the chunks would have to compare.
They let people quietly build a private workaround to do the comparing the model can't, and never notice it as a cost.
How to use it live. The moment an interviewer asks what a context window actually buys you, ask yourself: in this product, what real thing would only ever be caught by comparing an early part of a document against a later one? That's the concrete answer, not a token count.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Delegation flip: Ines hands first-pass cataloging down to junior staff, then takes it back once she finds cross-page inconsistencies the chunked tool was missing.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ines Kovac, senior archivist at Kestrel Rare Books Library, fourteen years in, who inherited FolioMind's rollout from a retired colleague.
3 · THE HABIT
What did Ines stop doing because FolioMind worked?
Tap to flip
ANSWER
Skimming a junior cataloger's finished entry against the original folder before filing it. By month six she'd stopped entirely.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Trusting junior staff's FolioMind-assisted entries fully, versus taking every large folder back to personally reread cover to cover.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Scoping FolioMind's context window to 8,000 tokens at launch to keep each call fast and cheap, back when nearly every folder ran under twenty pages.
6 · THE NUMBER
Fill in the blank: of 30 large folders Ines personally reread, ___ percent had a real cross-page inconsistency FolioMind had missed.
Tap to flip
ANSWER
40 percent (12 of 30), on folders the chunked 8,000-token version had already drafted entries for.
7 · THE REPLAY
Same 30 folders, 128,000-token window live. What changes?
Tap to flip
ANSWER
29 of 30 correctly flag the cross-page issue the chunked model missed, for only five extra minutes per folder.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Vosberg Port Authority's CargoLens, a workaround flip: Grace's team keeps a private hand-copied spreadsheet to track a container's weight across chunks the model itself couldn't compare.
Check yourself Score: 0 / 0
Multiple choice
1. Why did FolioMind's chunked version miss the contradiction between a 1932 letter and a 1935 event in the same folder?
A. The model wasn't trained on historical documents.
B. The scans of the two pages were low quality.
C. The two pages landed in different chunks, and no single call ever saw both at once.
D. Junior catalogers didn't read the draft entries closely enough.
Show hint
Look at the "identify the flip" step and the knowledge spark about chunked models.
Show answer
C. Each chunk was summarized accurately on its own; the contradiction only exists in the comparison between two chunks, which the 8,000-token window never held at the same time.
True or false
2. True or false: FolioMind's chunk summaries were factually wrong about the pages they actually processed.
True
False
Show hint
Think about what the direct answer says a small context window does versus doesn't do.
Show answer
False. Each individual chunk summary was accurate. The failure was never within a single chunk, it was in the comparison across chunks the model never got to make.
Fill in the blank
3. Fill in the blank: FolioMind's original context window was ___ tokens, roughly ___ pages of scanned correspondence.
Show hint
Look at "the choice I would take back."
Show answer
8,000 tokens; about 11 pages. A reasonable size when most folders ran under twenty pages, and too small once forty-to-eighty-page folders became the norm.
Short answer, where it wouldn't matter
4. Name a document type at Kestrel where the small context window was never actually a problem, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Short acquisition memos, almost always under ten pages, well within the original 8,000-token window, with no cross-page comparison ever needed.
Short answer, apply it yourself
5. Think of a long document you've had to work through yourself, a contract, a syllabus, a lease. What's one thing you'd only catch by comparing an early part against a later part?
Show hint
Think about a definition, a date, or a number that appears more than once.
Show answer
Model answer: A lease that defines "the premises" narrowly on page one, then refers to a shared storage area as included on page eleven. Only reading both at once reveals the contradiction.
Short answer, work the number
6. If only 10 percent of Kestrel's folders ran over forty pages instead of a larger share, would the context-window upgrade still be worth its extra cost?
Show hint
Weigh the per-call cost increase against how often a long folder actually comes through.
Show answer
Model answer: Likely still yes, but the case gets weaker. At 10 percent of volume, routing only those folders to the larger window, instead of upgrading every call, would capture most of the benefit at a fraction of the added cost.
Before you close the answer
Why this works
Tests whether you can translate a technical spec, context window size, into a concrete product capability, catching a cross-reference, instead of reciting a definition.
Follow-up traps
"Couldn't you just have the chunked model pass a running summary forward to the next chunk?" Response: it helps, but a running summary is itself a lossy compression, it can drop the exact detail, like a name's precise spelling, that the next chunk would need to catch a contradiction against.
"Isn't a bigger context window just strictly better, so why not always use it?" Response: no, it costs more and runs slower per call, and most of Kestrel's real documents, the short memos, never needed it in the first place.
If pressed
The 128,000-token upgrade still has its own ceiling, around 90 pages of scanned correspondence at Kestrel's OCR density, so the library flags any folder near that edge for a manual page-count check before trusting a single-call read.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.