Describe the difference between a context window and a model's memory.
Rivenote joins a video call, listens the whole way through, and writes up the notes and the action items so nobody has to. Tegan Sculthorpe built its live in call feature, Ask Rivenote, so a team could ask it a question mid meeting and get an answer pulled from everything said so far. This is the one Tuesday call that taught her the feature had two different failure modes wearing the same words, "I don't see a record of that."
- Treat context window and memory as two different systems, never as one feature with two settings.Why: a hard token limit that resets every call and a deliberately saved fact that survives a new one are not the same mechanism just because both feel like "remembering."
- Test any "it forgot" symptom in a brand new session with no shared history, not a rephrase in the same call.Why: only that test tells you whether you're looking at a window problem or a memory problem, and the two need completely different fixes.
- Never assume a large context window means nothing gets dropped on a long call.Why: Bramknell's ceiling was 128,000 tokens, and the running transcript plus every earlier live answer crossed it by 12:50 p.m., under four hours in.
- Make the drop visible instead of silent.Why: a warning at the ceiling gives a person the chance to catch what's about to fall out, instead of finding out thirty four minutes into a re debate.
- Widen what a memory system is allowed to capture, not just how big the window is.Why: only 61% of decisions spoken out loud on long calls ever made it into memory at all, because the save step needed an exact trigger phrase.
- Leave short calls alone.Why: under ninety minutes, the whole conversation stays inside the window the entire time, so tightening the save rule there just adds noise for a problem that never happens.
How to answer this, stage by stage
Nobody is grading whether you can recite the word "token." They're grading whether you can tell two things apart that a product's own marketing happily blurs together.
Let's learn
Rivenote is a program that sits in on a video call, listens the whole time, and hands the team a clean set of notes and a list of who owes what once the call ends. Before it, someone on the team had to type notes by hand and email a summary afterward, and it usually took about twenty five minutes to write up properly, longer if the meeting ran long or the notetaker had to step out.
Once Rivenote shipped, that twenty five minutes dropped close to zero. The bigger change was Ask Rivenote, the live in call assistant Tegan Sculthorpe built. Type a question into the call at any point, get an answer pulled from everything said so far that meeting. For a year and a half, it was the most used thing Rivenote had ever shipped. Teams stopped keeping their own scratch notes during a call entirely, because asking the assistant was faster than scrolling back through a shared doc.
Then came a Tuesday. Bramknell Logistics, a regional freight company, ran its quarterly planning call as one continuous video meeting, five hours and ten minutes, no restart, Rivenote present the whole way through. Forty minutes in, the team made a real call: move the Ohio distribution contract to Vantage Line, and stop renewing with Bell Fleet. Nobody used a formal phrase for it. Someone just said, "yeah, let's go with Vantage then, we're done with Bell Fleet," and the room moved on.
Here is the turn. The extra minutes it took the assistant to answer wrong were never the real problem. The real problem is that "I don't see a record of this" sounds exactly the same whether it means "this genuinely hasn't been decided yet" or "this fell out of what I can currently see." A manager who had stepped away and rejoined at 1:47 asked Ask Rivenote whether Ohio had already been settled. It answered, "I don't see a decision on this yet in what I have so far." Nobody in the room could tell that was a blind spot instead of the truth. So they spent the next thirty four minutes re arguing a decision that had already been made, and the regional director very nearly picked up the phone to reopen the conversation with Bell Fleet's account rep before someone remembered it had already been settled that morning.
What it costs at its worst: twelve people sat through thirty four minutes of that re debate, about six and a half hours of combined time on one call. And it could have been worse. If the director's call to Bell Fleet had gone through, Bramknell would have been renegotiating with a vendor it had verbally dropped that same morning, in front of the vendor it had verbally chosen instead. A confidently worded "I don't see a decision on this" from a tool people trust is more dangerous than silence, because silence makes a person ask around. A wrong answer that sounds sure of itself gets believed.
What I would leave alone: any call under about ninety minutes. The entire conversation stays inside the token window the whole time on a call that short, so the assistant never needs memory to answer a question about something said earlier that same meeting. Tightening the trigger phrase rule there would only add noise to a problem that structurally can't happen yet.
The lesson: a feature that quietly behaves differently depending on how long the call has been running is a feature nobody actually designed on purpose. We built one mechanism, a window, and let people believe it was two, a window and a memory, because we never made the line between them visible to anyone using it.
Now here is the same thing as a story
Say the short version out loud in an interview. Read this one when you want to feel exactly how the same four words, "I don't see a record," can mean two completely different things.
Tegan Sculthorpe spent four years as an operations analyst before she ever wrote a line of product code, sitting in on client planning calls with a paper legal pad, writing down every decision the second someone said it out loud. She'd learned the hard way that a decision left unwritten for even an hour was a decision that would get re argued in three weeks by someone who genuinely didn't remember agreeing to it. She joined Rivenote two years ago to build, in software, the thing she used to be by hand.
She built Ask Rivenote herself. For a year and a half it was the best loved thing the company had shipped. Teams used it constantly, live, mid call: did we already cover the budget line, what number did Priya give earlier, who's supposed to own the intake form. It got good enough, fast enough, that a habit formed nobody had planned for. First, people still kept a personal scratch doc open during calls, just in case. Then they stopped opening it. Then, on most teams, the doc quietly stopped existing at all. Why keep your own notes when the assistant already has better ones?
Bramknell Logistics' quarterly planning call ran long that Tuesday, the way it always did, one continuous session from nine in the morning. Forty minutes in, the regional director made the Ohio call out loud, in the middle of an unrelated sentence about a warehouse lease, and the room moved straight on to the next topic. Nobody wrote it down. Nobody needed to. Ask Rivenote had it.
Somewhere around 12:50, with nobody watching for it, the running transcript, plus every earlier answer Ask Rivenote had given that morning stacked into its own chat history, crossed 128,000 tokens. Nothing announced it. No banner. No dropped notification. The pipeline just started sending the model the newest chunk that fit, and quietly left the oldest part behind, the part that held the Ohio decision.
At 1:47, a warehouse ops manager who'd stepped out to take a call rejoined and asked, out loud, into the meeting, "did we already decide on Ohio or are we still going back and forth." Ask Rivenote answered instantly and confidently: "I don't see a decision on this yet in what I have so far." The manager had no way to know that sentence meant "it fell off the back of what I can currently see" instead of "this genuinely hasn't happened." Neither did anyone else in the room. So they started over. Thirty four minutes of re arguing vendor pros and cons that had already been settled, until the regional director, halfway through drafting a message to Bell Fleet's account rep to reopen talks, stopped and said, "wait, didn't we already do this?"
Tegan opened the incident the next morning expecting a clean, single answer: the window had filled, case closed, ship a bigger one. She almost stopped there. What made her keep going was a nagging question: even inside the window, before it filled, the live save step should have been catching real decisions and writing them to memory on the spot, the same way it always had on shorter calls. Had it caught this one, or not.
She checked the boring explanations first. The transcript around 9:38 was clean, correctly spelled, nothing garbled by the transcription engine. Bramknell's workspace had the memory feature switched on, confirmed, because Ask Rivenote had correctly pulled a fact from a meeting two weeks earlier for a different question that same morning, before the window filled. So the plumbing worked. Something behavioral, not broken, was actually going on.
That left three real candidates. One: the live window genuinely explained the 1:47 answer, since the ceiling crossed at 12:50, before the question was ever asked. Two: the decision had been saved to memory just fine, and the manager's exact phrasing simply didn't match how it got stored, a plain search problem. Three: the decision was never captured anywhere at all, because the live save step only fires on an exact trigger phrase, and "yeah, let's go with Vantage, we're done with Bell Fleet" never matched the pattern it was listening for.
The evidence test was the one that actually mattered. The following Monday, Tegan opened a brand new session, no shared history, no continuation of Tuesday's call at all, and asked about the Ohio decision four different ways: plainly, by vendor name, by contract number, by date. All four came back empty. That single test ruled out candidate two on the spot. A retrieval mismatch would have surfaced under at least one phrasing. Nothing did, which meant the fact was never written down in the first place. Candidate three, confirmed. Candidate one still stood on its own, proven by nothing more exotic than the clock: the ceiling crossed at 12:50, the question landed at 1:47, an hour later.
The decision Tegan would take back sat in a design review nine months before launch, the one where the team built the trigger phrase gate for the live save step, because ungated extraction in early testing had turned every stray suggestion into a saved "decision," and the feature drowned in noise nobody trusted. Gating it was the right call at the time. Most decisions on most calls did get phrased the way the gate expected, because most calls were short enough that people spoke a little more formally, a little more on the record.
Run the same Tuesday again, with one change: whenever the running transcript nears the ceiling, before anything drops, a second pass, less strict than the live gate, sweeps the part about to fall out and writes any sentence that reads like a decision into memory on its own judgment, flagged as "worth a second look" rather than treated as certain. Same five hour call. At 1:47, the ops manager asks the same question. Ask Rivenote answers in about four seconds: "yes, decided 9:38 this morning, Ohio moves to Vantage Line." No re debate. No half drafted message to Bell Fleet.
What I'd tell myself, sitting in that design review nine months earlier: I built a gate to keep the feature honest, and never once asked what would back it up on the one day the gate itself would miss something real.
TRACE, so "it forgot" never gets said again
Not a way to prove the window was the villain. TRACE is what actually separates two structurally different failures that happen to share one apologetic sentence.
Three things worth saying plainly, since interviewers push here. Tegan considered a second fix before the one she shipped: just make the trigger phrase gate stricter, so it only ever caught unmistakably formal decisions, and train the team to speak more formally when something mattered. She rejected it, because it puts the burden on twelve people changing how they talk in a five hour call, instead of on the product catching real language the way people actually use it. The AI specific failure worth naming by name is silent context truncation: a running conversation quietly outgrows what a model can see, with no error and no visible sign, and the fix is not a bigger window, since any window eventually fills on a long enough call, it's a visible warning plus a second, less strict capture pass that runs before anything drops. The trade off, accepted on purpose: that second pass is deliberately looser than the live gate, which means it will occasionally flag something as a "decision" that wasn't quite one, a small accuracy cost, in exchange for never again saying "I don't see a record" about something that was, in fact, decided out loud in the same room an hour earlier.
And if you want to be sure it really works, try it somewhere else
Same five letters, a veterinary clinic instead of a freight company, and this time the whole call barely runs twenty minutes. The confusion still shows up, because it was never really about how long the call was.
Thistlefield builds an ambient scribe for veterinary clinics: it listens during an appointment and drafts the visit note, and it's meant to carry a patient's history forward automatically, visit to visit, weeks or months apart. Griet Kinnaird runs product there, and hit a version of Tegan's exact confusion eleven months into the job, on a case where the "call" itself was never anywhere close to any token ceiling.
Three weeks earlier, a dog's owner had mentioned, mid sentence, that the dog had a mild reaction to amoxicillin once as a puppy. The vet was interrupted before finishing that part of the note, and it never made it into the signed, finalized chart. Today, a different vet at the same clinic, treating an ear infection, asked Thistlefield's assistant whether this patient had any known drug reactions. It said no record found. The vet nearly prescribed amoxicillin.
Mapped onto TRACE: the timeline shows the fact spoken at minute fourteen of visit one, and the gap discovered at minute six of visit two, three weeks later. The recut is the same clean split, context window against memory, except here the window was never the issue at all, since a twenty minute visit never comes close to any ceiling. The assumption Griet ruled out first was that the raw audio simply didn't exist anymore; it did, sitting in a searchable archive nobody's live workflow ever queried. The cause candidates: a note taking error, a memory bug, or memory doing exactly what it was built to do, reading only the signed record. The evidence test gave the same shape of answer as Bramknell's: pulling the raw transcript directly, bypassing the assistant's memory path entirely, found the sentence in seconds, proving the fact was never lost to a technical failure. It was excluded on purpose by a rule nobody had revisited since the day it was written.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the reframe: it's not "did it forget," it's "was this ever saved on purpose, and does the fact survive a brand new session."
Cost: no time to trace a real incident. Ask one question instead: last time something "wasn't there," did a fresh, unrelated session ever get tried, or only a rephrase in the same conversation.
The model got better, for real: say the next model ships with a context window ten times the size. The live truncation risk on a five hour call mostly disappears. The memory capture gap doesn't shrink at all, because it was never about the window's size in the first place.
Where people run it wrong.
They treat a big context window as proof that nothing will ever get dropped, instead of proof that it takes longer to happen.
They test "does it remember" by asking again in the same conversation, and mistake context window persistence for real memory.
They fix a memory gap by shipping a bigger window, when the actual problem is what gets written down in the first place.
How to use it live. When an interviewer asks this cold, buy two seconds by asking one thing back: "are we talking about the same conversation, or a fact coming back in a brand new one?" That question alone is usually exactly what a question shaped like this one is listening for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't memory just a context window with a longer reach?" Response: no. A context window resets to empty at the start of every new session. Memory, when it's real, doesn't; it's a separate store that gets loaded back in on purpose, on any call, old or new.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on The AI literacy baseline every PM needs
- #1 Explain what a token is and why a PM should care about it.
- #3 What is the practical difference between prompting, RAG and fine-tuning for a product decision?
- #4 Explain hallucination in one paragraph a sales team could repeat accurately.
- #5 What does temperature control and when would you lower it in a product?
- #6 Describe what an embedding is and one product feature it makes possible.
- #7 Explain the difference between latency and throughput and which one your users feel.