Context & Memory Interview Questions (2026)
Covers the context window — the fixed budget every request shares — and agent memory, the separate system that makes anything persist beyond it. See also all interview topics. These assume you already know the concept — if this is unfamiliar, read the linked concept page first; the questions test judgment on top of the concept, not the concept itself.
Context Window
Your chat app has been running for hours. The model suddenly seems to forget something from early in the conversation. What's actually happening?
The context window filled up, and the earliest turns got pushed out to make room for new ones. The model isn't forgetting in any psychological sense — it simply never receives that text anymore. Nothing about a conversation is stored on the model's side between requests; if the application didn't resend it, it isn't there.
Your model now supports a context window ten times larger than before. Should you just retrieve everything relevant and stuff it all in, instead of carefully picking the best chunks?
Not automatically. Fitting is necessary but not sufficient — research on long-context use shows models use information near the start or end of a prompt more reliably than information buried in the middle, and every extra token still costs money and adds latency even when it fits comfortably. A bigger window raises the ceiling; it doesn't remove the reason to be selective about what you send.
A request errors out with "prompt too long" instead of just answering with less context. Is that a bug?
No — some systems deliberately fail loudly rather than silently drop part of the prompt, and that's the safer behavior. The alternative, quietly truncating, can drop an early instruction, a fact from a retrieved document, or a tool result the answer depended on — and the application keeps responding as if nothing were missing.
Your RAG pipeline retrieves 50 chunks for a question that only needed 3. What's the actual cost of over-retrieving, beyond it being wasteful?
Every retrieved chunk counts against the same context window budget as the question and the conversation history, so over-retrieving crowds out room for other things, costs more per request, and can bury the 3 chunks that actually mattered among 47 that didn't — which is exactly the situation the lost-in-the-middle effect makes worse.
Agent Memory
A conversation is still well within its context window, with nothing dropped yet. Does that count as the system having memory?
No. That's just the current request still holding everything so far in its budget — nothing has been deliberately decided as worth keeping. The moment the session ends, none of it persists unless a separate memory system explicitly saved it first. Memory is what makes something survive on purpose, not what happens to still fit right now.
A team builds a memory system that saves the full raw transcript of every conversation, indiscriminately, to be retrieved later. What's the risk?
It relocates RAG's own over-retrieval problem into memory. A future request can pull back too many saved memories, most irrelevant to the current question, crowding out the ones that actually matter — the same "more retrieved context isn't automatically a better context" problem, just sourced from past conversations instead of a document store.
An agent's memory includes something read from an untrusted document weeks ago, in a session that's long since ended. Why does that still matter today?
Because injected text saved into memory doesn't need its original source to still be around to fire again — a future request that retrieves that memory reintroduces the same untrusted content, now with no obvious link back to where it came from. Stored memory needs the same untrusted treatment on read as anything freshly fetched, not a pass because it already made it into storage.