The Decoder/What is a context window?
The Decoder • numbers checked 30 September 2026What is a context window?
A context window is the most text an AI model can hold in view at once, counted in tokens: your messages, its replies, any pasted files and its hidden instructions all have to fit inside it.

The context window is the model's working desk. Everything it can see while writing a reply has to be on that desk: the conversation so far, the document you pasted, the instructions the vendor added, and the reply itself. The desk has a fixed size, measured in tokens. When a conversation grows past it, the oldest pages fall off the edge, which is why a long chat starts forgetting what you said at the beginning.
Key facts
- 1,000,000tokens in the current Claude, GPT and Gemini flagship windows, about 555,000 to 750,000 words
- 200,000tokens in the smaller and older models, about 150,000 words
- 128,000tokens the largest windows can write back in one reply
- 0memory the model keeps between messages: the whole conversation is sent again every turn
- Anthropic models overview (1M default, 200k on Haiku 4.5, 128k output), OpenAI GPT-6 Astra model page (1,050,000), Google Gemini 3.1 Pro and 3.8 Flash model pages (1,048,576); checked 30 September 2026
Why does ChatGPT forget what I told it earlier?
Because it never remembered in the first place. A chat app sends the entire conversation to the model on every message. The model reads all of it, writes one reply, and keeps nothing. On the next message the app sends everything again, including the new reply. That works until the pile is bigger than the window. Then the app has to drop something, and it drops the oldest part. Your first instruction is gone, not because the model forgot it, but because it was never sent.
The memory features some apps now offer are a separate notebook the app keeps for you and quietly pastes back in. They help with facts about you. They do not make the window bigger.
| Model | Context window (tokens) | Max reply (tokens) | Roughly, in words |
|---|---|---|---|
| Claude Opus 5.5, Sonnet 5.5 | 1,000,000 | 128,000 | ~555,000 |
| Claude Haiku 4.5 | 200,000 | 64,000 | ~150,000 |
| OpenAI GPT-6 Astra (API) | 1,050,000 | 128,000 | ~750,000 |
| OpenAI GPT-5 (API) | 400,000 | 128,000 | ~300,000 |
| Google Gemini 3.1 Pro, 3.8 Flash | 1,048,576 | 65,536 | ~750,000 |
| DeepSeek V4 Pro, Flash | 1,000,000 | 384,000 | ~750,000 |
How long can a conversation be before the AI forgets?
Longer than most people will ever reach on the flagship models, and shorter than you would think on the small ones. A 1 million token window holds roughly 555,000 to 750,000 words depending on the vendor’s tokenizer, which is six or seven novels. A 200,000 token window holds about 150,000 words, a long novel. In practice the wall arrives sooner, because every pasted PDF, every long reply and the vendor’s own hidden instructions all count against the same total, and because the chat apps keep a working window smaller than the model’s maximum on the cheaper plans.
Quality also fades before the hard limit. Models reason best about what is near the start and near the end of the window and skim the middle. A fresh chat with a short summary of what matters usually beats a 300 message thread.
What is a 1 million token context window good for?
Whole things. A full codebase, a year of meeting notes, a long contract with all of its exhibits, a book manuscript with the editor’s letter. The point of a large window is that the model can hold the entire object and answer questions that need two distant parts of it at once, instead of you deciding which pages to paste. It is the cheap alternative to building a retrieval system for a single document set you will use a few times.
The cost is real, though. The model reads every token in the window on every turn, so a chat with a 400,000 token document loaded costs that much input each time you ask a follow up.
Does uploading a PDF use up the context window?
Yes, the text of it does. A 10 page PDF is around 5,000 words, so between 7,000 and 9,000 tokens depending on the model, and it sits in the window for the rest of the conversation. Images and scanned pages cost more, because they are handled as pictures. If you only need one chapter, paste the chapter. If you need to ask about a whole shelf of documents over weeks, that is the job for RAG, which fetches the right pages per question instead of loading everything.
Comments
No comments yet. Yours can be the first.