Context window
Why it matters
When a task outgrows the window, the oldest material drops out of view and the model carries on as if it had never seen it. Even inside the window, models often use the start and end of a long input better than the middle, so a fact buried halfway can be missed.
Cost matters too. Providers charge per token on each call, so a bloated context makes the same answer slower and dearer. A bigger window is not a reason to fill it.
How to apply it
- Estimate what a task needs before loading everything available.
- Paste the relevant sections of a document, not the whole file, for a narrow question.
- Put the most important instructions at the start or the end of a long prompt.
- Summarise or drop earlier parts of a long conversation once they stop mattering, or start a fresh chat with a short summary.
- Treat a model that suddenly forgets something said earlier as a sign the window has filled.
- Fetch reference material on demand with retrieval instead of carrying it in every call.
What it is
A context window is a model's working memory for one task. It holds everything the model can consider at once: the instructions, any documents pasted in, the conversation so far and the answer it is writing. It is measured in tokens, small chunks of text. In English a token is roughly three-quarters of a word, so 100,000 tokens is around 75,000 words. Modern models range from tens of thousands of tokens to more than a million.
Chat tools resend the whole conversation with each message. A long chat therefore fills the window message by message.
Common mistakes
- Assuming a model remembers every earlier chat. Each session starts with an empty window unless a tool stores memory.
- Forgetting that the answer also uses window space.
- Trusting that a fact is used simply because it fits.