Context window

Definition
The amount of text, prompt, conversation and documents together, an AI model can hold in mind at once, measured in tokens.

Why it matters

When a task outgrows the window, the oldest material drops out of view and the model carries on as if it had never seen it. Even inside the window, models often use the start and end of a long input better than the middle, so a fact buried halfway can be missed.

Cost matters too. Providers charge per token on each call, so a bloated context makes the same answer slower and dearer. A bigger window is not a reason to fill it.

How to apply it

  • Estimate what a task needs before loading everything available.
  • Paste the relevant sections of a document, not the whole file, for a narrow question.
  • Put the most important instructions at the start or the end of a long prompt.
  • Summarise or drop earlier parts of a long conversation once they stop mattering, or start a fresh chat with a short summary.
  • Treat a model that suddenly forgets something said earlier as a sign the window has filled.
  • Fetch reference material on demand with retrieval instead of carrying it in every call.

What it is

A context window is a model's working memory for one task. It holds everything the model can consider at once: the instructions, any documents pasted in, the conversation so far and the answer it is writing. It is measured in tokens, small chunks of text. In English a token is roughly three-quarters of a word, so 100,000 tokens is around 75,000 words. Modern models range from tens of thousands of tokens to more than a million.

Chat tools resend the whole conversation with each message. A long chat therefore fills the window message by message.

Common mistakes

  • Assuming a model remembers every earlier chat. Each session starts with an empty window unless a tool stores memory.
  • Forgetting that the answer also uses window space.
  • Trusting that a fact is used simply because it fits.
Worked example

Suppose a growth agency of seven people holds a weekly call with each of its twelve clients. Twelve full transcripts pasted into one chat would fill the window before the question even arrives, so the team starts from Fathom, which records each call and generates a summary and action items. The lead then pastes only the twelve summaries, about 300 words each, into a single chat and asks which clients mention churn risk. The answer covers all twelve. Twelve hour-long transcripts would run to well over 100,000 words, far more than one chat handles comfortably. Each claim in the reply is checked against its own call before it goes into a client report, so the shorter input also keeps the answer accountable.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Context engineering

    Deciding what goes into the window.

  2. Article

    Prompt engineering

    Writing the instruction inside it.

  3. Article

    Hallucination

    A failure a crowded or empty window can cause.