Skip to content
aviral gupta

// Lesson 1 of 3 · ~20 min · Beginner

Tokens and context windows

After this lesson you can work out what a model can “see” on any turn of a conversation, and plan a long chat so the important parts never fall out.

You will be able to

  • Name everything that counts toward the context window
  • Calculate how input grows turn by turn
  • Plan a long-document chat that leaves room for the answer
  1. Warm-up · Activity 1 of 7

    Before a language model works with your prompt, the text is split into…

  2. Predict · Activity 2 of 7

    You paste a 300-page contract into a chat that has a long system prompt, and ask for a summary. Which of these count toward the context window of that request? Pick all that apply.

    Select all that apply.

  3. Practice · Activity 3 of 7

    Worked example. Turn 1: you send 2,000 tokens and the model replies with 500. Turn 2: you send 300 tokens. Ignoring caching and trimming, roughly how many input tokens does the model receive on turn 2?

  4. Practice · Activity 4 of 7

    Your turn, with the support removed. Continue the same chat: the model replied to turn 2 with 400 tokens, and on turn 3 you send 200. How many input tokens now?

    tokens
  5. Practice · Activity 5 of 7

    Match each term to what it means.

  6. Brain teaser · Activity 6 of 7

    Brain teaser. A model has a 1,000,000-token context window. You load a 900,000-token codebase and ask for a 150,000-token rewrite in a single reply. What happens?

  7. Apply · Activity 7 of 7

    Mini-task. You want to ask many questions about a long report in one chat. Write a short plan (4–6 lines) for keeping the chat useful from the first question to the last.

    Check your work against this list

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

The window is the model’s working memory

A context window is all the text a model can reference while it generates a response, and that includes the response itself. It is working memory, not long-term memory: nothing outside the window exists for the model on that request.

Conversations accumulate

In a chat, every earlier message and reply is sent again with each new turn, so input grows as the conversation goes on. When it approaches the limit, the application has to drop, trim or summarise something. Better that you decide what than leave it to chance.

More is not always better

Accuracy and recall can degrade as the token count grows, which Anthropic’s documentation calls “context rot”. Curating what goes into the window matters as much as how much fits. Token counts differ between languages and between models, so measure with the provider’s token counter for the model you will use instead of guessing from word counts.

Sources

Last reviewed September 28, 2026