appzasGamesToolsDevLearnReferenceNetworkTimeCalculatorsCompareLLM prices

Tokens and the context window

Lesson 2 of 410 minBeginner

What a token is, how many words fit in a context window and why the same text costs different amounts.

Models do not read letters or words. They read tokens: pieces of text that can be a whole short word, part of a longer one, a punctuation mark or a space.

A useful rule of thumb

For English text, one token is roughly 4 characters or about 0.75 of a word, so 1,000 tokens is around 750 words. The exact figure depends on the text and the model. Code, numbers and non-English languages usually need more tokens per word.

Different models count differently

Each model family has its own tokenizer, so the same text is a different number of tokens depending on the provider. Anthropic states that Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text. That is why our cost calculator has a switch for it.

The context window

The context window is the maximum number of tokens a model can handle in one request, counting your input and its output together. Larger windows let you include long documents or long histories. They also let you spend much more per request.

  • Input too long: the request is rejected or truncated.
  • A separate maximum output limit caps how long the answer can be.
  • Setting a sensible output limit protects you from unexpectedly long, costly replies.

A quick estimate

A 3,000-word article is about 4,000 tokens. Summarizing it into 300 words uses roughly 4,000 input tokens and 400 output tokens.

Test yourself

Answer all the questions, then check them. Finish with every answer right to mark the lesson as done.

1. About how many English words is 1,000 tokens?
2. What counts toward the context window?
3. Why can the same text cost different amounts at two providers?

Key terms