appzasGamesToolsDevLearnReferenceNetworkTimeCalculatorsCompareLLM prices

How LLM API pricing works

Lesson 3 of 412 minBeginner

Input vs output prices, cached input, the batch discount and a worked example.

Providers publish prices per million tokens, in US dollars. A model listed at "$2 per million input tokens" costs $2 to process one million input tokens, or $0.002 for 1,000.

Input and output are priced differently

Output tokens cost several times more than input tokens. A model priced at $2 input and $10 output charges five times more for what it writes than for what it reads. Long answers cost more than long prompts.

Other price components

  • Cached input (prompt caching): if a repeated start of your prompt was processed recently, reading it again is billed at a fraction of the input price. Useful for long, fixed instructions or documents.
  • Batch API: asynchronous jobs that can wait are discounted. Anthropic documents a 50% discount on input and output.
  • Long prompts: some models, such as Gemini Pro, charge more once a prompt passes 200,000 tokens.
  • Tools: web search and similar server-side tools may add fees.

Worked example

A support assistant handles 100,000 requests a month. Each request sends 2,000 input tokens and gets 500 output tokens. With a model priced at $2 input and $10 output per million tokens:

per request = (2,000 × $2 + 500 × $10) / 1,000,000
            = $0.009
per month   = 100,000 × $0.009 = $900

Run the same job offline in batch at a 50% discount and it is $450. A much smaller model priced around $0.05 and $0.40 per million tokens would cost about $30 for the same volume. Whether the cheaper one is good enough is a quality question, not a pricing one.

CarefulPrices change often. The numbers here are examples from our comparison dated 2026-10-04. Always check the provider's pricing page.

Test yourself

Answer all the questions, then check them. Finish with every answer right to mark the lesson as done.

1. A model costs $3 per million input tokens. How much for 500,000 input tokens?
2. Why does a long answer cost more than a long prompt of the same length?
3. Using the worked example, what is the monthly cost at $2 / $10 per million with 100,000 requests of 2,000 in and 500 out?
4. When does the Batch API discount make sense?

Key terms