How LLM API pricing works
Input vs output prices, cached input, the batch discount and a worked example.
Providers publish prices per million tokens, in US dollars. A model listed at "$2 per million input tokens" costs $2 to process one million input tokens, or $0.002 for 1,000.
Input and output are priced differently
Output tokens cost several times more than input tokens. A model priced at $2 input and $10 output charges five times more for what it writes than for what it reads. Long answers cost more than long prompts.
Other price components
- Cached input (prompt caching): if a repeated start of your prompt was processed recently, reading it again is billed at a fraction of the input price. Useful for long, fixed instructions or documents.
- Batch API: asynchronous jobs that can wait are discounted. Anthropic documents a 50% discount on input and output.
- Long prompts: some models, such as Gemini Pro, charge more once a prompt passes 200,000 tokens.
- Tools: web search and similar server-side tools may add fees.
Worked example
A support assistant handles 100,000 requests a month. Each request sends 2,000 input tokens and gets 500 output tokens. With a model priced at $2 input and $10 output per million tokens:
per request = (2,000 × $2 + 500 × $10) / 1,000,000
= $0.009
per month = 100,000 × $0.009 = $900Run the same job offline in batch at a 50% discount and it is $450. A much smaller model priced around $0.05 and $0.40 per million tokens would cost about $30 for the same volume. Whether the cheaper one is good enough is a quality question, not a pricing one.
CarefulPrices change often. The numbers here are examples from our comparison dated 2026-10-04. Always check the provider's pricing page.
Test yourself
Answer all the questions, then check them. Finish with every answer right to mark the lesson as done.