Estimate your bill before you build
A repeatable method to budget an LLM feature and keep costs under control.
You can estimate an LLM bill with five numbers and a calculator. Do it before you write the feature.
Five steps
- Measure a typical request. Count input tokens (instructions, documents, history) and output tokens for a realistic example. Providers offer token counters, and ours has rules of thumb.
- Estimate volume. Requests per day times 30. Be pessimistic.
- Account for caching. What share of the input is the same long prefix each time? That part can be billed at the cached rate.
- Add a margin for retries. Failed or repeated calls cost money. Ten to twenty percent is a reasonable start.
- Compare models. Enter the numbers in the calculator and look at the cheapest few. Test whether they are good enough on your own examples.
Keep it under control
- Set a maximum output length for every request.
- Trim conversation history and send only what the model needs.
- Use a smaller model for easy tasks and a stronger one only when needed.
- Set budget alerts and spending limits in your provider account.
- Handle rate limit errors (HTTP 429) with exponential backoff instead of tight retry loops.
Common surprises
- Chat history grows each turn, so cost per message rises over a conversation.
- Long system prompts are paid on every request unless cached.
- An agent that loops can use thousands of tokens before a human notices.
NoteCheapest is not best. Evaluate quality on your own data before you choose a model for cost alone.
Test yourself
Answer all the questions, then check them. Finish with every answer right to mark the lesson as done.