LLM bills are not mysterious — they are multiplication. requests × tokens × price. What separates a $200/month feature from a $20,000/month one is usually three decisions: output length, caching, and model routing. This handbook prices each one.
1. Tokens Are the Unit of Everything
Providers bill per million tokens, split into input and output. The rule of thumb for English is ~4 characters per token (¾ of a word); code tokenizes denser, and CJK text costs roughly 1 token per character — the same prompt in Chinese routinely costs 30%+ more than its English translation.
The asymmetry that drives budgets: output tokens cost 3–5× input tokens, because generation is serial while prompt processing is parallel. A feature that reads 2,000 tokens and answers 200 is cheap; the same feature answering 2,000 verbose tokens costs ten times more on the output side alone.
GPT Token CounterPaste the real prompt and measure tokens before you price anything — measured density beats guessed density every time.→2. Measure Token Density Before You Model Costs
Every cost model built on assumed token counts is fiction. The honest workflow: take 20 real production prompts, count their tokens, take the p50 and p95, and model both. Chat features cluster tightly; RAG features have fat tails — the p95 prompt with five retrieved chunks may cost 8× the p50.
// Monthly LLM budget model
cost = requests/day * 30.44
* ( in_tokens * price_in
+ out_tokens * price_out ) / 1_000_000
//
// Levers, in order of ROI:
// 1. shrink output tokens (output costs 3-5x input)
// 2. cache stable prefixes (50-90% off cached input)
// 3. batch/async APIs (~50% off, latency for discount)
// 4. right-size the model (route easy traffic to small)3. Context Caching: The 50–90% Discount Nobody Budgets
If your prompts share a stable prefix — system prompt, few-shot examples, tool definitions — prompt caching reuses the provider’s processed state. Cached input tokens are billed at 10–50% of list price, and time-to-first-token drops as a bonus. For an agent with an 8,000-token system prompt called 50,000 times a day, caching converts the dominant line item into a rounding error.
Design for the cache: put stable content first, volatile content last; keep the prefix byte-identical (whitespace changes bust the cache); and route long-document Q&A through cache-friendly prompt layouts instead of re-sending the document per turn.
4. Batch APIs and Model Routing: The Other Two Halves
Non-interactive work — evals, backfills, nightly summarization — belongs in batch APIs, where output runs at roughly half price in exchange for minutes-to-hours latency. Interactive traffic should never pay batch prices, and batch traffic should never pay interactive prices; mixing them is the most common unforced error in LLM budgeting.
Then route by difficulty: classify or extract with the small model, reserve the flagship for genuinely hard reasoning. A router that sends 80% of traffic to a model 10× cheaper cuts the blended bill by more than any prompt optimization ever will.
LLM API Cost CalculatorMonthly cost across models from your tokens-per-request and requests-per-day — compare the routing scenarios side by side.→5. Embeddings: The Line Item You Can Stop Worrying About
Embedding models price at $0.02–$0.13 per million tokens — a billion-token corpus costs tens of dollars to embed once. For most RAG systems the embedding API is the smallest line on the invoice; the real costs live in vector storage (proportional to dimensions) and retrieval compute. Budget embeddings once, then stop optimizing the wrong line.
Embedding Price CalculatorFirst-pass and monthly re-embedding costs across providers — sized against the storage bill it actually competes with.→The Budget Review, Monthly
- Recompute from measured p50/p95 tokens, never assumed ones.
- Check the output/input ratio — if it drifts above plan, output discipline slipped.
- Verify cache hit rate; a busted prefix is a silent 2–10× tax on input.
- Move everything non-interactive to batch; audit quarterly.
- Re-price against current list prices each quarter — LLM pricing falls fast enough that last quarter’s model may now be the expensive one.