Per-request and monthly API costs across 17 models, side by side.
| Model | In $/1M | Out $/1M | Per request | Per day | Per month |
|---|---|---|---|---|---|
| DeepSeek V4 Flash (off-peak)cheapest | $0.21 | $0.63 | $9.2e-4 | $0.09 | $3 |
| GPT-5.6 Luna | $0.2 | $1.2 | $1.4e-3 | $0.14 | $4 |
| DeepSeek V4 Flash (peak) | $0.42 | $1.25 | $1.8e-3 | $0.18 | $6 |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | $2.6e-3 | $0.26 | $8 |
| DeepSeek V4 Pro (off-peak) | $0.63 | $1.88 | $2.8e-3 | $0.28 | $8 |
| DeepSeek V4 Pro (peak) | $1.25 | $3.75 | $5.5e-3 | $0.55 | $17 |
| Claude Haiku 4.5 | $1 | $5 | $6.0e-3 | $0.6 | $18 |
| GLM-5.2 | $1.4 | $4.4 | $6.3e-3 | $0.63 | $19 |
| Qwen3.8 Max | $2 | $6 | $8.8e-3 | $0.88 | $27 |
| Gemini 3.6 Flash | $1.5 | $7.5 | $9.0e-3 | $0.9 | $27 |
| Claude Sonnet 5 | $2 | $10 | $0.012 | $1.2 | $37 |
| GPT-5.6 Terra | $2 | $12 | $0.0136 | $1.36 | $41 |
| Gemini 3.1 Pro | $2 | $12 | $0.0136 | $1.36 | $41 |
| Kimi K3 | $3 | $15 | $0.018 | $1.8 | $55 |
| Claude Opus 5 | $5 | $25 | $0.03 | $3 | $91 |
| GPT-5.6 Sol | $5 | $30 | $0.034 | $3.4 | $103 |
| Claude Fable 5 | $10 | $50 | $0.06 | $6 | $183 |
💰 Prices are uncached-input rates per 1M tokens (tokencost.app, checked 2026-08). Cached prompts run 50-90% cheaper on input; batch/async APIs roughly halve output costs. Monthly assumes 30.44 days.
Rows are sorted by projected monthly cost — the cheapest model for your exact token mix floats to the top with a green highlight. Input and output prices are uncached per-million rates from tokencost.app, re-checked 2026-08.
Output tokens cost 3-15× input tokens on most models, so a chatty assistant writing 2k tokens per call dwarfs a long prompt of cheap input. Trimming max_tokens is often the biggest single lever.
Cached prompts (same system prompt repeatedly) run 50-90% cheaper on input at OpenAI/Anthropic/DeepSeek. Batch APIs halve output costs when you can wait minutes instead of seconds.
A hobby app at 100 requests/day on a mid-tier model typically lands under $30/month. The same volume on frontier models can exceed $300 — the table makes the 10× gap visible before you commit.
The LLM API Cost Calculator lets you figure out llm cost calculatorinstantly, without reaching for a spreadsheet or doing the math by hand. Whether you're planning a budget, checking a loan, or working through homework, the tool applies the correct formula behind the scenes and shows the result the moment you enter your numbers.
Unlike a static chart or table, this calculator adapts to your exact inputs. You can adjust any value and see the outcome update in real time, which makes it easy to compare scenarios — for example, "what if the rate were 1% lower?" or "what if I paid an extra $50 a month?"
Common uses: people reach for this tool when they need to find a how much does gpt api cost per month, claude vs gpt api price comparison, llm api cost estimation tool, or token cost calculator openai anthropic.
Browser-based tools like this one have a few real advantages over installed software or manual methods:
The LLM API Cost Calculator is based on the following formula:
Cost per request = Tin ÷ 1,000,000 × Pin + Tout ÷ 1,000,000 × Pout Daily cost = requests/day × Cost per request Monthly cost = Daily cost × 30 With cached input: Cost = (Tin,cached × Pcache + Tin,fresh × Pin + Tout × Pout) ÷ 1,000,000
Variables: Tin = Input tokens per request (tokens) Tout = Output tokens per request (tokens) Pin = Input price ($ per 1M tokens) Pout = Output price ($ per 1M tokens) Tin,cached = Input tokens served from the cache (tokens) Tin,fresh = Input tokens billed at full price (tokens) Pcache = Cached-input price ($ per 1M tokens) requests/day = Requests sent per day Monthly cost = Projected cost for a 30-day month ($)
API pricing is quoted per million tokens, so divide each token count by 1,000,000 and multiply by its side's rate, then scale by request volume. Cached input is billed at a deep discount, so splitting input into cached and fresh portions captures that saving. Output tokens typically cost several times more than input, which is why long replies dominate the bill.
Worked example: Step 1: model prices $3 per 1M input and $15 per 1M output; a request uses 2,000 input and 800 output tokens. Step 2: input = 2,000 ÷ 1,000,000 × 3 = $0.006; output = 800 ÷ 1,000,000 × 15 = $0.012 → $0.018 per request. Step 3: at 500 requests/day: 500 × 0.018 = $9.00/day → × 30 = $270/month. Step 4: with half the input cached at $0.30 per 1M: (1,000 × 0.30 + 1,000 × 3 + 800 × 15) ÷ 1,000,000 = $0.0153 per request → 500 × 0.0153 × 30 = $229.50/month. Result: caching trims the monthly bill from $270 to about $229.50, a 15% saving.
More tools you might find useful