Estimate monthly and annual API spend across 18 OpenAI, Anthropic, Google, Meta and DeepSeek models — before the invoice arrives.
| Model | In $/1M | Out $/1M | Per call | Per month | Per year |
|---|---|---|---|---|---|
| GPT-4.1 Nano OpenAI | $0.1 | $0.4 | $0.0003 | $30.00 | $360 |
| Gemini 2.0 Flash Google | $0.1 | $0.4 | $0.0003 | $30.00 | $360 |
| Llama 4 Scout Meta | $0.1 | $0.4 | $0.0003 | $30.00 | $360 |
| GPT-4o Mini OpenAI | $0.15 | $0.6 | $0.0004 | $45.00 | $540 |
| Gemini 2.5 Flash Google | $0.15 | $0.6 | $0.0004 | $45.00 | $540 |
| Llama 4 Maverick Meta | $0.2 | $0.8 | $0.0006 | $60.00 | $720 |
| DeepSeek Chat DeepSeek | $0.27 | $1.1 | $0.0008 | $82.00 | $984 |
| GPT-4.1 Mini OpenAI | $0.4 | $1.6 | $0.0012 | $120 | $1440 |
| DeepSeek Reasoner DeepSeek | $0.55 | $2.19 | $0.0016 | $165 | $1974 |
| Claude 3.5 Haiku Anthropic | $0.8 | $4 | $0.0028 | $280 | $3360 |
| o4-mini OpenAI | $1.5 | $6 | $0.0045 | $450 | $5400 |
| GPT-4.1 OpenAI | $2 | $8 | $0.0060 | $600 | $7200 |
| Gemini 2.5 Pro Google | $1.25 | $10 | $0.0063 | $625 | $7500 |
| GPT-4o OpenAI | $2.5 | $10 | $0.0075 | $750 | $9000 |
| Claude 4 Sonnet Anthropic | $3 | $15 | $0.0105 | $1050 | $12,600 |
| Claude 3.7 Sonnet Anthropic | $3 | $15 | $0.0105 | $1050 | $12,600 |
| o3 OpenAI | $10 | $40 | $0.0300 | $3000 | $36,000 |
| Claude 4 Opus Anthropic | $15 | $75 | $0.0525 | $5250 | $63,000 |
LLM API prices are per 1M tokens as of June 2026 and may change without notice. Token estimates use model-specific chars-per-token ratios and may vary ±10-20% from actual counts.
An LLM API bill is two separate meters, not one. Input tokens and output tokens are priced independently, and output is typically 3–5× the input rate — so a model that looks cheap on its headline number can lose once your completions get long. Monthly spend is calls × ((input ÷ 1M × inPrice) + (output ÷ 1M × outPrice)). Because every provider sets both rates independently, the ranking flips depending on your input-to-output ratio: a retrieval pipeline that stuffs 8,000 tokens of context into every call is dominated by input price, while a code-generation workload is dominated by output price.
Output tokens usually cost more per million than input tokens. That is why the cheapest model for a chat product is often not the cheapest model for a summarisation or code-generation pipeline.
Priya runs a support chatbot: 100,000 calls a month, 1,000 input tokens each (system prompt + retrieved help-centre articles), 500 output tokens for the reply.
Why did the cow cross the road?
No signups, no data sold. The core of every tool is free forever — the optional Pro plan adds batch processing, unlimited downloads, white-label exports and an ad-free experience.
☕Support me on Ko-fi— keep tools free100% of proceeds go towards hosting & building more free tools.