llm pricing comparison
LLM Pricing Comparison 2026 — API Costs
2026 LLM pricing comparison for Claude API pricing, OpenAI, Gemini, and DeepSeek on one workload. Use the LLM cost calculator (token cost / OpenAI pricing calculator) for prompt-level bills; this page is the map. List rates below are OmniKit catalog v2026-09-11.
Last checked 2026-09-12
Editorial note
This brief models deepseek-flash and gpt-4o-mini and claude-sonnet-5 and gemini-3.8-flash on the same workload so finance and engineering can argue from one spreadsheet row. List prices change; the point is relative ordering and sensitivity to output length. Pair the mini calculator below with your own prompt logs, then export a fuller scenario from the LLM Cost Calculator before you commit spend. Updated 2026-09-12.
- DeepSeek V4.1 Flash typically sets the floor; Claude Sonnet 5 often sets the ceiling.
- GPT-4o Mini and Gemini 3.8 Flash land in the cheap/mid band depending on output length.
- Revisit prices monthly — list rates move (catalog v2026-09-11).
- Re-run after prompt changes — output length shifts monthly $ fast.
List prices (catalog v2026-09-11)
| Model | Input $/1M | Output $/1M | Notes |
|---|---|---|---|
| DeepSeek V4.1 Flash (deepseek-flash) | 0.30 | 1.20 | Peak cache-miss. Off-peak $0.15 / $0.60. |
| GPT-4o Mini | 0.15 | 0.60 | OpenAI list. Cached input can be lower on the vendor page. |
| Gemini 3.8 Flash | 0.75 | 3.75 | Intro list. Steps up to $1.50 / $7.50 after the intro window. |
| Claude Sonnet 5 | 2.00 | 10.00 | Anthropic list. Cache and batch discounts not applied here. |
Mini calculator
Estimate these models
What changed for operators
List prices move quickly. Cache discounts, batch APIs, gateways (OpenRouter, OpenCode Zen, Command Code), and open-weight self-hosting changed the break-even math versus 2024. Treat any static blog table as stale — re-run with your RPS. OmniKit stores peak or current intro rates so estimates lean conservative.
Methodology
This hub quotes OmniKit’s versioned catalog (apps/api model_pricing.v1.json, version 2026-09-11). Dual-rate models use the higher uncached or post-intro figure unless notes say the intro is still current. Gateways are compared after you run the calculator, not in this static table. We do not invent negotiated contract rates.
Worked example: 1M input + 200k output
One million input tokens plus 200k output, one shot (not monthly traffic): DeepSeek Flash peak: $0.30 + $0.24 = $0.54. GPT-4o Mini: $0.15 + $0.12 = $0.27. Gemini 3.8 Flash intro: $0.75 + $0.75 = $1.50. Claude Sonnet 5: $2.00 + $2.00 = $4.00. Mini wins this mix because output is short. Longer completions pull Gemini and Claude up faster.
Worked example: 10M input + 2M output
Ten million input plus two million output (a heavy chatbot day): DeepSeek Flash peak: $3.00 + $2.40 = $5.40. GPT-4o Mini: $1.50 + $1.20 = $2.70. Gemini 3.8 Flash intro: $7.50 + $7.50 = $15.00. Claude Sonnet 5: $20.00 + $20.00 = $40.00. Scale with requests per second in the calculator; do not multiply these one-shot totals by 30 unless that is actually your monthly token volume.
Worked example: output-heavy completions
Same 1M input but 1M output. DeepSeek Flash: $0.30 + $1.20 = $1.50. GPT-4o Mini: $0.15 + $0.60 = $0.75. Gemini intro: $0.75 + $3.75 = $4.50. Claude Sonnet 5: $2.00 + $10.00 = $12.00. Output price, not input, usually decides chatbot envelopes.
How to use this calculator pair
Start with equal prompts across models, note monthly $, then stress-test output length. If one model answers longer, cost can invert even when input rates look cheaper. Open the LLM cost calculator for token-level OpenAI, Claude, and Gemini API pricing, then check OpenRouter / Zen / Command Code on the same ids.
How OmniKit estimates these costs
OmniKit tokenizes representative prompts locally when possible and multiplies by versioned list prices (per million input/output tokens). Traffic is modeled from requests per second and days per month so you can compare vendors on the same workload. Cache write fees, batch multipliers, and enterprise discounts are not assumed unless you set them in the related savings tools.
Use this page for directional planning, then confirm with your provider invoice and a quality eval on your domain. For a full planning path, see the LLM cost planning guide. Last verified against catalog v2026-09-11 (page checked 2026-09-12).
Sources
Vendor list pages can change without notice. OmniKit’s calculator uses a dated catalog; these links are the public originals.
- OpenAI API pricing · checked 2026-09-11
- Anthropic Claude API pricing · checked 2026-09-11
- Google Gemini developer pricing · checked 2026-09-11
- DeepSeek API pricing · checked 2026-09-11
FAQ
How current are the prices?
Last verified 11 September 2026 in OmniKit catalog v2026-09-11 (GET /v1/pricing/models). Dual-rate models use the higher peak or uncached list unless a current intro window is noted. Re-run the calculator after vendor announcements.
Why four models?
Enough spread for a strategy review without overwhelming the chart: DeepSeek Flash, GPT-4o Mini, Claude Sonnet 5, Gemini 3.8 Flash. Add more ids in the full LLM cost calculator.
What about open source?
Add the vLLM A100 profile in the full LLM cost calculator for a self-host reference point. Self-host TCO is GPU hours, not only token list price.
Does OmniKit update prices automatically from vendors?
No live scrape. Prices follow a versioned JSON catalog. Operators should re-check official OpenAI, Anthropic, Google, and DeepSeek pricing pages after major launches.
How do I estimate ChatGPT API cost for 1 million tokens?
Multiply tokens by the matching input or output list rate. One million GPT-4o Mini input tokens at $0.15/1M is $0.15; one million output tokens at $0.60/1M is $0.60. Real traffic mixes both. Use the calculator for your prompts.
Related comparisons
Go deeper in the full calculator
Add more models, tune prompts, and export CSV for finance reviews.
Open LLM Cost Calculator