Prompt Cache Savings
Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models.
- 1Fill inputs
- 2Run
- 3Copy / export
Inputs
Required fields on the left
0.5 ≈ half-price reads · 0.1 ≈ Anthropic-style
Free local math from list prices. Pair with LLM Cost Calculator or Prompt Token Diff.
Results
Appears after you run
Price the cache hit
Enter a reusable prefix size, hit rate, and traffic to see monthly $ saved vs full-price input.
From inputs to a decision
Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models.
- 01
Enter cached prefix tokens and per-request dynamic tokens.
- 02
Set hit rate and cache-read multiplier (fraction of list input price).
- 03
Add traffic (RPS) and models to compare.
- 04
Export the monthly savings table.
Tips for better results
- Anthropic-style cache reads are often ~10% of input; OpenAI-style discounts are commonly ~50% — set the multiplier to match your provider.
- Savings scale with prefix size and hit rate more than with output tokens.
- Use the same prompt pack and RPS when comparing models.
Frequently asked questions
- Does this include cache write fees?
- No. V1 models steady-state reads only. If your provider charges extra for the first write, treat that as a one-time setup cost outside this table.
- What hit rate should I assume?
- Start at 0.7–0.9 for sticky system prompts. Lower it if prefixes rotate often or traffic is cold.
Related tools
Keep measuring in the same cluster — or jump to the next decision.
- LLM Cost CalculatorSide-by-side monthly token cost across DeepSeek, OpenAI, Anthropic, Gemini, and vLLM.Open →
- Prompt Token DiffCompare two prompts on tokens and monthly $ across models.Open →
- Batch vs Realtime CostCompare monthly spend on realtime list prices vs discounted batch APIs.Open →
- RAG Cost EstimatorProject monthly embedding + retrieval + generation spend for a RAG stack.Open →