Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models.
Required fields on the left
0.5 ≈ half-price reads · 0.1 ≈ Anthropic-style
Free local math from list prices. Pair with LLM Cost Calculator or Prompt Token Diff.
Appears after you run
Price the cache hit
Enter a reusable prefix size, hit rate, and traffic to see monthly $ saved vs full-price input.
Prompt cache savings estimates how much you save when a long system prompt or tool schema is cached at a discounted input rate. It compares monthly spend with versus without cache hits. It models steady-state reads, not the first cache-write fee, and the discount multiplier is yours to set per provider.
Anthropic-style cache reads are often about 10% of input; OpenAI-style discounts are commonly about 50%. Set the multiplier to match the contract you actually have.
Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models.
Enter cached prefix tokens, per-request dynamic tokens, hit rate, cache-read multiplier, RPS, and models. Export the monthly savings table. Writes are out of scope for v1 — treat the first write as a one-time cost outside this page.
Enter cached prefix tokens and per-request dynamic tokens.
Set hit rate and cache-read multiplier (fraction of list input price).
Add traffic (RPS) and models to compare.
Export the monthly savings table.
No. It models steady-state cache reads. If your provider charges extra for the first write, put that in a separate setup line. This table is for the repeating discount after the prefix is warm.
Usually not on the same request. Use this page for realtime sticky prefixes. Use Batch vs Realtime for deferred volume at a batch multiplier. Interactive chat stays on cache; overnight jobs stay on batch.
Prompt cache savings estimates how much you save when a long system prompt or tool schema is cached at a discounted input rate. It compares monthly spend with versus without cache hits. It models steady-state reads, not the first cache-write fee, and the discount multiplier is yours to set per provider.
No. V1 models steady-state reads only. If your provider charges extra for the first write, treat that as a one-time setup cost outside this table.
Start at 0.7–0.9 for sticky system prompts. Lower it if prefixes rotate often or traffic is cold.
Project monthly $ saved when a reusable prompt prefix hits the provider cache.
Keep measuring in the same cluster — or jump to the next decision.
Prompt Cache Savings lives at omnikitapp.net/tools/prompt-cache-savings. Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models. Use it as an operator page, then confirm invoices, policies, or published facts on the source. OmniKit does not sell detector evasion or ranking guarantees.