← Tools

B2B / Engineering

Free tool

Prompt Cache Savings

Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models.

  1. 1Fill inputs
  2. 2Run
  3. 3Copy / export

Inputs

0.5 ≈ half-price reads · 0.1 ≈ Anthropic-style

Free local math from list prices. Pair with LLM Cost Calculator or Prompt Token Diff.

Results

Price the cache hit

Enter a reusable prefix size, hit rate, and traffic to see monthly $ saved vs full-price input.

How it works

From inputs to a decision

Estimate how much you save when a long system prompt or tool schema is cached at a discounted input rate. Compare monthly spend with vs without cache hits across catalog models.

  1. 01

    Enter cached prefix tokens and per-request dynamic tokens.

  2. 02

    Set hit rate and cache-read multiplier (fraction of list input price).

  3. 03

    Add traffic (RPS) and models to compare.

  4. 04

    Export the monthly savings table.

Tips for better results

  • Anthropic-style cache reads are often ~10% of input; OpenAI-style discounts are commonly ~50% — set the multiplier to match your provider.
  • Savings scale with prefix size and hit rate more than with output tokens.
  • Use the same prompt pack and RPS when comparing models.

Frequently asked questions

Does this include cache write fees?
No. V1 models steady-state reads only. If your provider charges extra for the first write, treat that as a one-time setup cost outside this table.
What hit rate should I assume?
Start at 0.7–0.9 for sticky system prompts. Lower it if prefixes rotate often or traffic is cold.

Keep measuring in the same cluster — or jump to the next decision.