Prompt caching savings

Cached Input Pricing Calculator

Use this cached input pricing calculator to estimate savings when part of your prompt is reused across many LLM requests.

Calculate Cached Input Pricing Cost

Estimate how cached input pricing can reduce LLM API costs for repeated system prompts, instructions and shared context.

No API key required. Your transcript stays in your browser.

Characters: 0Words: 0Estimated input tokens: 0Estimated token count

Output tokens are estimated based on the selected summary type and input length.

Number of AI summarization requests expected each month.

Advanced Settings
i

Tokens used by the system prompt or recurring instructions in every interaction.

i

Percentage of input tokens expected to use provider prompt caching.

Compare the Same Workload Across Models

Compare model pricing, per-interaction cost, and monthly difference using the same current calculator values.

Paste sample content above to compare model costs.

Cached input helps repeated prompts

Some models offer lower pricing for cached or reused input tokens. This matters for apps that send the same system instructions, policy text or tool schema on many requests.

Enter stable prompt content as system instruction tokens and adjust cached input percentage in Advanced Settings to estimate the impact.

Reusable inputGood caching candidate?Reason
System promptYesOften repeated on every request
Tool schemaYesUsually stable across calls
User messageNoUsually different each request

Cached Input Pricing Calculator FAQ

What is cached input pricing?

It is discounted pricing for prompt tokens that a provider can reuse from previous requests.

Which tokens should I mark as cached?

Only stable repeated tokens, such as recurring instructions or shared context. Do not mark unique user content as cached.

Do all models support cached input pricing?

No. The calculator only applies cached-input cost where the selected model has cached pricing configured.