AI Engineering · 6 MIN
How do you control LLM costs with prompt caching?
Prompt caching, batching, model routing, and output limits can cut LLM bills without hurting quality. Here is how to apply each and measure the savings.
You control LLM costs by reducing how many tokens you pay full price for. Prompt caching does this by reusing a stable prompt prefix across requests, and both Anthropic and OpenAI bill cached input at a fraction of the normal input rate. Batching, model routing, and output limits stack on top. The order matters: measure first, cache what repeats, route easy work to smaller models, and batch anything that is not urgent. This post covers each lever and how to confirm it worked.
- Cost is tokens times price, so every lever either cuts tokens or moves them to a cheaper tier.
- Prompt caching reuses an unchanged prompt prefix. Provider docs describe cached reads priced well below normal input, with a small premium to write the cache.
- Cache hits require an identical prefix, so put stable content first and variable content last.
- Batch APIs trade latency for a lower price and suit evals, classification, and backfills.
- Route simple tasks to smaller models only after an eval confirms quality holds.
- Nactore instruments cost per successful task on every feature it ships, so savings are measured and not assumed.
Where does LLM spend actually come from?
Most bills are dominated by input tokens that repeat. A support assistant sends the same system prompt, policy documents, tool definitions, and examples with every request. An agent resends its growing conversation on every step. You pay full price for that repeated text each time unless you cache it.
Before optimizing, log tokens in, tokens out, and cost per request, then compute cost per successful task. A cheap request that fails and retries three times is not cheap. See LLM observability: what to log.
How does prompt caching work?
When a request begins with a prefix the provider has seen recently, it can reuse the already processed prefix instead of recomputing it. You pay a lower rate for the cached portion and usually get lower latency.
The details differ by provider, so check current docs before you plan.
| Aspect | Anthropic | OpenAI |
|---|---|---|
| How you enable it | Add cache_control to the request or to specific content blocks | Enabled by default on supported models |
| What matches | Identical prefix up to a breakpoint, in order: tools, system, messages | Identical prefix |
| Minimum size | Varies by model, stated in the docs | Stated in the docs, 1,024 tokens on recent models |
| Pricing shape | Cache writes cost a premium over base input, cache reads cost a fraction | Cached input discounted, write pricing varies by model |
| How you verify | cache_read_input_tokens and cache_creation_input_tokens in usage | cached_tokens in usage details |
| Lifetime | Default 5 minutes, optional 1 hour, refreshed on reuse | Minutes to hours depending on model and settings |
Sources: Anthropic's prompt caching documentation and OpenAI's prompt caching guide. Exact multipliers and thresholds change by model, so read the current page before you model savings.
How do you structure prompts so the cache hits?
Caching only works on an exact prefix match, so prompt layout is the whole game.
- Put stable content first. Tool definitions, system instructions, policy text, and few-shot examples go at the top.
- Put variable content last. The user's message, retrieved passages, and timestamps go after the stable block.
- Never edit the prefix casually. Anthropic's docs note that changing tool definitions invalidates every cache level below them. A single reordered tool breaks it.
- Keep dynamic values out of the prefix. A current timestamp or request ID in the system prompt guarantees a miss every time.
- Mind the lifetime. If traffic is bursty, a short cache window may expire between requests, and a longer window or a warm-up request may pay off.
After you ship caching, graph the cache hit rate from the usage fields. A feature that "has caching on" but sits at a near-zero hit rate usually has a timestamp or a reordered tool list in its prefix.
What other levers reduce cost?
- Batch processing. OpenAI's Batch API is documented at half the synchronous price with a 24-hour completion window, and suits evals, classification, and embedding large repositories. Anything that does not need an instant answer is a candidate.
- Model routing. Send easy, high-volume tasks to a smaller model and escalate hard cases. Anthropic's model selection guide describes starting efficiency-first and upgrading only where evals show a gap, and also describes multi-model patterns where a lower-cost model does bulk work. For the selection method, see choosing an LLM model for production.
- Effort and reasoning settings. Some models expose a parameter that trades intelligence for latency and cost. Anthropic's guide notes that tuning it is often a better lever than switching models.
- Output limits. Cap
max_tokens, ask for terse structured output, and avoid asking the model to restate the input. - Retrieval discipline. Send the three best passages, not the twenty nearest. More context costs more and can lower accuracy.
- Response caching. For truly identical requests, cache the final answer in your own store and skip the model call.
How do you prove the savings are real?
Run the change through your eval set first so you know quality held, then compare cost per successful task before and after, not only price per token. Check three numbers in staging: cache hit rate, average input cost per request, and eval pass rate. If the pass rate drops, the saving is not a saving. The method for this is in why AI features need evals before production.
Frequently asked questions
Does prompt caching change the answers?
It does not change the model's processing of your prompt. It reuses computation for an identical prefix. Always confirm with your own eval run after any prompt restructuring, because moving content around can change behavior.
What is the first thing to try?
Restructure prompts so stable content comes first, enable caching, and watch the cache-read fields in the usage data. It is usually the lowest-effort change.
Is a cheaper model always the answer?
No. A smaller model that needs retries or produces more escalations can cost more per successful task. Compare on your eval set.
Should we cache across users?
Shared prefixes such as system prompts and tool definitions are safe and valuable to cache. Never put one user's private data in a shared prefix. Check your provider's isolation rules.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.