−90%
Typical saving
Prompt caching
A stable prefix reused often enough that cache reads beat the one-time write premium.
Change it in Prompts
Insights audits your workloads once a week and hands back the changes worth making - with the math, the sampled requests, and the screen to make them on.
Free by default · Read-only · No code changes to see it
Weekly run · Aug 3 · 02:00 UTC
read-onlyFound this run
71% of that workload's bill, across 4 scopes.
$10,400/mo
Four changes. $10,400 a month. Nobody rewrote a line of routing code.
production chat + batch workload
41K req/wk · 7 Key + Model scopes analyzed
Illustrative example. Savings depend on your traffic mix, models, and prompt structure.
Each is detected at the grain it lives on - one Key, one model - and priced against your own sampled tokens at live rates.
−90%
Typical saving
A stable prefix reused often enough that cache reads beat the one-time write premium.
Change it in Prompts
−50%
Typical saving
Latency-tolerant volume moved to the discounted tier, with standard-tier fallback.
Change it with :flex or an alias
−60–85%
Typical saving
A cheaper model that held your quality bar on your own requests, scored by an LLM judge.
Change it in Keys
−20–40%
Typical saving
Prompt tokens trimmed by a compression engine without moving the output.
Change it in Prompts
Percentages are typical savings on the scopes where a leak is detected, not a guarantee - every card carries the figure measured on your own traffic.
One job, once a week, per Key + Model scope. Nothing to install, nothing to configure.
Step 1
A scope qualifies once it passes 1,000 text-generation requests in 7 days.
read-only · requests never modified
Step 2
Spend, cache potential, tier eligibility and - on Advanced - eval-verified swap and compression candidates.
frozen at generation · never re-computed
Step 3
Cards sorted by projected impact, normalized per week, each naming where the change is made.
ranked by projected impact · advisory only
A number you can't interrogate is a number you won't act on. Each card opens into the work behind it - what was measured, what's proposed, and where the change is made.
Sampled requests priced at live rates - including your negotiated ones - with the formula and its inputs shown.
The exact surface and the suggested change, spelled out: Keys for model swaps, Prompts for caching and compression, the :flex slug or an alias for Flex tier.
Figures never drift. Next week's run supersedes this one instead of mixing with it.
prompt_caching
prod-chat · claude-sonnet-4.6
Suggested change: mark the stable prefix with cache_control
Advanced insights spend real tokens on your behalf, so they're off until you say otherwise.
$0/ always
qualifies at 1,000+ req / wk
per Key + Model scope
Token cost· itemized on your invoice
qualifies at 10,000+ req or $50+ spend / wk
per Key + Model scope
Both tiers refresh on the same weekly run. Thresholds are indicative and may be tuned server-side.
Insights is a standing weekly audit of your LLM bill - instead of a quarterly cost-review scramble, the biggest opportunities surface every run with the evidence already assembled.
Stable system prompts reused thousands of times a day are prime caching territory. Insights finds the prefixes, sizes the reuse, and shows the caching math before you touch a prompt.
Summarization, enrichment, and other latency-tolerant jobs often qualify for Flex tier at roughly half the cost. Insights sizes the eligible volume and shows the switch - slug or alias with standard-tier fallback.
Teams default to the biggest model and rarely revisit. With Advanced cost insights, swap candidates are sampled on your real requests and scored by an LLM judge - so a downgrade is only recommended when quality holds.
Retrieved context and tool schemas inflate every prompt. Compression candidates trim the tokens that don't move the output, verified before they're ever recommended.
You don't. Free insights run automatically on qualifying scopes. The only switch on the page is Advanced cost insights, which is off by default.
No. Analysis is read-only and Insights never edits a request, a prompt, or a route. Every change is one you make yourself, versioned and reviewable.
After the next weekly run, once a Key + Model scope has passed 1,000 text-generation requests in the previous 7 days. A run that finds nothing says so.
Each sampled request is priced against its real tokens at live rates - including per-customer negotiated pricing - then normalized to a per-week rate. The formula and inputs sit on every card.
Judge and sample tokens at standard rates, itemized on your invoice. You see the disclosure before the toggle takes effect.
Never. Savings, candidates, window and confidence are frozen when the run generates them. The next run supersedes the last rather than mixing figures.
Route a week of qualifying traffic and your first recommendations arrive with the next run. Free insights included · read-only · no card required.