Insights

Cut 50–80% from the traffic you already route.

Insights audits your workloads once a week and hands back the changes worth making - with the math, the sampled requests, and the screen to make them on.

Free by default · Read-only · No code changes to see it

Four changes. $10,400 a month. Nobody rewrote a line of routing code.

production chat + batch workload
41K req/wk · 7 Key + Model scopes analyzed

$10,400
Saved per month
71%
Of that workload's bill
4
Changes to make
0
Requests modified

Illustrative example. Savings depend on your traffic mix, models, and prompt structure.

Recommendation types

Four leaks, priced weekly

Each is detected at the grain it lives on - one Key, one model - and priced against your own sampled tokens at live rates.

Free

−90%

Typical saving

Prompt caching

A stable prefix reused often enough that cache reads beat the one-time write premium.

Change it in Prompts

Free

−50%

Typical saving

Flex tier

Latency-tolerant volume moved to the discounted tier, with standard-tier fallback.

Change it with :flex or an alias

Advanced

−60–85%

Typical saving

Model swap

A cheaper model that held your quality bar on your own requests, scored by an LLM judge.

Change it in Keys

Advanced

−20–40%

Typical saving

Prompt compression

Prompt tokens trimmed by a compression engine without moving the output.

Change it in Prompts

Percentages are typical savings on the scopes where a leak is detected, not a guarantee - every card carries the figure measured on your own traffic.

How it works

1,000 requests in. A ranked feed out.

One job, once a week, per Key + Model scope. Nothing to install, nothing to configure.

Step 1

You route as usual

A scope qualifies once it passes 1,000 text-generation requests in 7 days.

Key + Model1,000+ req / 7 days

read-only · requests never modified

Step 2

Detectors scan the week

Spend, cache potential, tier eligibility and - on Advanced - eval-verified swap and compression candidates.

cachingflex tierswapcompression

frozen at generation · never re-computed

Step 3

You get a ranked feed

Cards sorted by projected impact, normalized per week, each naming where the change is made.

KeysPromptsAliases

ranked by projected impact · advisory only

Evidence

Every figure opens

A number you can't interrogate is a number you won't act on. Each card opens into the work behind it - what was measured, what's proposed, and where the change is made.

The measurement, not the conclusion

Sampled requests priced at live rates - including your negotiated ones - with the formula and its inputs shown.

A named next step

The exact surface and the suggested change, spelled out: Keys for model swaps, Prompts for caching and compression, the :flex slug or an alias for Flex tier.

Frozen at generation

Figures never drift. Next week's run supersedes this one instead of mixing with it.

prompt_caching

prod-chat · claude-sonnet-4.6

High confidence
Projected savings
$610 / wk
Stable prefix
6,214 tokens
Reuse
14× per session
Volume
41K req / wk
Formula
reused_tok × rate × 0.90
Measurement window
Jul 27 – Aug 2
Generated
Aug 3 · frozen

Suggested change: mark the stable prefix with cache_control

Tiers

Free by default. One toggle for the rest.

Advanced insights spend real tokens on your behalf, so they're off until you say otherwise.

Free insights

On by default

$0/ always

  • Prompt caching and Flex tier recommendations
  • Metadata only - volumes, tokens, spend, cache patterns
  • Weekly run with full evidence and savings math

qualifies at 1,000+ req / wk
per Key + Model scope

Get started for free

Advanced cost insights

Single toggle

Token cost· itemized on your invoice

  • Everything in Free, plus model swap and compression
  • An LLM judge scores candidates against your current model
  • Top 3 system prompts per scope sampled to comparison models

qualifies at 10,000+ req or $50+ spend / wk
per Key + Model scope

See what it costs

Both tiers refresh on the same weekly run. Thresholds are indicative and may be tuned server-side.

Use cases

Built for workloads with money on the table

Insights is a standing weekly audit of your LLM bill - instead of a quarterly cost-review scramble, the biggest opportunities surface every run with the evidence already assembled.

High-volume chat products

Stable system prompts reused thousands of times a day are prime caching territory. Insights finds the prefixes, sizes the reuse, and shows the caching math before you touch a prompt.

Batch and background pipelines

Summarization, enrichment, and other latency-tolerant jobs often qualify for Flex tier at roughly half the cost. Insights sizes the eligible volume and shows the switch - slug or alias with standard-tier fallback.

Over-provisioned model choices

Teams default to the biggest model and rarely revisit. With Advanced cost insights, swap candidates are sampled on your real requests and scored by an LLM judge - so a downgrade is only recommended when quality holds.

Long-context RAG and agents

Retrieved context and tool schemas inflate every prompt. Compression candidates trim the tokens that don't move the output, verified before they're ever recommended.

FAQ

Six questions, short answers

You don't. Free insights run automatically on qualifying scopes. The only switch on the page is Advanced cost insights, which is off by default.

No. Analysis is read-only and Insights never edits a request, a prompt, or a route. Every change is one you make yourself, versioned and reviewable.

After the next weekly run, once a Key + Model scope has passed 1,000 text-generation requests in the previous 7 days. A run that finds nothing says so.

Each sampled request is priced against its real tokens at live rates - including per-customer negotiated pricing - then normalized to a per-week rate. The formula and inputs sit on every card.

Judge and sample tokens at standard rates, itemized on your invoice. You see the disclosure before the toggle takes effect.

Never. Savings, candidates, window and confidence are frozen when the run generates them. The next run supersedes the last rather than mixing figures.

Find out what your traffic is trying to tell you

Route a week of qualifying traffic and your first recommendations arrive with the next run. Free insights included · read-only · no card required.