From your own traffic
Every recommendation is generated from a weekly analysis of your actual requests - real volumes, real prompts, real provider behavior. No generic best-practice checklists.
Insights analyzes your FastRouter workloads once a week and turns them into evidence-backed cost recommendations - each with the projected savings, the evidence behind it, and exactly where to make the change.
Free insights on by default · Read-only - Insights never modifies your traffic
Recommendations
Weekly run · Aug 3Savings found
$2,140/wk
Recommendations
3 open
Scopes analyzed
7
Enable prompt caching on prod-chat · claude-sonnet-4.6
Stable 6.2K-token prefix reused 14× per session
Move batch summarization to Flex tier
78% of requests tolerate queueing · est. $410/wk
Swap classification traffic to a lighter model
Eval-verified: quality parity on 200 sampled requests
You already have the metrics. Insights reads them for you - and hands back specific actions with a dollar figure attached, plus the evidence behind each one.
Every recommendation is generated from a weekly analysis of your actual requests - real volumes, real prompts, real provider behavior. No generic best-practice checklists.
Insights recommends; you decide. Every card opens into the sampled requests, savings math, and confidence rationale - then points to the exact surface where the change is made.
Prompt caching and Flex tier insights are included and enabled automatically. Advanced cost insights - model-swap and compression candidates verified by automatic evaluations - are a separate opt-in with clear billing disclosure.
A weekly job analyzes every qualifying Key + Model scope, freezes what it finds as a point-in-time snapshot, and publishes it to your Insights feed with the run metadata on top.
Step 1 · Your traffic
Step 2 · FastRouter
Step 3 · Back to you
Savings, candidate models, measurement period, and confidence are fixed when the run generates them - figures never silently drift under you. The next weekly run supersedes the previous one, so you always see the latest run only, with its metadata above every card.
Run metadata
SnapshotEach type is scoped to the grain where it was detected - a specific Key + Model. All recommendations are advisory: Insights shows you the evidence and the projected savings, and you make the change where it belongs.
A stable prompt prefix is reused enough that cache reads (up to 90% off) beat the one-time write premium. The card points you to the exact prompt and prefix, with the reuse rate measured from your traffic.
Latency-tolerant traffic qualifies for a discounted service tier. The card shows the eligible volume and savings math, and how to switch - the :flex slug or a Virtual Model Alias with standard-tier fallback.
A cheaper model matches your quality bar on your own traffic - priced against each sampled request's real tokens, and verified by automatic evaluations under Advanced cost insights before it's ever recommended.
Recommendation types
3 typesA card is only as good as the evidence behind it. Each flyout lays out what was measured, what's proposed, and precisely where in FastRouter you make the change.
Sampled requests priced at live rates, the prefix or volume analysis behind the finding, and the measurement window it was drawn from.
Model swaps point to Keys, prompt caching points to Prompts, and Flex tier points to the :flex slug or Virtual Model Aliases - with the suggested change spelled out.
Changes to model choice, prompt structure, and routing are versioned, reviewable edits you make in the product - Insights never modifies anything itself.
Recommendation detail · prompt_caching
CostSuggested change: mark the stable prefix with cache_control
Every card carries a projected figure frozen at generation and normalized to a per-week rate - with the formula, inputs, and measurement window shown, so the number can be checked, not just believed.
Cost projections price each sampled request's actual tokens against the live pricing table - including per-customer negotiated rates - at generation time.
High, Medium, or Low, fixed per run, with a one-paragraph explanation of why - so you can weigh a strong signal differently from a thin one.
Per-week normalization means a figure generated this run reads the same way as one from a month ago - no re-basing needed to compare opportunities.
flex_tier · prod-batch
High confidenceFree insights run automatically on qualifying traffic. Advanced cost insights are off by default and enabled with a single toggle - because they spend real tokens on your behalf, and you should choose that.
| How the tiers compare | Free insightsIncludedPrompt caching · Flex tier | Advanced cost insightsOpt-inAuto-evals · Compression |
|---|---|---|
| Coverage | ||
| Recommendation types | Prompt caching and Flex tier | Everything in Free, plus eval-verified model swaps and prompt compression candidates |
| Quality verification | Not included | LLM judge scores candidate outputs against your current model |
| Enablement & qualification | ||
| Enablement | On by default | Single toggle, off by default, with billing disclosure |
| Scope qualifies at | 1,000+ text-generation requests / week per Key + Model | 10,000+ requests or $50+ spend / week per Key + Model |
| What's analyzed | Metadata: volumes, tokens, spend, and cache patterns | Additionally: top 3 system prompts per qualifying scope, sampled to comparison models |
| Cost | ||
| What you pay | Free | Judge and sample tokens, billed at standard rates and itemized on your invoice |
Both tiers refresh on the same weekly run. Qualification thresholds are indicative and may be tuned server-side.
From chat products to background pipelines, Insights turns the traffic you already route into concrete, evidence-backed savings.
Stable system prompts reused thousands of times per day are prime caching territory. Insights finds the prefixes, sizes the reuse, and shows the caching math before you touch a prompt.
Summarization, enrichment, and other latency-tolerant jobs often qualify for Flex tier at roughly half the cost. Insights sizes the eligible volume and shows the switch - slug or alias with standard-tier fallback.
Teams default to the biggest model and rarely revisit. With Advanced cost insights, model-swap candidates are sampled on your real requests and scored by an LLM judge - so a downgrade is only recommended when quality holds.
Insights is a standing weekly audit of your LLM bill. Instead of a quarterly cost-review scramble, the biggest opportunities surface every run, ranked by projected savings, with the evidence already assembled.
Free insights - prompt caching and Flex tier recommendations - are enabled automatically for your org. Analysis runs once a week: your first set of recommendations appears after the next weekly run, and a Key + Model scope needs at least 1,000 text-generation requests in the past 7 days to qualify for analysis. The only setting is the Advanced cost insights toggle, which adds eval-verified model-swap and compression recommendations and is off by default.
No. Analysis is read-only against your request metadata, and all recommendations are advisory - Insights never modifies your routing, prompts, or requests. Any change is one you make yourself in the product, as a normal versioned, reviewable edit.
Advanced cost insights verify model-swap and prompt compression candidates by making sample requests to comparison models and scoring the outputs with an LLM judge. Those judge and sample tokens are billed at standard rates and itemized on your invoice. The toggle is off by default, requires explicit enablement with the billing disclosure shown, and only runs on scopes exceeding 10,000 requests or $50 spend per week.
Each recommendation type has an explicit formula computed at generation time against live pricing - for example, prompt caching weighs prefix tokens, reuse rate, and the cache-read discount against the one-time write premium, and Flex tier multiplies eligible volume by cost per request and the tier discount. Figures are normalized to a per-week rate and frozen in the snapshot, with the formula, inputs, and measurement window shown in the flyout so the projection can be verified against your own data.
Each recommendation names its surface. Model swaps are made in Keys; prompt caching is enabled in Prompts by marking the stable prefix with cache_control, with prompt versions retained so the change can be reverted; and Flex tier is switched by appending the :flex slug or configuring a Virtual Model Alias with standard-tier fallback. Review the evidence, make the change, and the next weekly run measures the new baseline.
Weekly. Each run produces a fresh set of point-in-time snapshots that supersede the previous run, and the run metadata - last generated, measurement window, refresh cadence - is shown above every card. If a run finds nothing new, the feed says exactly that.
Free insights are already included with FastRouter. Route a week of qualifying traffic and your first recommendations arrive with the next run.