Insights

Your traffic already knows where the savings are

Insights analyzes your FastRouter workloads once a week and turns them into evidence-backed cost recommendations - each with the projected savings, the evidence behind it, and exactly where to make the change.

Free insights on by default · Read-only - Insights never modifies your traffic

Insights

Recommendations

Weekly run · Aug 3

Savings found

$2,140/wk

Recommendations

3 open

Scopes analyzed

7

Cost

Enable prompt caching on prod-chat · claude-sonnet-4.6

Stable 6.2K-token prefix reused 14× per session

Cost

Move batch summarization to Flex tier

78% of requests tolerate queueing · est. $410/wk

Cost

Swap classification traffic to a lighter model

Eval-verified: quality parity on 200 sampled requests

Measurement window: Jul 27 – Aug 2FreeAdvanced
Why Insights

Recommendations, not dashboards

You already have the metrics. Insights reads them for you - and hands back specific actions with a dollar figure attached, plus the evidence behind each one.

From your own traffic

Every recommendation is generated from a weekly analysis of your actual requests - real volumes, real prompts, real provider behavior. No generic best-practice checklists.

Advisory, with the evidence attached

Insights recommends; you decide. Every card opens into the sampled requests, savings math, and confidence rationale - then points to the exact surface where the change is made.

Free by default, advanced on toggle

Prompt caching and Flex tier insights are included and enabled automatically. Advanced cost insights - model-swap and compression candidates verified by automatic evaluations - are a separate opt-in with clear billing disclosure.

How it works

From your traffic to a ranked feed

A weekly job analyzes every qualifying Key + Model scope, freezes what it finds as a point-in-time snapshot, and publishes it to your Insights feed with the run metadata on top.

Step 1 · Your traffic

Qualifying scopes

Key + Model1,000+ req / 7 days
  • Analysis runs per Key + Model scope, qualifying at 1,000+ text-generation requests in the past 7 days.
  • Read-only - your requests are never modified by analysis.

Step 2 · FastRouter

Weekly analysis run

prompt cachingflex tiermodel swap*
  • Detectors scan spend, cache potential, and service-tier eligibility across your traffic.
  • Each finding is frozen at generation: savings, evidence, and confidence fixed at that moment.
  • *Model-swap candidates are verified by auto-evals under Advanced cost insights.

Step 3 · Back to you

Ranked recommendations

KeysPromptsAliases
  • Cards ranked by projected impact, normalized to a per-week rate.
  • Evidence in every flyout: sampled requests, before/after figures, confidence rationale.
  • Each card names the surface where you make the change - Keys, Prompts, or Virtual Model Aliases.

Frozen snapshots, refreshed weekly

Savings, candidate models, measurement period, and confidence are fixed when the run generates them - figures never silently drift under you. The next weekly run supersedes the previous one, so you always see the latest run only, with its metadata above every card.

  • Point-in-time by design. Every figure is frozen at generation and never re-computed under you.
  • Superseded, not contradicted. Each run replaces the previous one instead of mixing old and new figures.
  • Honest empty states. A run that found nothing says so, and a brand-new account is told its first run is scheduled.

Run metadata

Snapshot
Last generated
Aug 3, 2026 · 02:00 UTC
Measurement window
Jul 27 – Aug 2
Scopes analyzed
7 Key + Model scopes
Recommendations
3 open
Refresh cadence
Weekly
Recommendation types

Three ways your spend can shrink

Each type is scoped to the grain where it was detected - a specific Key + Model. All recommendations are advisory: Insights shows you the evidence and the projected savings, and you make the change where it belongs.

Prompt caching

A stable prompt prefix is reused enough that cache reads (up to 90% off) beat the one-time write premium. The card points you to the exact prompt and prefix, with the reuse rate measured from your traffic.

Flex tier

Latency-tolerant traffic qualifies for a discounted service tier. The card shows the eligible volume and savings math, and how to switch - the :flex slug or a Virtual Model Alias with standard-tier fallback.

Model swap

A cheaper model matches your quality bar on your own traffic - priced against each sampled request's real tokens, and verified by automatic evaluations under Advanced cost insights before it's ever recommended.

Recommendation types

3 types
  • prompt_cachingReused prefixes → cache readsCost
  • flex_tierDiscounted tier for tolerant trafficCost
  • model_swapCheaper model, same quality barCost
With Advanced cost insights:auto-evalsprompt compression
On the roadmap:latencyerror reductioncache reordering
Evidence & next step

Every recommendation shows its work

A card is only as good as the evidence behind it. Each flyout lays out what was measured, what's proposed, and precisely where in FastRouter you make the change.

The measurement, not just the conclusion

Sampled requests priced at live rates, the prefix or volume analysis behind the finding, and the measurement window it was drawn from.

A named next step

Model swaps point to Keys, prompt caching points to Prompts, and Flex tier points to the :flex slug or Virtual Model Aliases - with the suggested change spelled out.

Deliberate by design

Changes to model choice, prompt structure, and routing are versioned, reviewable edits you make in the product - Insights never modifies anything itself.

Recommendation detail · prompt_caching

Cost
Scopeprod-chat · claude-sonnet-4.6
Projected savings$610 / wk
ConfidenceHigh
PStable prefix: 6,214 tokens system + tool schemas
RReuse: 14× per session across 41K req/wk
$Cache reads −90% vs one-time write premium
Open in Prompts

Suggested change: mark the stable prefix with cache_control

Measurement

Projected figures you can interrogate

Every card carries a projected figure frozen at generation and normalized to a per-week rate - with the formula, inputs, and measurement window shown, so the number can be checked, not just believed.

Real formulas, live pricing

Cost projections price each sampled request's actual tokens against the live pricing table - including per-customer negotiated rates - at generation time.

Confidence with a rationale

High, Medium, or Low, fixed per run, with a one-paragraph explanation of why - so you can weigh a strong signal differently from a thin one.

Comparable across time

Per-week normalization means a figure generated this run reads the same way as one from a month ago - no re-basing needed to compare opportunities.

flex_tier · prod-batch

High confidence
Projected savings
$410 / wk
Eligible volume
78% of 41K req/wk
Formula
eligible_vol × cost/req × flex_discount
Measurement window
Jul 27 – Aug 2
Generated
Aug 3, 2026 · frozen at generation
Next step
:flex slug or Virtual Model Alias
Tiers

Free insights, one advanced toggle

Free insights run automatically on qualifying traffic. Advanced cost insights are off by default and enabled with a single toggle - because they spend real tokens on your behalf, and you should choose that.

Free insights compared with Advanced cost insights
How the tiers compareFree insightsIncludedPrompt caching · Flex tierAdvanced cost insightsOpt-inAuto-evals · Compression
Coverage
Recommendation typesPrompt caching and Flex tierEverything in Free, plus eval-verified model swaps and prompt compression candidates
Quality verificationNot includedLLM judge scores candidate outputs against your current model
Enablement & qualification
EnablementOn by defaultSingle toggle, off by default, with billing disclosure
Scope qualifies at1,000+ text-generation requests / week per Key + Model10,000+ requests or $50+ spend / week per Key + Model
What's analyzedMetadata: volumes, tokens, spend, and cache patternsAdditionally: top 3 system prompts per qualifying scope, sampled to comparison models
Cost
What you payFreeJudge and sample tokens, billed at standard rates and itemized on your invoice

Both tiers refresh on the same weekly run. Qualification thresholds are indicative and may be tuned server-side.

Use cases

Built for workloads with money on the table

From chat products to background pipelines, Insights turns the traffic you already route into concrete, evidence-backed savings.

High-volume chat products

Stable system prompts reused thousands of times per day are prime caching territory. Insights finds the prefixes, sizes the reuse, and shows the caching math before you touch a prompt.

Batch and background pipelines

Summarization, enrichment, and other latency-tolerant jobs often qualify for Flex tier at roughly half the cost. Insights sizes the eligible volume and shows the switch - slug or alias with standard-tier fallback.

Over-provisioned model choices

Teams default to the biggest model and rarely revisit. With Advanced cost insights, model-swap candidates are sampled on your real requests and scored by an LLM judge - so a downgrade is only recommended when quality holds.

Growing spend, no time to audit

Insights is a standing weekly audit of your LLM bill. Instead of a quarterly cost-review scramble, the biggest opportunities surface every run, ranked by projected savings, with the evidence already assembled.

FAQ

Insights questions, answered

Free insights - prompt caching and Flex tier recommendations - are enabled automatically for your org. Analysis runs once a week: your first set of recommendations appears after the next weekly run, and a Key + Model scope needs at least 1,000 text-generation requests in the past 7 days to qualify for analysis. The only setting is the Advanced cost insights toggle, which adds eval-verified model-swap and compression recommendations and is off by default.

No. Analysis is read-only against your request metadata, and all recommendations are advisory - Insights never modifies your routing, prompts, or requests. Any change is one you make yourself in the product, as a normal versioned, reviewable edit.

Advanced cost insights verify model-swap and prompt compression candidates by making sample requests to comparison models and scoring the outputs with an LLM judge. Those judge and sample tokens are billed at standard rates and itemized on your invoice. The toggle is off by default, requires explicit enablement with the billing disclosure shown, and only runs on scopes exceeding 10,000 requests or $50 spend per week.

Each recommendation type has an explicit formula computed at generation time against live pricing - for example, prompt caching weighs prefix tokens, reuse rate, and the cache-read discount against the one-time write premium, and Flex tier multiplies eligible volume by cost per request and the tier discount. Figures are normalized to a per-week rate and frozen in the snapshot, with the formula, inputs, and measurement window shown in the flyout so the projection can be verified against your own data.

Each recommendation names its surface. Model swaps are made in Keys; prompt caching is enabled in Prompts by marking the stable prefix with cache_control, with prompt versions retained so the change can be reverted; and Flex tier is switched by appending the :flex slug or configuring a Virtual Model Alias with standard-tier fallback. Review the evidence, make the change, and the next weekly run measures the new baseline.

Weekly. Each run produces a fresh set of point-in-time snapshots that supersede the previous run, and the run metadata - last generated, measurement window, refresh cadence - is shown above every card. If a run finds nothing new, the feed says exactly that.

Find out what your traffic is trying to tell you

Free insights are already included with FastRouter. Route a week of qualifying traffic and your first recommendations arrive with the next run.