Insights

Your traffic already knows where the savings are

Insights analyzes your FastRouter workloads once a week and turns them into evidence-backed recommendations on cost, latency, and errors - each with the projected impact, the evidence behind it, and exactly where to make the change.

Free insights on by default · Read-only - Insights never modifies your traffic

Insights

Recommendations

Weekly run · Aug 3

Savings found

$2,140/wk

P95 latency

1.8s

Error rate

0.9%

Cost

Enable prompt caching on prod-chat · claude-sonnet-4.6

Stable 6.2K-token prefix reused 14× per session

Cost

Move batch summarization to Flex tier

78% of requests tolerate queueing · est. $410/wk

Latency

Route gpt-5.5 through fastest provider

p95 6.0s → 3.5s on alternate provider

Errors

Add failover for gemini-3.1-pro workloads

3 timeout spikes in measurement window

Measurement window: Jul 27 – Aug 2FreeAdvanced
Why Insights

Recommendations, not dashboards

You already have the metrics. Insights reads them for you - and hands back specific actions with a dollar or performance figure attached, plus the evidence behind each one.

From your own traffic

Every recommendation is generated from a weekly analysis of your actual requests - real volumes, real prompts, real provider behavior. No generic best-practice checklists.

Advisory, with the evidence attached

Insights recommends; you decide. Every card opens into the sampled requests, before/after figures, and confidence rationale - then points to the exact surface where the change is made.

Free by default, advanced on toggle

Cost, latency, and error insights are included and enabled automatically. Advanced cost insights - model-swap and compression candidates verified by automatic evaluations - are a separate opt-in with clear billing disclosure.

How it works

From your traffic to a ranked feed

A weekly job analyzes every qualifying Key + Model scope, freezes what it finds as a point-in-time snapshot, and publishes it to your Insights feed with the run metadata on top.

Step 1 · Your traffic

Qualifying scopes

Key + Model1,000+ req / 7 days
  • Analysis runs per Key + Model scope, qualifying at 1,000+ text-generation requests in the past 7 days.
  • Read-only - your requests are never modified by analysis.

Step 2 · FastRouter

Weekly analysis run

costlatencyerrorsauto-evals*
  • Six detectors scan spend, cache potential, tier eligibility, provider latency, and error patterns.
  • Each finding is frozen at generation: savings, evidence, and confidence fixed at that moment.
  • *Auto-evals run only with Advanced cost insights enabled.

Step 3 · Back to you

Ranked recommendations

KeysPromptsAliases
  • Cards ranked by projected impact, normalized to a per-week rate.
  • Evidence in every flyout: sampled requests, before/after figures, confidence rationale.
  • Each card names the surface where you make the change - Keys, Prompts, or Virtual Model Aliases.

Frozen snapshots, refreshed weekly

Savings, candidate models, measurement period, and confidence are fixed when the run generates them - figures never silently drift under you. The next weekly run supersedes the previous one, so you always see the latest run only, with its metadata above every card.

  • Point-in-time by design. Every figure is frozen at generation and never re-computed under you.
  • Superseded, not contradicted. Each run replaces the previous one instead of mixing old and new figures.
  • Honest empty states. A run that found nothing says so, and a brand-new account is told its first run is scheduled.

Run metadata

Snapshot
Last generated
Aug 3, 2026 · 02:00 UTC
Measurement window
Jul 27 – Aug 2
Scopes analyzed
7 Key + Model scopes
Recommendations
4 new · 2 open
Refresh cadence
Weekly
Recommendation types

Six detectors across cost, latency, and errors

Each type is scoped to the grain where it was detected. All recommendations are advisory - Insights shows you the evidence and the projected impact, and you make the change where it belongs.

Four cost detectors

model_swap finds cheaper models that match your quality bar, prompt_caching and cache_reordering turn reused prefixes into cache reads, and flex_tier moves latency-tolerant traffic to a discounted tier.

One latency detector

low_latency spots another provider serving the same model measurably faster, comparing current vs proposed at p95/p90/p75 with a suggested Lowest Latency alias configuration.

One error detector

error_reduction flags recurring timeouts and errors with the reference logs attached, and suggests a Priority Routing setup: healthy provider first, fallback on failure.

Recommendation types

6 detectors
  • model_swapCheaper model, same quality barCost
  • prompt_cachingReused prefixes → cache readsCost
  • cache_reorderingStatic blocks before dynamicCost
  • flex_tierDiscounted tier for tolerant trafficCost
  • low_latencyFastest provider for the modelLatency
  • error_reductionHealthy-first priority routingErrors
With Advanced cost insights:auto-evalsprompt compression
Evidence & next step

Every recommendation shows its work

A card is only as good as the evidence behind it. Each flyout lays out what was measured, what's proposed, and precisely where in FastRouter you make the change.

The measurement, not just the conclusion

Sampled requests, current vs proposed figures - p95/p90/p75 for latency, reference error logs with request IDs for reliability, per-request pricing for cost - and the measurement window.

A named next step

Model swaps point to Keys, caching and reordering point to Prompts, and routing recommendations point to Virtual Model Aliases with the suggested strategy and provider order spelled out.

Deliberate by design

Changes to model behavior, prompt structure, and routing are versioned, reviewable edits you make in the product - Insights never modifies anything itself.

Recommendation detail · low_latency

Latency
Scopeprod-chat · openai/gpt-5.5
Projected impactp95 −41%
ConfidenceHigh
Aprovider-a (current) p95 6.0s · p90 4.8s · p75 3.1s
Bprovider-b (proposed) p95 3.5s · p90 2.9s · p75 2.2s
Open Virtual Model Aliases

Suggested setup: Lowest Latency strategy, provider-b first

Measurement

Projected figures you can interrogate

Every card carries a projected figure frozen at generation and normalized to a per-week rate - with the formula, inputs, and measurement window shown, so the number can be checked, not just believed.

Real formulas, live pricing

Cost projections price each sampled request's actual tokens against the live pricing table - including per-customer negotiated rates - at generation time.

Confidence with a rationale

High, Medium, or Low, fixed per run, with a one-paragraph explanation of why - so you can weigh a strong signal differently from a thin one.

Comparable across time

Per-week normalization means a figure generated this run reads the same way as one from a month ago - no re-basing needed to compare opportunities.

flex_tier · prod-batch

High confidence
Projected savings
$410 / wk
Eligible volume
78% of 41K req/wk
Formula
eligible_vol × cost/req × flex_discount
Measurement window
Jul 27 – Aug 2
Generated
Aug 3, 2026 · frozen at generation
Next step
:flex slug or Virtual Model Alias
Tiers

Free insights, one advanced toggle

Free insights run automatically on qualifying traffic. Advanced cost insights are off by default and enabled with a single toggle - because they spend real tokens on your behalf, and you should choose that.

Free insights compared with Advanced cost insights
How the tiers compareFree insightsIncludedCost · Latency · ErrorsAdvanced cost insightsOpt-inAuto-evals · Compression
Coverage
Recommendation typesPrompt caching, cache reordering, Flex tier, low latency, error reductionEverything in Free, plus eval-verified model swaps and prompt compression candidates
Quality verificationNot includedLLM judge scores candidate outputs against your current model
Enablement & qualification
EnablementOn by defaultSingle toggle, off by default, with billing disclosure
Scope qualifies at1,000+ text-generation requests / week per Key + Model10,000+ requests or $50+ spend / week per Key + Model
What's analyzedMetadata: volumes, tokens, latency, errors, cache patternsAdditionally: top 3 system prompts per qualifying scope, sampled to comparison models
Cost
What you payFreeJudge and sample tokens, billed at standard rates and itemized on your invoice

Both tiers refresh on the same weekly run. Qualification thresholds are indicative and may be tuned server-side.

Use cases

Built for workloads with money on the table

From chat products to background pipelines, Insights turns the traffic you already route into concrete cost, latency, and reliability wins.

High-volume chat products

Stable system prompts reused thousands of times per day are prime caching territory. Insights finds the prefixes, sizes the reuse, and shows the caching math before you touch a prompt.

Batch and background pipelines

Summarization, enrichment, and other latency-tolerant jobs often qualify for Flex tier at roughly half the cost. Insights sizes the eligible volume and shows the switch - slug or alias with standard-tier fallback.

Latency-sensitive experiences

The same model can vary by seconds at p95 across providers. Insights spots the gap on your traffic and recommends the Lowest Latency alias configuration to close it.

Reliability-critical workloads

Recurring timeouts and error spikes surface as failover recommendations with the reference logs attached - request IDs, status, provider, time - plus the suggested priority-routing setup.

FAQ

Insights questions, answered

Free insights - cost, latency, and error recommendations - are enabled automatically for your org. Analysis runs once a week: your first set of recommendations appears after the next weekly run, and a Key + Model scope needs at least 1,000 text-generation requests in the past 7 days to qualify for analysis. The only setting is the Advanced cost insights toggle, which is off by default.

No. Analysis is read-only against your request metadata, and all recommendations are advisory - Insights never modifies your routing, prompts, or requests. Any change is one you make yourself in the product, as a normal versioned, reviewable edit.

Advanced cost insights verify model-swap and prompt compression candidates by making sample requests to comparison models and scoring the outputs with an LLM judge. Those judge and sample tokens are billed at standard rates and itemized on your invoice. The toggle is off by default, requires explicit enablement with the billing disclosure shown, and only runs on scopes exceeding 10,000 requests or $50 spend per week.

Each recommendation type has an explicit formula computed at generation time against live pricing - for example, prompt caching weighs prefix tokens, reuse rate, and the cache-read discount against the one-time write premium, and Flex tier multiplies eligible volume by cost per request and the tier discount. Figures are normalized to a per-week rate and frozen in the snapshot, with the formula, inputs, and measurement window shown in the flyout so the projection can be verified against your own data.

Each recommendation names its surface. Model swaps are made in Keys; prompt caching and cache reordering are made in Prompts, where prompt versions are retained so a change can be reverted; and Flex tier, low latency, and error reduction changes are made in Virtual Model Aliases (or, for Flex, by appending the :flex slug), with the suggested strategy and provider order shown in the flyout. Review the evidence, make the change, and the next weekly run measures the new baseline.

Weekly. Each run produces a fresh set of point-in-time snapshots that supersede the previous run, and the run metadata - last generated, measurement window, refresh cadence - is shown above every card. If a run finds nothing new, the feed says exactly that.

Find out what your traffic is trying to tell you

Free insights are already included with FastRouter. Route a week of qualifying traffic and your first recommendations arrive with the next run.