From your own traffic
Every recommendation is generated from a weekly analysis of your actual requests - real volumes, real prompts, real provider behavior. No generic best-practice checklists.
Insights analyzes your FastRouter workloads once a week and turns them into evidence-backed recommendations on cost, latency, and errors - each with the projected impact, the evidence behind it, and exactly where to make the change.
Free insights on by default · Read-only - Insights never modifies your traffic
Recommendations
Weekly run · Aug 3Savings found
$2,140/wk
P95 latency
1.8s
Error rate
0.9%
Enable prompt caching on prod-chat · claude-sonnet-4.6
Stable 6.2K-token prefix reused 14× per session
Move batch summarization to Flex tier
78% of requests tolerate queueing · est. $410/wk
Route gpt-5.5 through fastest provider
p95 6.0s → 3.5s on alternate provider
Add failover for gemini-3.1-pro workloads
3 timeout spikes in measurement window
You already have the metrics. Insights reads them for you - and hands back specific actions with a dollar or performance figure attached, plus the evidence behind each one.
Every recommendation is generated from a weekly analysis of your actual requests - real volumes, real prompts, real provider behavior. No generic best-practice checklists.
Insights recommends; you decide. Every card opens into the sampled requests, before/after figures, and confidence rationale - then points to the exact surface where the change is made.
Cost, latency, and error insights are included and enabled automatically. Advanced cost insights - model-swap and compression candidates verified by automatic evaluations - are a separate opt-in with clear billing disclosure.
A weekly job analyzes every qualifying Key + Model scope, freezes what it finds as a point-in-time snapshot, and publishes it to your Insights feed with the run metadata on top.
Step 1 · Your traffic
Step 2 · FastRouter
Step 3 · Back to you
Savings, candidate models, measurement period, and confidence are fixed when the run generates them - figures never silently drift under you. The next weekly run supersedes the previous one, so you always see the latest run only, with its metadata above every card.
Run metadata
SnapshotEach type is scoped to the grain where it was detected. All recommendations are advisory - Insights shows you the evidence and the projected impact, and you make the change where it belongs.
model_swap finds cheaper models that match your quality bar, prompt_caching and cache_reordering turn reused prefixes into cache reads, and flex_tier moves latency-tolerant traffic to a discounted tier.
low_latency spots another provider serving the same model measurably faster, comparing current vs proposed at p95/p90/p75 with a suggested Lowest Latency alias configuration.
error_reduction flags recurring timeouts and errors with the reference logs attached, and suggests a Priority Routing setup: healthy provider first, fallback on failure.
Recommendation types
6 detectorsA card is only as good as the evidence behind it. Each flyout lays out what was measured, what's proposed, and precisely where in FastRouter you make the change.
Sampled requests, current vs proposed figures - p95/p90/p75 for latency, reference error logs with request IDs for reliability, per-request pricing for cost - and the measurement window.
Model swaps point to Keys, caching and reordering point to Prompts, and routing recommendations point to Virtual Model Aliases with the suggested strategy and provider order spelled out.
Changes to model behavior, prompt structure, and routing are versioned, reviewable edits you make in the product - Insights never modifies anything itself.
Recommendation detail · low_latency
LatencySuggested setup: Lowest Latency strategy, provider-b first
Every card carries a projected figure frozen at generation and normalized to a per-week rate - with the formula, inputs, and measurement window shown, so the number can be checked, not just believed.
Cost projections price each sampled request's actual tokens against the live pricing table - including per-customer negotiated rates - at generation time.
High, Medium, or Low, fixed per run, with a one-paragraph explanation of why - so you can weigh a strong signal differently from a thin one.
Per-week normalization means a figure generated this run reads the same way as one from a month ago - no re-basing needed to compare opportunities.
flex_tier · prod-batch
High confidenceFree insights run automatically on qualifying traffic. Advanced cost insights are off by default and enabled with a single toggle - because they spend real tokens on your behalf, and you should choose that.
| How the tiers compare | Free insightsIncludedCost · Latency · Errors | Advanced cost insightsOpt-inAuto-evals · Compression |
|---|---|---|
| Coverage | ||
| Recommendation types | Prompt caching, cache reordering, Flex tier, low latency, error reduction | Everything in Free, plus eval-verified model swaps and prompt compression candidates |
| Quality verification | Not included | LLM judge scores candidate outputs against your current model |
| Enablement & qualification | ||
| Enablement | On by default | Single toggle, off by default, with billing disclosure |
| Scope qualifies at | 1,000+ text-generation requests / week per Key + Model | 10,000+ requests or $50+ spend / week per Key + Model |
| What's analyzed | Metadata: volumes, tokens, latency, errors, cache patterns | Additionally: top 3 system prompts per qualifying scope, sampled to comparison models |
| Cost | ||
| What you pay | Free | Judge and sample tokens, billed at standard rates and itemized on your invoice |
Both tiers refresh on the same weekly run. Qualification thresholds are indicative and may be tuned server-side.
From chat products to background pipelines, Insights turns the traffic you already route into concrete cost, latency, and reliability wins.
Stable system prompts reused thousands of times per day are prime caching territory. Insights finds the prefixes, sizes the reuse, and shows the caching math before you touch a prompt.
Summarization, enrichment, and other latency-tolerant jobs often qualify for Flex tier at roughly half the cost. Insights sizes the eligible volume and shows the switch - slug or alias with standard-tier fallback.
The same model can vary by seconds at p95 across providers. Insights spots the gap on your traffic and recommends the Lowest Latency alias configuration to close it.
Recurring timeouts and error spikes surface as failover recommendations with the reference logs attached - request IDs, status, provider, time - plus the suggested priority-routing setup.
Free insights - cost, latency, and error recommendations - are enabled automatically for your org. Analysis runs once a week: your first set of recommendations appears after the next weekly run, and a Key + Model scope needs at least 1,000 text-generation requests in the past 7 days to qualify for analysis. The only setting is the Advanced cost insights toggle, which is off by default.
No. Analysis is read-only against your request metadata, and all recommendations are advisory - Insights never modifies your routing, prompts, or requests. Any change is one you make yourself in the product, as a normal versioned, reviewable edit.
Advanced cost insights verify model-swap and prompt compression candidates by making sample requests to comparison models and scoring the outputs with an LLM judge. Those judge and sample tokens are billed at standard rates and itemized on your invoice. The toggle is off by default, requires explicit enablement with the billing disclosure shown, and only runs on scopes exceeding 10,000 requests or $50 spend per week.
Each recommendation type has an explicit formula computed at generation time against live pricing - for example, prompt caching weighs prefix tokens, reuse rate, and the cache-read discount against the one-time write premium, and Flex tier multiplies eligible volume by cost per request and the tier discount. Figures are normalized to a per-week rate and frozen in the snapshot, with the formula, inputs, and measurement window shown in the flyout so the projection can be verified against your own data.
Each recommendation names its surface. Model swaps are made in Keys; prompt caching and cache reordering are made in Prompts, where prompt versions are retained so a change can be reverted; and Flex tier, low latency, and error reduction changes are made in Virtual Model Aliases (or, for Flex, by appending the :flex slug), with the suggested strategy and provider order shown in the flyout. Review the evidence, make the change, and the next weekly run measures the new baseline.
Weekly. Each run produces a fresh set of point-in-time snapshots that supersede the previous run, and the run metadata - last generated, measurement window, refresh cadence - is shown above every card. If a run finds nothing new, the feed says exactly that.
Free insights are already included with FastRouter. Route a week of qualifying traffic and your first recommendations arrive with the next run.