Kimi K3 vs GLM-5.3-Flash

Side-by-side pricing per provider, context window, benchmark scores and speed, all on one OpenAI-compatible API.

Metric
Kimi K3MoonshotAImoonshotai/kimi-k3
GLM-5.3-FlashZ.aiz-ai/glm-5.3-flash
Creator
MoonshotAI
Z.ai
Released
Jul 16, 2026
Aug 26, 2026
Context window
1.05M
1.31M (best)
Max output
1.05M (best)
131K
Input / 1MLowest provider price
$2.70$2.70–$3.00 across 4
$0.07$0.07–$0.15 across 5 (best)
Output / 1MLowest provider price
$13.50$13.50–$15.00 across 4
$0.25$0.25–$0.50 across 5 (best)
Providers
4
5 (best)
IntelligenceArtificial Analysis index
43.6 (best)
41.8
CodingArtificial Analysis index
76.2 (best)
71.5
LatencyMedian time to first token
7.4s
—
ThroughputMedian tokens per second
<1 t/s
—
Accepts
Text, Image
Text, Image, Video
Produces
Text
Text
Tool calling
Yes
Yes
Structured output
Yes
Yes
Reasoning
No
Yes

Kimi K3 vs GLM-5.3-Flash: summary

Kimi K3 and GLM-5.3-Flash are both available through FastRouter's OpenAI-compatible API, so switching between them is a change of model id, not a new integration.

Kimi K3, from MoonshotAI, has a 1,048,576-token context window, costs from $2.70/1M input and $13.50/1M output tokens across 4 providers and scores 43.6 on the Artificial Analysis Intelligence Index.

GLM-5.3-Flash, from Z.ai, has a 1,310,720-token context window, costs from $0.07/1M input and $0.25/1M output tokens across 5 providers and scores 41.8 on the Artificial Analysis Intelligence Index.

GLM-5.3-Flash is the cheapest on input, 39× cheaper than the next model, GLM-5.3-Flash has the largest context window, Kimi K3 scores highest for intelligence and Kimi K3 leads on coding.

Switch with one line

client = OpenAI(base_url="https://api.fastrouter.ai/api/v1", api_key="<FASTROUTER_API_KEY>") client.chat.completions.create(model="moonshotai/kimi-k3", messages=[...]) # or client.chat.completions.create(model="z-ai/glm-5.3-flash", messages=[...])
Full API docs

Frequently asked questions

More comparisons

Benchmark scores sourced from ArtificialAnalysis