Z.ai logo

Z.ai AI models

9 Z.ai models on FastRouter, all behind one OpenAI-compatible API. Compare pricing, context windows and benchmarks, then open any model for its providers and code samples.

Filter in catalog
Models
9
Largest context
1.31M
GLM-5.3-Flash
Lowest input /1M
$0.07
GLM-5.3-Flash
Top intelligence
44.8
GLM-5.3

All Z.ai models

Z.ai logo

GLM-5.3-FlashX is Z.ai's high-speed serving variant of GLM-5.3-Flash, delivering inference speeds of up to 200 tokens/s for faster responses. It inherits Flash's 320B-total/18B-active MoE design, native multimodality, hybrid sparse and linear attention architecture, and 1M-token context window. FlashX runs on a purpose-built serving stack (SGLang-based inference engine, W8A8 quantization, mixed INT8/FP8/BF16 cache quantization, and encode-prefill-decode disaggregation), reporting roughly 3x end-to-end throughput over Z.ai's initial baseline on the same hardware. It is suited for interactive coding agents, browser/computer-use loops, visual coding, and other multi-step workflows where latency compounds.

z-ai/glm-5.3-flashxSep 18, 2026
Context
1.05M
Price /1M
$0.37 in$1.25 out
Intel
—
Z.ai logo

GLM-5.3-Flash is Z.ai’s flagship natively multimodal GLM-5-series model with a 1M-token context window, accepting text plus images and video and producing text outputs. It targets fast, high-quality general and coding assistance with strong reasoning at “flash” latency and is released as open weights under the MIT license.

z-ai/glm-5.3-flashAug 26, 2026
Context
1.31M
Price /1M
$0.07 in$0.25 out
Intel
41.8
Z.AI logo
GLM-5.3#10 Intelligence

GLM-5.3 is Z.ai’s latest GLM-5-series flagship, a 743B-parameter text-only foundation model post-trained from GLM-5.2 for frontier-level coding, long-horizon agentic workflows, and defensive cybersecurity analysis. It inherits GLM-5’s 200K-token context window and up to 128K-token outputs, supports advanced reasoning and tool use, and is initially exposed through Z.ai’s coding plan and forthcoming API access.

z-ai/glm-5.3Aug 18, 2026
Context
1.05M
Price /1M
$0.56 in$2.50 out
Intel
44.8
Z.AI logo

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

z-ai/glm-5.2Jun 16, 202636s latency
Context
1.05M
Price /1M
$0.56 in$1.80 out
Intel
33.7
Z.AI logo

zai-org/GLM-5.1 is Z.AI's (formerly Zhipu AI) open-weight, instruction-tuned coding flagship model—a refreshed upgrade over GLM-5—excelling in agentic engineering, complex system programming, and long-horizon tasks with SOTA open-source performance approaching Claude Opus 4.6.

z-ai/glm-5.1Apr 7, 20267.5s latency
Context
203K
Price /1M
$1.05 in$3.50 out
Intel
26.1
Z.AI logo

GLM-5 is Z.ai’s flagship open-source foundation model, built for complex system design and long-horizon agent workflows. Aimed at expert developers, it delivers production-grade results on large-scale programming tasks and competes with top closed-source models. With strong agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 goes beyond code generation to help design, build, and execute complete systems end-to-end.

z-ai/glm-5Feb 11, 202628s latency
Context
205K
Price /1M
$0.60 in$2.08 out
Intel
27.9
Z.AI logo

GLM-Image is Z.AI's text-to-image generation model that quickly and accurately understands text descriptions to produce precise, personalized, high-quality images. It supports an 'hd' mode for more detailed and consistent output (~20s generation time) and a 'standard' mode optimized for faster generation (~5-10s), along with flexible custom resolutions from 1024px to 2048px per side.

z-ai/glm-imageJan 14, 2026
Context
4K
Price /1M
— in$0.01/img out
Intel
—
Z.AI logo
GLM 4.7#4 Math

GLM-4.7 is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

z-ai/glm-4.7Dec 22, 20253.8s latency
Context
203K
Price /1M
$0.40 in$1.75 out
Intel
22.2
Z.AI logo

GLM-4.6 is the latest iteration in the GLM series by Zhipu AI, designed as a large language model with about 355 billion parameters in a Mixture of Experts (MoE) architecture. It is optimized for various complex tasks including real-world coding, long-context processing, advanced reasoning, intelligent agent applications, and refined writing. The model features an expanded context window of 200,000 tokens (up from 128,000 in GLM-4.5), allowing it to handle longer and more complex interactions such as extensive documents or multi-turn conversations.

z-ai/glm-4.6Sep 30, 20258.4s latency
Context
203K
Price /1M
$0.43 in$1.74 out
Intel
18.5

Frequently asked questions

Model scores sourced from ArtificialAnalysis

Z.ai AI Models: API Pricing & Benchmarks | FastRouter.ai