Two GLM-5.3 models: a text-only flagship, and a multimodal Flash about 89% cheaper
Asked:
“latest ai modles by z.ai”
As of 2026-09-19, Z.ai’s latest generation covers 2 of 2 current GLM-5.3-generation models: GLM-5.3, the text-only flagship for complex software engineering and long-horizon agents, and GLM-5.3-Flash, the low-cost, natively multimodal option. Both offer a 1,000,000-token context; prices below are USD per million tokens.
API list prices in USD per million tokens. Dark bars: GLM-5.3 (flagship, text-only). Light bars: GLM-5.3-Flash (multimodal). Lower price does not by itself mean the better model — the flagship targets harder engineering and agent work.
Flash costs about 89% less than the flagship on both standard input ($0.15 vs $1.40) and output ($0.50 vs $4.40) per million tokens.
docs.z.ai
GLM-5.3-Flash is the first native multimodal model in the GLM-5 series, taking video, image, text and file inputs; GLM-5.3 is text-only.
dev.to
Both ship downloadable weights, but licences differ: Flash is open-licensed, while GLM-5.3 uses Z.ai’s own licence with a restriction on model-as-a-service businesses exceeding $10 billion revenue over 12 months.
yuwish.kr
Output limits are close: 131,072 tokens for GLM-5.3 versus 128,000 for Flash, on the same 1,000,000-token context.
docs.z.ai
| Model | Positioning | Inputs | Context | Max output | Input $/M | Cached $/M | Output $/M | Weights | Official | Corroboration |
Data: 2 current GLM-5.3-generation models from Z.ai, checked 2026-09-19. Prices are API list prices in USD per million tokens (input, cached input, output); context and output limits in tokens. Each row carries an official docs.z.ai source and independent corroboration (dev.to, yuwish.kr). Nothing cut for space.