Z.ai: GLM 5.3 Flash
z-ai/glm-5.3-flash
GLM 5.3 Flash (Z.ai) is a low-cost open-weight model priced near the bottom of the board, at $0.15/$0.5 per million input/output tokens, served by 24 independent hosts. For the money it delivers #35 of 256 on the TryAii Score. Its best result is head-to-head chat (#9 on Chatbot Arena Elo (Code)). It's reasonably quick at roughly 44-125 tokens/sec.
Prompt $/1M
$0.150
Completion $/1M
$0.500
Context
1,048,576 tok
Best provider
269 tok/s · CoreWeave
Specs
- Model id
- z-ai/glm-5.3-flash
- Display name
- Z.ai: GLM 5.3 Flash
- Modality
- text+image+video->text
- Tokenizer
- Other
- Instruct type
- —
- Context length
- 1,048,576
- Max completion
- 943,717
- Prompt $/1M
- $0.150
- Completion $/1M
- $0.500
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionyes
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| CoreWeavefastest | 269 |
| Modal | 177 |
| Parasail | 163 |
| BaseTen | 146 |
| Together | 140 |
| Reka | 136 |
| Sail Research | 133 |
| DekaLLM | 125 |
| Friendli | 125 |
| Venice | 112 |
| Inceptron | 109 |
| InferenceNet | 107 |
| Fireworks | 102 |
| DeepInfra | 90 |
| Morph | 83 |
| Crusoe | 78 |
| Decart | 62 |
| Relace | 57 |
| Artificial Analysis | 55 |
| Z.AI | 51 |
| DigitalOcean | 47 |
| SiliconFlow | 46 |
| Novita | 43 |
| AtlasCloud | 42 |
| StreamLake | 39 |
| GMICloud | 36 |
| Phala | 26 |
| OpenInference | 21 |
| Near AI | 20 |
| Wafer | 15 |