Z.ai: GLM 4.7 Flash
z-ai/glm-4.7-flash
GLM 4.7 Flash (Z.ai) is a low-cost open-weight model priced near the bottom of the board, at $0.06/$0.4 per million input/output tokens, served by 5 independent hosts. For the money it delivers #132 of 235 on the TryAii Score. Its best showing is #2 on Tau2-bench. It's reasonably quick at roughly 30-56 tokens/sec.
Prompt $/1M
$0.060
Completion $/1M
$0.400
Context
202,752 tok
Best provider
63 tok/s · DeepInfra
Specs
- Model id
- z-ai/glm-4.7-flash
- Display name
- Z.ai: GLM 4.7 Flash
- Modality
- text->text
- Tokenizer
- Other
- Instruct type
- —
- Context length
- 202,752
- Max completion
- 16,384
- Prompt $/1M
- $0.060
- Completion $/1M
- $0.400
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionno
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| DeepInfrafastest | 63 |
| Venice | 60 |
| Phala | 43 |
| Z.AI | 30 |
| Novita | 30 |
| Cloudflare | 26 |
| z-aifastest | — |
| Artificial Analysis | 0 |