Benchmarks

Z.ai: GLM 5.3 Flash

z-ai/glm-5.3-flash

GLM 5.3 Flash (Z.ai) is a low-cost open-weight model priced near the bottom of the board, at $0.15/$0.5 per million input/output tokens, served by 24 independent hosts. For the money it delivers #35 of 256 on the TryAii Score. Its best result is head-to-head chat (#9 on Chatbot Arena Elo (Code)). It's reasonably quick at roughly 44-125 tokens/sec.

Prompt $/1M
$0.150
Completion $/1M
$0.500
Context
1,048,576 tok
Best provider
269 tok/s · CoreWeave

Specs

Model id
z-ai/glm-5.3-flash
Display name
Z.ai: GLM 5.3 Flash
Modality
text+image+video->text
Tokenizer
Other
Instruct type
—
Context length
1,048,576
Max completion
943,717
Prompt $/1M
$0.150
Completion $/1M
$0.500
Free tier
no

Capabilities

  • Function callingno
  • JSON modeno
  • System promptno
  • Visionyes
  • Moderatedno

Benchmark scores

Each axis is one benchmark, normalized to the best model in the field.

No benchmark data for this model.

Speed by provider

Providertok/s
CoreWeavefastest269
Modal177
Parasail163
BaseTen146
Together140
Reka136
Sail Research133
DekaLLM125
Friendli125
Venice112
Inceptron109
InferenceNet107
Fireworks102
DeepInfra90
Morph83
Crusoe78
Decart62
Relace57
Artificial Analysis55
Z.AI51
DigitalOcean47
SiliconFlow46
Novita43
AtlasCloud42
StreamLake39
GMICloud36
Phala26
OpenInference21
Near AI20
Wafer15