GLM 5.3 Flash API
Fast multimodal model with enforced structured output. Reads text, images and video.
$0.150 per 1M input and $0.500 per 1M output. 1M context.
Call it in three lines
from openai import OpenAI
client = OpenAI(
base_url="https://api.tokenify.dev/v1",
api_key=os.environ["TOKENIFY_API_KEY"],
)
response = client.chat.completions.create(
model="z-ai/glm-5.3-flash",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)What it costs
A typical agent turn — a 4,000-token prompt with 2,500 of it cached, producing 600 tokens — costs $0.00060. A million of those turns costs $600.
- Input / 1M
- $0.150
- Output / 1M
- $0.500
- Cache read / 1M
- $0.030
How it compares
| Model | Input | Output | Cache read | Context |
|---|---|---|---|---|
| GLM 5.3 Flash | $0.150 | $0.500 | $0.030 | 1M |
| GLM 5.2 | $0.700 | $2.200 | $0.260 | 1M |
Peak rates, per 1M tokens, from the live catalogue. All models · the full rate card.
Questions about this model
Is this the same model as Zhipu's own API?
Yes — GLM 5.3 Flash is Zhipu's model, served through our endpoint rather than theirs. We do not fine-tune, quantise or substitute it. What differs is the price, the API formats on offer and how it is billed.
Can I use it with Claude Code?
Yes. GLM 5.3 Flash answers on the Anthropic Messages API as well as the OpenAI one, so Claude Code reaches it with three environment variables and no router. The Claude Code page has the setup and the compatibility results.
Do I need a subscription?
No. Billing is per token from prepaid credit, with no subscription, no seats and no minimum. Credit does not expire.
Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.