Tokenify

GLM 5.3 Flash API

Fast multimodal model with enforced structured output. Reads text, images and video.

$0.150 per 1M input and $0.500 per 1M output. 1M context.

Call it in three lines

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tokenify.dev/v1",
    api_key=os.environ["TOKENIFY_API_KEY"],
)

response = client.chat.completions.create(
    model="z-ai/glm-5.3-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

What it costs

A typical agent turn — a 4,000-token prompt with 2,500 of it cached, producing 600 tokens — costs $0.00060. A million of those turns costs $600.

Input / 1M
$0.150
Output / 1M
$0.500
Cache read / 1M
$0.030

How it compares

ModelInputOutputCache readContext
GLM 5.3 Flash$0.150$0.500$0.0301M
GLM 5.2$0.700$2.200$0.2601M

Peak rates, per 1M tokens, from the live catalogue. All models · the full rate card.

Questions about this model

Is this the same model as Zhipu's own API?

Yes — GLM 5.3 Flash is Zhipu's model, served through our endpoint rather than theirs. We do not fine-tune, quantise or substitute it. What differs is the price, the API formats on offer and how it is billed.

Can I use it with Claude Code?

Yes. GLM 5.3 Flash answers on the Anthropic Messages API as well as the OpenAI one, so Claude Code reaches it with three environment variables and no router. The Claude Code page has the setup and the compatibility results.

Do I need a subscription?

No. Billing is per token from prepaid credit, with no subscription, no seats and no minimum. Credit does not expire.

Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.