DeepSeek V4.1 Flash API
Fast multimodal model. Reads text, images and video; cheapest per token we sell.
$0.150 per 1M input and $0.600 per 1M output at peak, and $0.090 / $0.360 off-peak. 1M context. 50% below DeepSeek's published rate.
Call it in three lines
from openai import OpenAI
client = OpenAI(
base_url="https://api.tokenify.dev/v1",
api_key=os.environ["TOKENIFY_API_KEY"],
)
response = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)What it costs
A typical agent turn — a 4,000-token prompt with 2,500 of it cached, producing 600 tokens — costs $0.00060. A million of those turns costs $600.
- Input / 1M
- $0.300$0.150
- Output / 1M
- $1.200$0.600
- Cache read / 1M
- $0.006
50% offoff DeepSeek’s published rate on input and output. Cache reads pass through undiscounted — they already cost a fraction of a fresh token.
Off-peak pricing
Peak hours are Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8). Everything else — nights, weekends and the two hours over lunch — is off-peak, and a request is billed at the rate in force when it is made. Your request log records which rate applied.
- Input / 1M
- $0.150$0.090
- Output / 1M
- $0.600$0.360
- Cache read / 1M
- $0.003
How it compares
| Model | Input | Output | Cache read | Context |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.150 | $0.600 | $0.006 | 1M |
| DeepSeek V4 Pro | $0.660 | $1.980 | $0.044 | 1M |
| DeepSeek V4 Flash | $0.220 | $0.660 | $0.014 | 1M |
Peak rates, per 1M tokens, from the live catalogue. All models · the full rate card · how sellers of this model compare.
Against the other fast models
| Model | Input | Output | Cache read | API format |
|---|---|---|---|---|
| GPT-6 LunaOpenAI | $0.100 | $0.500 | $0.010 | OpenAI |
| Claude Haiku 4.5Anthropic | $1.000 | $5.000 | $0.100 | Anthropic |
| Gemini 3.8 FlashGoogle | $0.750 | $3.750 | $0.075 | |
| DeepSeek V4.1 FlashTokenify, peak | $0.150 | $0.600 | $0.006 | OpenAI and Anthropic |
| DeepSeek V4.1 FlashTokenify, off-peak | $0.090 | $0.360 | $0.003 | OpenAI and Anthropic |
Per 1M tokens. Our rows come from the live catalogue at build time. Other vendors’ prices were read from their own pages and last checked 2026-10-01: openai.com (2026-09-30), claude.com (2026-10-01), ai.google.dev (2026-10-01).
Gemini 3.8 Flash is on a promotional rate until 2026-12-31; it becomes $1.50 / $7.50 / $0.15 on 1 January 2027.
Questions about this model
Is this the same model as DeepSeek's own API?
Yes — DeepSeek V4.1 Flash is DeepSeek's model, served through our endpoint rather than theirs. We do not fine-tune, quantise or substitute it. What differs is the price, the API formats on offer and how it is billed.
What are the peak hours?
Peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8). Every other hour — nights, weekends and the two hours over lunch — is off-peak, and a request is billed at the rate in force when it is made. Your request log records which rate applied.
How much cheaper is this than DeepSeek's list price?
50% on input and output, blended three parts prompt to one part completion. Cache reads are not discounted: they are billed at DeepSeek's published cache rate, which is already a small fraction of a fresh token.
Can I use it with Claude Code?
Yes. DeepSeek V4.1 Flash answers on the Anthropic Messages API as well as the OpenAI one, so Claude Code reaches it with three environment variables and no router. The Claude Code page has the setup and the compatibility results.
Do I need a subscription?
No. Billing is per token from prepaid credit, with no subscription, no seats and no minimum. Credit does not expire.
Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.