Launched on August 13 after nearly four months in preview, DeepSeek-V4-Pro-0813 is available through the company’s app, web interface and API.

We’re launching DeepSeek-V4-Pro today!
We’re launching DeepSeek-V4-Pro today! Х

Stronger agent benchmarks, but not an outright lead

In DeepSeek’s own tests, V4 Pro scored 87.9 on Terminal Bench 2.1, close to Claude Fable 5’s 88.0. The company also reported substantial improvements on NL2Repo, CyberGym and DeepSWE.

These results do not establish a lead across all tasks. Most figures come from DeepSeek’s testing, while independent evaluator Artificial Analysis still ranks Fable 5 ahead on its broader Intelligence Index.

New API pricing for V4 Pro and V4 Flash. Source: DeepSeek
New API pricing for V4 Pro and V4 Flash. Source: DeepSeek

How the new API prices work

The revised pricing took effect on August 16 for both V4 Pro and V4 Flash. Off-peak requests cost half as much as peak-time requests. V4 Pro’s rates per million tokens are:

Token type Off-peak Peak
Cached input $0.022 $0.044
Uncached input $0.66 $1.32
Output $1.98 $3.96

Peak hours run from 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. All other periods, including weekends, receive off-peak pricing.

Compared with the previous rates of $0.435 for uncached input and $0.87 for output, those charges have increased by roughly 52–203% and 128–355%, respectively. The increase therefore depends on both the token category and when a request is processed.

V4 Pro remains substantially cheaper per token than Claude Fable 5, whose standard rates are $10 for input and $50 for output. That price gap does not necessarily translate into the same savings per completed task, where token usage and retries also matter.

DeepSeek expands beyond models

The release comes as DeepSeek expands its workforce and infrastructure. The company reportedly raised $7.4 billion in its first external funding round and has been pursuing another round targeting a valuation of roughly $74 billion.

DeepSeek is also recruiting chip engineers and developing its own inference chip to reduce reliance on suppliers such as Nvidia and Huawei. The project remains under development, rather than an immediate replacement for its existing hardware.