Umans DeepSeek V4 Pro
83.3tok/s
throughput · p50 · last 5 min
1.92s
TTFT · p50 · last 5 min
99.90%
uptime · 24h
DeepSeek V4 Pro, served from the official 0813 release: DeepSeek's flagship coding and reasoning MoE, built for long-horizon agentic coding and demanding tool-heavy workloads on a 1M-token context window. The 0813 release supersedes the April preview with substantially stronger agentic performance. Reasoning has three modes: non-think (none), think high (high, the default) and think max (max). Billed per token ($1.32 / $3.96 / $0.044 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability.
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/high/max
Weights
Trends
Speed over the last 90 days
peak 176.9 tok/s · Aug 13now 90.9 tok/s
90 days agopre-release before Aug 15, 2026today
now 1.30s
90 days agopre-release before Aug 15, 2026today
Changelog
Events for Umans DeepSeek V4 Pro
Aug 152026
Released pay-per-token: Umans DeepSeek V4 Pro Released
umans-deepseek-v4-pro-0813 joins the lineup as the long-context coding flagship: DeepSeek's official 0813 release of V4 Pro, the checkpoint that served here as the seat-gated pre-release lab since August 13, with a 1M context window and thinking at high effort by default (dial to max when a task deserves more). Billed per token: $1.32 / $3.96 / $0.044 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated and sunsets on August 23, 2026. The pre-release window's metrics stay on the model's status page as its pre-release period (before Aug 15). Served on our own GPU infrastructure with high availability.