Summary
DeepSeek moved V4-Pro from preview to GA across app, web, and API (deepseek-v4-pro), with native OpenAI Responses API support and a one-click Codex configuration. Thinking-effort levels (low/high/max) were added for V4-Pro and V4-Flash. Peak/off-peak pricing takes effect 2026-08-16T16:00Z, with off-peak at 50% of peak.
Why it matters
Responses-API compatibility makes DeepSeek a near drop-in inside OpenAI-native agent stacks. The confirmed peak/off-peak rates are a large increase over the current flat rates: V4-Pro output rises from $0.87 to $3.96/Mtok at peak and $1.98 off-peak, effective 2026-08-16T16:00Z. Budget-sensitive agent workloads now need time-of-day scheduling or a fresh cost review.
Technical details
| Terminal Bench 2 1 | 87.9 | ||||
|---|---|---|---|---|---|
| Hle Without Tools | 42.7 | ||||
| Hle With Tools | 60 | ||||
| Deepswe | 62.7 | ||||
| Cybergym | 83.3 | ||||
| Responses API Compatible | ✓ | ||||
| Off Peak Pricing Effective | 2026-08-16T16:00Z | ||||
| Pricing Usd Per Mtok |
|
Updates
2026-08-16 Exact peak/off-peak rates confirmed from official pricing docs during 2026-08-16T00:00Z run. Net effect is a substantial price increase, not a discount: even off-peak V4-Pro output is ~2.3x the current flat rate, and peak is ~4.6x. Peak windows align with China business hours (09:00-12:00 / 14:00-18:00 Beijing). Concurrency limits unchanged: flash 2500, pro 500.
Tags
deepseekv4-progaresponses-apipricingagents