Summary

DeepSeek moved V4-Pro from preview to GA across app, web, and API (deepseek-v4-pro), with native OpenAI Responses API support and a one-click Codex configuration. Thinking-effort levels (low/high/max) were added for V4-Pro and V4-Flash. Peak/off-peak pricing takes effect 2026-08-16T16:00Z, with off-peak at 50% of peak.

Why it matters
Responses-API compatibility makes DeepSeek a near drop-in inside OpenAI-native agent stacks. The confirmed peak/off-peak rates are a large increase over the current flat rates: V4-Pro output rises from $0.87 to $3.96/Mtok at peak and $1.98 off-peak, effective 2026-08-16T16:00Z. Budget-sensitive agent workloads now need time-of-day scheduling or a fresh cost review.
Technical details
Terminal Bench 2 1 87.9
Hle Without Tools 42.7
Hle With Tools 60
Deepswe 62.7
Cybergym 83.3
Responses API Compatible
Off Peak Pricing Effective 2026-08-16T16:00Z
Pricing Usd Per Mtok
Until 2026 08 16T16Z {"deepseek_v4_pro":{"input_cache_hit":0.003625,"input_cache_miss":0.435,"output":0.87},"deepseek_v4_flash":{"input_cache_hit":0.0028,"input_cache_miss":0.14,"output":0.28}}
From 2026 08 16T16Z {"peak_hours_utc":"01:00-04:00 and 06:00-10:00 (7h/day); all other hours off-peak","deepseek_v4_pro_peak":{"input_cache_hit":0.044,"input_cache_miss":1.32,"output":3.96},"deepseek_v4_pro_off_peak":{"input_cache_hit":0.022,"input_cache_miss":0.66,"output":1.98},"deepseek_v4_flash_peak":{"input_cache_hit":0.014,"input_cache_miss":0.44,"output":1.32},"deepseek_v4_flash_off_peak":{"input_cache_hit":0.007,"input_cache_miss":0.22,"output":0.66}}
Updates
2026-08-16 Exact peak/off-peak rates confirmed from official pricing docs during 2026-08-16T00:00Z run. Net effect is a substantial price increase, not a discount: even off-peak V4-Pro output is ~2.3x the current flat rate, and peak is ~4.6x. Peak windows align with China business hours (09:00-12:00 / 14:00-18:00 Beijing). Concurrency limits unchanged: flash 2500, pro 500.
Tags
deepseekv4-progaresponses-apipricingagents