DeepSeek will stop distinguishing between peak and off-peak pricing for its API on weekends starting August 23, with usage on Saturdays and Sundays now charged at the lower off-peak rate regardless of time of day, the Chinese AI lab said in a notice. The change, effective from midnight Beijing time, partially unwinds a new peak-pricing structure the company introduced just a week earlier, on August 16, that had drawn sharp criticism from developers.

That August 16 overhaul replaced DeepSeek's long-standing flat API rates with a peak-and-off-peak schedule for its V4-Flash and V4-Pro models, with peak hours running from 01:00 to 04:00 and 06:00 to 10:00 UTC — windows that correspond to Chinese business hours but fall largely outside standard working hours in the United States and Europe. Under the new schedule, V4-Pro output token costs rose from a flat $0.87 per million to $1.98 off-peak and $3.96 at peak; V4-Flash output rose from $0.28 to $0.66 off-peak and $1.32 at peak — increases developers calculated at up to 1,100 percent depending on model, token type and time of use.

The repricing followed the general availability launch of DeepSeek-V4-Pro on August 16, which brought major agent-workflow upgrades and flexible reasoning-effort controls, alongside native support for OpenAI's Responses API format. Even at the new, higher rates, DeepSeek's models remain dramatically cheaper than Western frontier alternatives — V4-Flash output, for instance, runs roughly 23 to 45 times cheaper than GPT-5.5 at comparable context lengths, according to independent pricing trackers.

image.png

Developer reaction to the August 16 change was pointed, with some on social media noting that DeepSeek appeared to be effectively subsidising US and European daytime usage while raising costs disproportionately for developers operating in or near Asian time zones. The weekend exemption announced August 23 addresses one slice of that criticism, though weekday peak pricing remains in effect for now, with DeepSeek stating that charges incurred before the August 23 cutover would continue to be settled under the prior pricing standard.

The pricing shifts arrive against a backdrop of surging demand for Chinese open-weight models: OpenRouter usage data from late July showed Chinese models processing 28.13 trillion tokens against 4.38 trillion for US models over a comparable window — more than a sixfold gap that has intensified scrutiny of how sustainable China's aggressively low AI pricing strategy actually is as usage scales.

For developers building on DeepSeek's API, the episode is a reminder that even the AI industry's most aggressive low-cost providers are not immune to the underlying economics of serving increasingly capable, increasingly demanded models — and that pricing structures introduced with little warning can shift again just as quickly, making workload scheduling and multi-provider fallback strategies an increasingly standard part of production AI engineering rather than an optional optimisation.