DeepSeek Word Limit by Model (2026)

DeepSeek changed the economics of frontier AI in early 2025 and stayed cheap. Here's what each current model accepts, what it costs, and why V4.1 Flash at $0.30 per million input tokens is eating the market.

Quick Answer

DeepSeek V4 Pro and DeepSeek V4.1 Flash (the newest model, released September 10, 2026) both accept 1,000,000 tokens (~750K words) with up to 384K tokens of output. They are the only two models the DeepSeek API serves now: the older V3.2 (128K) and R1 (64K) models behind the deepseek-chat and deepseek-reasoner names were discontinued on July 24, 2026. At $0.30 per million input tokens (peak rate), V4.1 Flash is roughly 17x cheaper than Claude Opus 5.5 for equivalent context, which is why it dominates high-volume workloads.

DeepSeek context windows by model

ModelInput tokensMax outputReleased
DeepSeek V4 Pro1,000,000384,000Aug 2026
DeepSeek V4.1 Flash1,000,000384,000Sep 2026

Specs from DeepSeek API documentation, October 2026. The older V3.2, R1, V3.1 and V3 models are no longer served by the API and are not listed. DeepSeek moved to peak/off-peak billing on August 16, 2026; off-peak requests run about half the listed rate.

The DeepSeek pricing story

DeepSeek R1's launch in January 2025 is often called the "DeepSeek moment" because it demonstrated ChatGPT-level reasoning at a fraction of the training and API cost. The pricing held. Current rates:

ModelInput / 1MCache hit / 1MOutput / 1M
DeepSeek V4 Pro$1.32$0.044$3.96
DeepSeek V4.1 Flash$0.30$0.006$1.20

For reference: Claude Opus 5.5 is $4 per million input tokens. DeepSeek V4.1 Flash is $0.30 and V4 Pro is $1.32 at peak rates. That is a 17x and 4x difference respectively. Output costs $1.20 per million tokens on V4.1 Flash and $3.96 on V4 Pro.

The cache hit discount is the detail that actually changes the math. If your prompts share a common prefix (system prompt, tool definitions, a reference document), cached input tokens cost 90% less. A production app with a well-structured system prompt sees effective input costs below $0.05 per million tokens on V4. That is approaching commodity pricing.

The V4 output budget and the retired R1

DeepSeek R1 (January 2025) was unusual for giving 64K tokens of output on top of a 64K input window, and the older V3.2 capped output at 8K. Both were served through the deepseek-reasoner and deepseek-chat names, which DeepSeek discontinued on July 24, 2026. Today's V4 Pro and V4.1 Flash allow up to 384K tokens of output.

This matters for reasoning workloads. If you're asking a model to solve a multi-step problem with step-by-step working shown, the reasoning chain itself can consume 10K-40K tokens, so a generous output ceiling is what keeps long answers from being cut off.

When DeepSeek V4 is the right pick

V4 scores 81% on SWE-bench Verified (vs V3's 69%) and holds its own against GPT-5 on general benchmarks. At $0.30 input, it's the best price-to-quality ratio on the market for most production workloads. Specifically strong for:

  • Any high-volume workflow where per-token cost dominates (classification, extraction, basic Q&A)
  • Code agents and coding assistants (V4's coding scores are competitive with GPT-4 tier)
  • Long-document summarization with the 1M token context
  • Multilingual applications (DeepSeek was trained heavily on Chinese and English, strong on both)
  • Startups that cannot afford Claude Opus pricing but need frontier-tier quality

Where V4 is not the right pick: workloads that genuinely need the deepest reasoning (use V4 Pro or Claude Opus), vision-heavy tasks (DeepSeek's multimodal is behind GPT-4o and Gemini), or enterprise environments with data-residency concerns about servers hosted in mainland China.

The statelessness gotcha

DeepSeek's API is stateless. There is no persistent conversation memory. Every multi-turn chat requires re-sending the full conversation history in each API call. For long sessions, this is expensive even at DeepSeek's low rates because token volume grows quadratically with turn count.

Workarounds: use context caching aggressively for shared prefixes, summarize older turns instead of replaying verbatim, or use conversation-summarization techniques to compress the history. The API is powerful but you're responsible for managing context yourself.

See DeepSeek cost estimates for your actual prompt

Our AI prompt counter shows token count and input cost across DeepSeek and 9 other models

AI Prompt Word Counter

FAQ

What is DeepSeek's word limit?

V4 Pro and V4.1 Flash accept about 750,000 words (1M tokens) each, with up to 384K tokens of output. The older V3.2 (128K) and R1 (64K) API models were discontinued on July 24, 2026.

Why is DeepSeek so much cheaper than Claude or GPT?

Lower training costs (DeepSeek pioneered efficient MoE and FP8 training), simpler deployment, and a deliberate pricing strategy to capture market share. The quality is genuinely frontier-tier; the pricing isn't an accident.

Can I trust DeepSeek with sensitive data?

Their privacy policy states data may be stored on servers in mainland China. For sensitive or regulated data, check whether that meets your compliance requirements. Alternatively, self-host DeepSeek weights (they're open-source under MIT License) on your own infrastructure.

Can I still use DeepSeek V3.2 or R1 through the API?

No. The deepseek-chat and deepseek-reasoner names (V3.2 and R1) were discontinued on July 24, 2026. The API now serves DeepSeek V4 Pro and V4.1 Flash.

How long can DeepSeek's output be?

V4 Pro and V4.1 Flash allow up to 384K tokens per response. The older V3.2 capped output at 8,000 tokens. For very long outputs, chunk the task or use continue prompts.

Related Tools