How Many Words Can AI Remember?
By Munir Afridi · Updated September 2026 · 12 min read
Every AI provider advertises its context window in tokens, a unit nobody thinks in. This page converts the 2026 lineup into words, the unit you actually write in, so you can compare ChatGPT, Claude, and Gemini the way a writer would.
Quick Answer
The biggest AI context windows in September 2026 hold about 750,000 to 787,500 words. GPT-6 Astra leads at 1.05 million tokens (about 787,500 words). Claude Opus 5.5, Claude Sonnet 5, and Claude Fable 5.1 tie with Gemini 3.1 Pro's input side at 1 million tokens (about 750,000 words). The smallest mainstream model, Claude Haiku 4.5, holds 200,000 tokens, about 150,000 words. All figures use the standard 0.75-words-per-token ratio for English prose.
How many words can ChatGPT, Claude, and Gemini remember in 2026?
Here is the full 2026 lineup, ranked by word capacity. "Words" is the context window in tokens multiplied by 0.75, the standard ratio for English text used throughout this site's tokens to words converter. Release dates are when each model or tier first shipped.
| Model | Provider | Context window | ≈ Words | Released |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 1.05M tokens | 787,500 | Sep 3, 2026 |
| Claude Fable 5.1 | Anthropic | 1M tokens | 750,000 | 2026 |
| Claude Opus 5.5 | Anthropic | 1M tokens | 750,000 | Sep 22, 2026 |
| Gemini 3.1 Pro | 1M in / 64K out | 750,000 in | 2026 | |
| Claude Sonnet 5 | Anthropic | 1M tokens | 750,000 | 2026 |
| Gemini 3.8 Flash | 1M in / 66K out | 750,000 in | 2026 | |
| GPT-6 Sol | OpenAI | 1.05M tokens | 787,500 | Sep 3, 2026 |
| GPT-6 Luna | OpenAI | 1.05M tokens | 787,500 | Sep 3, 2026 |
| Claude Haiku 4.5 | Anthropic | 200K tokens | 150,000 | 2025 |
Sources: Claude Platform Docs (platform.claude.com, Sep 2026); OpenAI, "GPT-6 Astra: A new generation of intelligence" and "Introducing GPT-6 Sol and Luna" (openai.com, Sep 3, 2026); Google DeepMind Gemini 3.1 Pro model card (deepmind.google, 2026); Google AI for Developers Gemini API pricing (ai.google.dev, Sep 2026). Word figures rounded to the nearest 500 and use 0.75 words per token.
Two things stand out. First, every 2026 flagship clusters tightly between 750,000 and 787,500 words. Context windows stopped being the headline race they were in 2024 and 2025, when models jumped from 8,000 tokens to 200,000 in about two years. Second, the budget tiers, GPT-6 Luna, Gemini 3.8 Flash, and Claude Haiku 4.5, no longer mean a smaller window across the board. Luna matches Astra's full 787,500 words for a fraction of the price, which is the more useful story for anyone paying per token.
What is a token, and why isn't it a word?
A token is the unit a language model actually reads, and it usually is not a whole word. Common short words like "the", "and", or "cat" are typically 1 token each. Longer or less common words split into pieces: "unbelievable" might become "un", "believ", and "able", three tokens for one word. Punctuation and spaces can count as partial tokens too.
Across typical English prose, those splits and mergers average out to about 0.75 words per token, or roughly 1.33 tokens per word. That is the ratio behind every word figure on this page, and it is the same ratio our 1 million tokens to words guide uses. It holds well for plain English; code, non-English languages, and heavy formatting all tokenize differently, usually less efficiently, which is why a coding session fills a context window faster than a plain-text one of the same word count.
Want the exact count for something you are about to paste into a prompt? The AI prompt word counter shows live token and word counts across models as you type, and the tokens to words converter does the math for any token count you already have.
How fast have AI context windows grown?
Six years ago, a model forgetting the start of a conversation was the norm, not the exception. GPT-3 launched in 2020 with a 2,048-token window, about 1,500 words, roughly three pages. The climb from there to 750,000-plus words took a series of jumps, not a steady line.
| Model (launch year) | Context window | ≈ Words |
|---|---|---|
| GPT-3 (2020) | 2,048 tokens | ~1,500 |
| ChatGPT / GPT-3.5 (2022) | 4,096 tokens | ~3,000 |
| GPT-4 (2023) | 8,192 tokens (32K variant) | ~6,000 (~24,000) |
| Claude 2 (2023) | 100,000 tokens | ~75,000 |
| GPT-4 Turbo (2023) | 128,000 tokens | ~96,000 |
| Claude 3 (2024) | 200,000 tokens | ~150,000 |
| Gemini 1.5 Pro (2024) | 1,000,000 tokens | ~750,000 |
| 2026 flagships (this page) | 1M to 1.05M tokens | 750,000 to 787,500 |
The biggest single jump was Gemini 1.5 Pro's move to a 1 million token window in 2024, roughly 500 novels' worth of memory added in one release. Since then, growth has flattened. Every 2026 flagship in the table above sits within 5% of 750,000 words, which suggests providers have found the window size that covers nearly every real document a person hands an AI, and are now competing on price and reasoning quality instead of raw capacity.
How many book pages is each AI's context window?
Words are still abstract at six figures. Pages make it concrete. Using the same 275-words-per-page convention as our pages-to-words guides, here is what each context window holds in printed book pages and in average 80,000-word novels.
| Context window | Words | Book pages (~275 wpp) | 80,000-word novels |
|---|---|---|---|
| GPT-6 (Astra / Sol / Luna), 1.05M tokens | 787,500 | ~2,864 | ~9.8 |
| Claude Opus 5.5 / Sonnet 5 / Fable 5.1, 1M tokens | 750,000 | ~2,727 | ~9.4 |
| Gemini 3.1 Pro / 3.8 Flash input, 1M tokens | 750,000 | ~2,727 | ~9.4 |
| Claude Haiku 4.5, 200K tokens | 150,000 | ~545 | ~1.9 |
In practice, this means a single GPT-6 or Claude prompt can hold almost 10 novels, or a small library of quarterly reports, contracts, or codebase files, in one pass. Even Haiku 4.5, the smallest window here, still covers roughly two full novels, which is more than enough for most single documents. If you want to check how a specific manuscript or PDF stacks up in pages, the words to pages tool converts any word count using your own font and spacing.
Which AI model is cheapest for long documents?
Context window size and price move independently. A model family often reuses one window across several price tiers, so a smaller bill does not have to mean a smaller memory. Here is the full 2026 price list per million tokens, cheapest first.
| Model | Context window | Input / 1M tok | Output / 1M tok |
|---|---|---|---|
| GPT-6 Luna | 787,500 words | $0.10 | $0.50 |
| Gemini 3.8 Flash | 750,000 words | $0.75 | $3.75 |
| Claude Haiku 4.5 | 150,000 words | $1.00 | $5.00 |
| Claude Sonnet 5 | 750,000 words | $2.00 | $10.00 |
| GPT-6 Sol | 787,500 words | $2.00 | $10.00 |
| Gemini 3.1 Pro (≤200K) | 750,000 words | $2.00 | $12.00 |
| Claude Opus 5.5 | 750,000 words | $4.00 | $20.00 |
| GPT-6 Astra / Claude Fable 5.1 | 787,500 / 750,000 words | $10.00 | $50.00 |
Sources: Claude Platform Docs (Sep 2026); OpenAI GPT-6 Astra and Sol/Luna announcements (openai.com, Sep 3, 2026); Google AI for Developers Gemini API pricing (ai.google.dev, Sep 2026). Gemini 3.1 Pro rates rise to $4.00 input / $18.00 output per million tokens past 200K tokens in one request.
GPT-6 Luna is the standout: the full 787,500-word GPT-6 window at $0.10 per million input tokens, 100 times cheaper than Astra for the identical context length. That gap is the real lesson in 2026 pricing. Before picking a model by window size alone, check whether a cheaper tier in the same family already gives you that window.
How much does it cost to process a 100,000-word document with each model?
A 100,000-word manuscript, thesis, or codebase export is about 133,333 tokens at the 0.75-words-per-token ratio. Here is what one input pass costs on each model, smallest bill first.
| Model | Cost to input 100,000 words |
|---|---|
| GPT-6 Luna | ~$0.01 |
| Gemini 3.8 Flash | ~$0.10 |
| Claude Haiku 4.5 | ~$0.13 |
| Claude Sonnet 5 | ~$0.27 |
| GPT-6 Sol | ~$0.27 |
| Gemini 3.1 Pro | ~$0.27 |
| Claude Opus 5.5 | ~$0.53 |
| GPT-6 Astra / Claude Fable 5.1 | ~$1.33 |
The spread is over 100x between the cheapest and most expensive input cost for the exact same 100,000 words. For a one-off summary or extraction task, that difference is pocket change either way. Run it as a daily pipeline over thousands of documents, and the model you pick changes the monthly bill by orders of magnitude. Paste your own word count into the AI prompt word counter to price it against any of these models directly.
Does a bigger context window mean a smarter AI?
No. Context window measures memory, not intelligence. A model with a 750,000-word window can still misread a document if it reasons poorly, and a shorter-window model can outperform a longer one on the same task if it reasons better. Treat window size as a hard ceiling on how much you can hand the model in one go, not as a quality score.
Window size also is not the same as how well a model uses that window. Independent long-context evaluations have repeatedly found that recall weakens for information buried in the middle of a very long prompt, a pattern researchers call "lost in the middle." Feeding a model its full 750,000-word limit does not guarantee it weighs every word equally. For anything you need answered precisely, put the most important facts near the start or the end of the prompt, and verify pulled numbers against the source rather than trusting recall alone.
There is also a cost to using the whole window even when a model handles it well. Every token in a prompt gets billed and re-processed on each turn of a conversation, so a 700,000-word prompt is slower and pricier to work with than a 50,000-word one, even on a model rated for both. In practice, most people never come close to the ceiling: a novel manuscript, a year of email, or an entire codebase for a small app all fit inside 200,000 words, well under every model in this comparison. Reach for the largest windows when you genuinely need them, not as a default.
How do you check whether your document fits?
Three steps, and you do not need to know a single token count by heart.
- Get your word count. Paste the document into the word counter for an exact number, or use the words to pages tool if you only know the page count.
- Convert to tokens if you need the model's own unit. Divide your word count by 0.75, or let the tokens to words converter do it, to get an approximate token count for any model's API limit.
- Compare against the table above. If your document is under 150,000 words, every model here handles it. Past that, you need at least a Claude, Gemini, or GPT-6 tier with a 1 million-token class window, and past 787,500 words you will need to split the document across multiple prompts regardless of model.
Remember that a prompt also has to leave room for the model's reply. A 780,000-word input on a 787,500-word window leaves almost nothing for the answer, so treat these numbers as a ceiling for input plus output combined, not input alone, unless a provider explicitly lists separate input and output limits the way Gemini does.
Which AI should you use for long documents?
Match the model to the job rather than defaulting to whichever has the biggest number.
- A single massive file (a full book, a year of chat logs, a large codebase). GPT-6 Astra or Claude Fable 5.1 give the most headroom at 750,000 to 787,500 words in one prompt.
- Routine long documents (contracts, reports, theses, transcripts). Claude Sonnet 5, GPT-6 Sol, or Gemini 3.1 Pro cover almost any real file at a fraction of the flagship price.
- High-volume, budget-sensitive work (batch summarizing, tagging, extraction). GPT-6 Luna or Gemini 3.8 Flash keep the full million-token-class window at the lowest per-word cost in this lineup.
- Short prompts, fast replies (chat, quick edits, short-form content). Claude Haiku 4.5's 150,000-word window is still more than 500 pages, plenty for anything short of a full manuscript, and it is the fastest model here.
Whichever model you use, check your prompt length before you hit send. The word counter and AI prompt word counter both show a live count, and per-model detail lives on the Claude word limit, ChatGPT word limit, and Gemini word limit pages, or the full AI tools hub.
Frequently Asked Questions
How many words can ChatGPT remember at once?
GPT-6 Astra, OpenAI’s flagship model as of September 2026, holds a 1.05 million token context window, about 787,500 words. The cheaper GPT-6 Sol and GPT-6 Luna tiers share that same window at lower prices per token, so the word limit does not change, only the cost.
How many words can Claude remember at once?
Claude Opus 5.5, Claude Sonnet 5, and Claude Fable 5.1 each use a 1 million token context window, about 750,000 words. Claude Haiku 4.5, Anthropic’s fast and cheap tier, is smaller at 200,000 tokens, about 150,000 words.
How many words can Gemini remember at once?
Gemini 3.1 Pro and Gemini 3.8 Flash both accept about 1 million tokens of input, roughly 750,000 words. Output is capped much lower on both, around 64,000 to 66,000 tokens per response, about 48,000 to 49,500 words.
Which AI has the biggest context window in 2026?
OpenAI’s GPT-6 family (Astra, Sol, and Luna) has the largest window among mainstream models at 1.05 million tokens, about 787,500 words. Claude’s three main models and Gemini’s input side tie at 1 million tokens, about 750,000 words.
What is a token, and why is it not the same as a word?
A token is the chunk of text a model actually processes, often a word piece rather than a whole word. Across English prose, 1 token averages about 0.75 words, so 1,000 tokens is roughly 750 words. Short common words are usually 1 token; longer or unusual words can split into 2 or more.
How many book pages is a 1 million token context window?
About 2,727 pages, using the publishing-industry norm of roughly 275 words per typical book page. 750,000 words divided by 275 works out to just under 2,730 pages, or roughly nine 80,000-word novels stacked together.
Does a bigger context window cost more money?
Not necessarily. GPT-6 Sol shares GPT-6 Astra’s full 1.05 million token window but charges 80% less per token. Price depends on which tier of a model family you call, not on the window size alone, so check the per-model rate rather than assuming bigger always costs more.
Which AI model should you use for a long document?
For a single very long file, GPT-6 Astra or Claude Fable 5.1 give the most headroom at 750,000 to 787,500 words. For routine long documents where cost matters, Claude Sonnet 5, GPT-6 Sol, or Gemini 3.8 Flash cover almost any real-world file at a fraction of the price.
Sources: Claude Platform Docs, platform.claude.com/docs/en/about-claude/models/overview (accessed Sep 2026). OpenAI, "GPT-6 Astra: A new generation of intelligence" and "Introducing GPT-6 Sol and Luna," openai.com (Sep 3, 2026). Google DeepMind, Gemini 3.1 Pro model card, deepmind.google/models/model-cards/gemini-3-1-pro (2026). Google AI for Developers, Gemini API pricing, ai.google.dev/gemini-api/docs/pricing (Sep 2026). Word conversions by WordCounterTool at 0.75 words per token, rounded to the nearest 500. Prices and context windows change often; figures reflect what each provider published as of September 2026.