Context window sizes by model
Every number is taken from the manufacturer's documentation as of 7 August 2026, and manufacturers change these values.
Anthropic models, as documented on 7 August 2026
| Model |
Context window |
Max output |
Source |
| Claude Opus 5 |
1M tokens |
128k tokens |
Anthropic |
| Claude Sonnet 5 |
1M tokens |
128k tokens |
Anthropic |
| Claude Opus 4.6 |
1M tokens |
128k tokens |
Anthropic |
| Claude Sonnet 4.6 |
1M tokens |
128k tokens |
Anthropic |
| Claude Opus 4.5 |
200k tokens |
64k tokens |
Anthropic |
| Claude Sonnet 4.5 |
200k tokens |
64k tokens |
Anthropic |
| Claude Haiku 4.5 |
200k tokens |
64k tokens |
Anthropic |
OpenAI models, as documented on 7 August 2026
| Model |
Context window |
Max output |
Source |
| GPT-5.6 Sol |
1,050,000 |
128,000 |
OpenAI |
| GPT-5.6 Terra |
1,050,000 |
128,000 |
OpenAI |
| GPT-5 |
400,000 |
128,000 |
OpenAI |
| GPT-4o |
128,000 |
16,384 |
OpenAI |
| GPT-4o mini |
128,000 |
16,384 |
OpenAI |
| o3 |
200,000 |
100,000 |
OpenAI |
| o3-pro |
200,000 |
100,000 |
OpenAI |
| o4-mini |
200,000 |
100,000 |
OpenAI |
Google and xAI models, as documented on 7 August 2026
| Model |
Context window |
Max output |
Source |
| Gemini 3.1 Pro Preview |
1,048,576 |
65,536 |
Google |
| Gemini 3.6 Flash |
1,048,576 |
65,536 |
Google |
| Gemini 2.5 Pro |
1,048,576 |
65,536 |
Google |
| Gemini 2.5 Flash |
1,048,576 |
65,536 |
Google |
| grok-4.5 |
500,000 |
not documented |
xAI |
| grok-4.3 |
1M |
not documented |
xAI |
Footnotes that change the picture
- For GPT-5.6 Sol and Terra, requests over 272,000 input tokens cost double the input price and 1.5 times the output price according to OpenAI's documentation, applying to the entire request. The window is available, but priced differently above this limit.
- For GPT-5.6 Terra, the documented context window is 1,050,000, of which the maximum input length is 922,000 tokens.
- Anthropic documents 128k output for the synchronous Messages API; up to 300k output tokens are possible via the Message Batches API using a beta header.
- Anthropic points out that the tokenizer introduced with Opus 4.7 generates around 30 percent more tokens for the same text than previous generations. Window sizes of different generations are therefore not one-to-one comparable.
gemini-3.1-pro-preview is documented as a preview. For Grok, the documentation only specifies the context window, no maximum output length.
A bigger window does not mean a free window
The window size is the upper limit and not the consumption. Everything in the window is transmitted again in every subsequent request of the same session. Information on the measured costs of a single paste can be found in the article what a paste costs.