Model rules
Context windows, output caps, prices per million tokens and what each model tends to refuse. Prices move and limits change, so every row carries the date someone last checked it. If a number is wrong, say so and it gets fixed.
| Model | Context | Max output | In / out per 1M | Modalities | Good at | Tends to refuse | Checked |
|---|---|---|---|---|---|---|---|
| GPT-5 OpenAI | 400K | 128K | $1.25 / $10 | text, image, audio | Tool calling, structured output, agentic loops | Weapons, malware, explicit sexual content, medical and legal advice framed as… | May 2026 |
| GPT-5 mini OpenAI | 400K | 128K | $0.25 / $2 | text, image | Cheap structured extraction at volume | Same categories, stricter on borderline security questions | May 2026 |
| Claude Opus Anthropic | 200K | 64K | $15 / $75 | text, image | Long code review, editing, following long instruction lists | Weapons and malware detail, sexual content involving minors, impersonating a named living… | May 2026 |
| Claude Sonnet Anthropic | 200K | 64K | $3 / $15 | text, image | The default for daily work: fast, cheap enough, follows constraints | Same categories as Opus, with slightly more caution on security tooling | May 2026 |
| Claude Haiku Anthropic | 200K | 32K | $0.8 / $4 | text, image | Classification, extraction, high volume | Same policy categories, and it declines ambiguous requests more readily | May 2026 |
| Gemini Pro Google | 1000K | 65K | $1.25 / $10 | text, image, audio, video | Very long documents, video and PDF sets | Weapons, self-harm instruction, election-specific claims in some regions | May 2026 |
| Gemini Flash Google | 1000K | 65K | $0.3 / $2.5 | text, image, audio, video | Long input at low cost | Same categories as Pro | May 2026 |
| Grok 4 xAI | 256K | 64K | $3 / $15 | text, image | Recent events, informal drafting | Fewer refusals than the others by design, which cuts both ways | May 2026 |
| Llama 3.3 70B Meta | 128K | 8K | Local / Local | text | Private data, offline work, fine-tuning | Base weights refuse very little; behaviour depends on the instruct tuning you run | May 2026 |
| Llama 3.2 8B Meta | 128K | 8K | Local / Local | text | Running on a laptop, classification, drafting | Depends on the tuning | May 2026 |
| Mistral Large Mistral AI | 128K | 32K | $2 / $6 | text | Fast extraction and classification | Standard safety categories, comparatively permissive on security research | May 2026 |
| Mistral Small Mistral AI | 128K | 32K | $0.2 / $0.6 | text | High volume classification | Same categories | May 2026 |
Prices are list prices for the standard API tier in US dollars, before batch or cache discounts. Local models show no price because the cost is your hardware.
Something out of date here?
A wrong context window sends people down the wrong path. Corrections take a minute and go live after review.