Loading...
Compare pricing, context windows, capabilities, and latency across major LLM providers. Find the right model for your use case.
Last updated: August 2026. Prices and specifications from official provider documentation.
| Model | Provider ⇅ | Context ⇅ | Input Price ⇅ | Output Price ⇅ | Latency ⇅ | Capabilities | Best For |
|---|---|---|---|---|---|---|---|
Claude 3 Opus 2024-03 | Anthropic | 200K max out: 4K | $15.00 per 1M tokens | $75.00 per 1M tokens | 🐢 Slow ~1200ms TTFT | VisionReasoning | Complex analysisNuanced writing+1 |
Claude 3.5 Haiku 2024-11 | Anthropic | 200K max out: 8K | $0.80 per 1M tokens | $4.00 per 1M tokens | ⚡ Very Fast ~140ms TTFT | VisionReasoning | Fast lightweight tasksCode completion+2 |
Claude 3.5 Sonnet 2024-10 | Anthropic | 200K max out: 8K | $3.00 per 1M tokens | $15.00 per 1M tokens | 🚀 Fast ~350ms TTFT | VisionReasoning | CodingTechnical writing+2 |
Claude 3.7 Sonnet 2025-02 | Anthropic | 200K max out: 64K | $3.00 per 1M tokens | $15.00 per 1M tokens | 🚀 Fast ~320ms TTFT | VisionReasoning | Hybrid reasoningComplex coding+2 |
Codestral 2024-05 | Mistral | 256K max out: 8K | $0.30 per 1M tokens | $0.90 per 1M tokens | 🚀 Fast ~200ms TTFT | VisionReasoning | Code completion & generationFill-in-the-Middle (FIM)+1 |
DeepSeek R1 2025-01 | DeepSeek | 64K max out: 8K | $0.55 per 1M tokens | $2.19 per 1M tokens | ⏱️ Medium ~1200ms TTFT | VisionReasoning | Open reasoningMath & logic synthesis+1 |
DeepSeek V3 2024-12 | DeepSeek | 64K max out: 8K | $0.14 per 1M tokens | $0.28 per 1M tokens | 🚀 Fast ~250ms TTFT | VisionReasoning | General codingUltra cost-effective throughput+1 |
Gemini 1.5 Flash 2024-05 | 1.0M max out: 8K | $0.07 per 1M tokens | $0.30 per 1M tokens | ⚡ Very Fast ~130ms TTFT | VisionReasoning | High-throughput summarizationLow-cost indexing | |
Gemini 1.5 Pro 2024-05 | 2.1M max out: 8K | $1.25 per 1M tokens | $5.00 per 1M tokens | ⏱️ Medium ~600ms TTFT | VisionReasoning | Large audio/video analysisLong documents+1 | |
Gemini 2.0 Flash 2025-02 | 1.0M max out: 8K | $0.10 per 1M tokens | $0.40 per 1M tokens | ⚡ Very Fast ~110ms TTFT | VisionReasoning | Real-time multimodal1M token context+1 | |
Gemini 2.0 Flash Thinking 2024-12 | 1.0M max out: 66K | $0.10 per 1M tokens | $0.40 per 1M tokens | ⏱️ Medium ~900ms TTFT | VisionReasoning | Fast visual & math reasoningLong context thinking | |
Gemini 2.0 Pro Experimental 2025-02 | 2.1M max out: 8K | $1.25 per 1M tokens | $5.00 per 1M tokens | 🚀 Fast ~400ms TTFT | VisionReasoning | 2M context codingWorld knowledge+1 | |
GPT-4 Turbo 2023-11 | OpenAI | 128K max out: 4K | $10.00 per 1M tokens | $30.00 per 1M tokens | ⏱️ Medium ~600ms TTFT | VisionReasoning | Legacy workloadsJSON mode |
GPT-4o 2024-05 | OpenAI | 128K max out: 16K | $2.50 per 1M tokens | $10.00 per 1M tokens | 🚀 Fast ~250ms TTFT | VisionReasoning | Multimodal vision & audioProduction APIs+1 |
GPT-4o Mini 2024-07 | OpenAI | 128K max out: 16K | $0.15 per 1M tokens | $0.60 per 1M tokens | ⚡ Very Fast ~120ms TTFT | VisionReasoning | Cost-effective tasksHigh-volume classification+1 |
Llama 3.1 405B 2024-07 | Meta | 128K max out: 4K | $3.00 per 1M tokens | $3.00 per 1M tokens | ⏱️ Medium ~700ms TTFT | VisionReasoning | Synthetic data generationTeacher model for fine-tuning+1 |
Llama 3.1 8B 2024-07 | Meta | 128K max out: 4K | $0.05 per 1M tokens | $0.05 per 1M tokens | ⚡ Very Fast ~80ms TTFT | VisionReasoning | Edge & local deploymentFine-tuning on private datasets+1 |
Llama 3.3 70B 2024-12 | Meta | 128K max out: 4K | $0.70 per 1M tokens | $0.80 per 1M tokens | 🚀 Fast ~220ms TTFT | VisionReasoning | Self-hosting on vLLMEnterprise data sovereignty+1 |
Mistral Large 2 2024-07 | Mistral | 128K max out: 8K | $2.00 per 1M tokens | $6.00 per 1M tokens | 🚀 Fast ~300ms TTFT | VisionReasoning | Multilingual EU legal/financeStructured JSON outputs+1 |
Mistral Small 3 2025-01 | Mistral | 32K max out: 8K | $0.20 per 1M tokens | $0.60 per 1M tokens | ⚡ Very Fast ~120ms TTFT | VisionReasoning | Fast lightweight tasksCost-sensitive European workloads |
o1 2024-12 | OpenAI | 200K max out: 100K | $15.00 per 1M tokens | $60.00 per 1M tokens | 🐢 Slow ~4000ms TTFT | VisionReasoning | Deep reasoningScientific analysis+1 |
o1-mini 2024-09 | OpenAI | 128K max out: 66K | $1.10 per 1M tokens | $4.40 per 1M tokens | ⏱️ Medium ~1500ms TTFT | VisionReasoning | STEM tasksCoding+1 |
o3-mini 2025-01 | OpenAI | 200K max out: 100K | $1.10 per 1M tokens | $4.40 per 1M tokens | ⏱️ Medium ~1200ms TTFT | VisionReasoning | Math & STEM reasoningCompetitive coding+1 |
Qwen 2.5 72B 2024-09 | Qwen | 128K max out: 8K | $0.35 per 1M tokens | $0.40 per 1M tokens | 🚀 Fast ~250ms TTFT | VisionReasoning | Multilingual accuracySTEM & mathematics+1 |
Qwen 2.5 7B 2024-09 | Qwen | 128K max out: 8K | $0.05 per 1M tokens | $0.05 per 1M tokens | ⚡ Very Fast ~70ms TTFT | VisionReasoning | Task-specific distillationEdge inference+1 |
Qwen 2.5 Coder 32B 2024-11 | Qwen | 128K max out: 8K | $0.20 per 1M tokens | $0.20 per 1M tokens | 🚀 Fast ~180ms TTFT | VisionReasoning | Self-hosted code assistantAgentic tools+1 |
GPT-4o Mini, Gemini 2.0 Flash, or DeepSeek V3 offer excellent price-to-performance ratios for high-volume tasks.
Claude 3.7 Sonnet, o3-mini, o1, or DeepSeek R1 excel at multi-step reasoning, mathematics, and complex analysis tasks.
Gemini 2.0 Pro / 1.5 Pro (2M context) and Claude 3.7 Sonnet (200K) handle extensive documents and codebases effectively.
Claude 3.7 Sonnet, GPT-4o, and Qwen 2.5 Coder provide state-of-the-art function calling and tool use for autonomous agent applications.