AI Model Pricing Guide

Because tokens cost money and you're not made of it | Updated July 27, 2026

✨ Claude Opus 5 launched (Jul 24) — tops Artificial Analysis Intelligence Index at 61 pts (ahead of Fable 5 & GPT-5.6 Sol). Near-Fable-5 intelligence at half the price: $5/$25 per 1M. ARC-AGI-3 30.2% (4x GPT-5.6 Sol). New default on Claude Max.
Gemini 3.5 Pro launched (Jul 17) — 2M-token context window (largest at the frontier), Deep Think reasoning mode. $1.25/$10 per 1M — 4x cheaper input than GPT-5.6 Sol. Rebuilt from scratch after original run failed.
Kimi K3 (Moonshot AI, Jul 16) — 2.8T-param open-weight model, $3/$15 per 1M, 1M context, native multimodal. World's largest open-weight model. Full weights public Jul 27. Shook US chip stocks on release.
Qwen 3.8 Max (Alibaba, Jul 19) — 2.4T-param MoE, claims 'second only to Fable 5'. Multimodal. Preview via Token Plan subscription; open weights promised 'soon'.
GPT-5.6 now GA (Jul 9) — Sol $5/$30, Terra $2.50/$15, Luna $1/$6. Beats Fable 5 on Agents' Last Exam by 13 pts at 1/4 the cost. Government-coordinated rollout.
Grok 4.5 launched (Jul 8) — $2/$6 per 1M, Opus-class intelligence at 80 TPS. 4.2x more token-efficient than Opus 4.8. Available in Grok Build, Cursor, and SpaceXAI console.
Meta's first paid API: Muse Spark 1.1 ($1.25/$4.25) — ~25% cheaper than OpenAI/Anthropic. Agentic model from Meta Superintelligence Labs. US preview only, $20 free credits.
CHEAPEST: Qwen 3.6 Plus ($0.10/$0.30) / Llama 4 Scout ($0.15/$0.55) / Gemini 2.5 Flash-Lite ($0.10/$0.40)BEST VALUE: Claude Opus 5 ($5/$25, near-Fable-5 at half price) / Gemini 3.5 Pro ($1.25/$10, 2M ctx) / Grok 4.5 ($2/$6)SMARTEST: Claude Opus 5 (#1 AA Index 61pts) / GPT-5.6 Sol ($5/$30) / Claude Fable 5 ($10/$50) — Opus 5 matches Fable 5 on most benchmarks at half the costNEW THIS WEEK: Claude Opus 5 (Jul 24) • Gemini 3.5 Pro (Jul 17) • Kimi K3 (Jul 16) • Qwen 3.8 Max (Jul 19) • GPT-5.6 GA (Jul 9) • Grok 4.5 (Jul 8)
O

OpenAI

The OG of AI APIs. GPT kicked off the revolution and they're still leading.

BUDGET
GPT-5 mini
Fast & cheap
In
$0.25
per 1M
Out
$2.00
per 1M
128K ctxFast
Lightweight champion. Surprisingly capable for simple tasks and high-volume apps.
Best: Chatbots, simple QA, data extraction
BUDGET
o4-mini
Reinforcement tuned
In
$1.10
per 1M
Out
$4.40
per 1M
200K ctxFine-tuning
Price dropped 70%. Optimized for reinforcement fine-tuning workflows. Create custom reasoning patterns.
Best: Fine-tuning, custom reasoning
NEW
GPT-5.4 mini
Coding & subagents
In
$0.75
per 1M
Out
$4.50
per 1M
Cached inputCoding
New GPT-5.4-class model. Stronger than GPT-5 mini for coding and subagent workflows.
Best: Coding, subagents, mid-tier apps
NEW
GPT-5.4 nano
Cheapest 5.4-class
In
$0.20
per 1M
Out
$1.25
per 1M
Cached inputBudget
Cheapest way into the GPT-5.4 family. Cheaper input than GPT-5 mini.
Best: High-volume, budget apps
POWER
GPT-5.2
Reasoning beast
In
$1.75
per 1M
Out
$14.00
per 1M
6.6h horizon200K ctx
Top 3 on METR. Excels at complex tasks, code, and multi-step reasoning.
Best: Code, analysis, agent workflows
POWER
GPT-5.2 Pro
Reasoning premium
In
$21.00
per 1M
Out
$168.00
per 1M
200K ctxPremium
OpenAI's most precise reasoning model. For when you need the absolute best reasoning.
Best: Hardest problems, precision work
NEW
GPT-5.6 Terra
Balanced 5.6 — 2x cheaper than Sol
In
$2.50
per 1M
Out
$15.00
per 1M
GA Jul 9Balanced
Competitive with GPT-5.5 at 2x lower cost. Outperforms Fable 5 at ~1/16 the cost. The everyday work model of the 5.6 family.
Best: General purpose, production apps
NEW
GPT-5.6 Luna
Fastest 5.6 — lowest cost
In
$1.00
per 1M
Out
$6.00
per 1M
GA Jul 9Fast
Outperforms Opus 4.8 on coding agent index. Nearly matches GPT-5.5 peak at less than half the cost. The high-volume tier of the 5.6 family.
Best: High-volume, cost-sensitive apps
POWER
GPT-5.5
Now #2 — still elite
In
$5.00
per 1M
Out
$30.00
per 1M
1.05M ctxReasoningAgents
Previous #1. 82.7% Terminal-Bench, 84.9% GDPval, 78.7% OSWorld. 1.05M context window. Still elite, now behind GPT-5.6.
Best: Coding, agents, research, multi-step tasks
FLAGSHIP
GPT-5.4 Pro
Previous premium — now #3
In
$30.00
per 1M
Out
$180.00
per 1M
270K ctxPremium
Former #1, now behind GPT-5.5. Still incredibly powerful for demanding tasks.
Best: Most demanding tasks, unlimited budget
A

Anthropic

Safety-first company. Claude is beloved by developers for being genuinely helpful.

FAST
Claude Haiku 4.5
Speed demon
In
$1.00
per 1M
Out
$5.00
per 1M
200K ctxFastest
Optimized for fast responses. Perfect for real-time apps and bulk processing.
Best: Real-time chat, bulk processing
BEST
Claude Sonnet 4.6
Previous Sonnet — still solid
In
$3.00
per 1M
Out
$15.00
per 1M
Balanced200K ctx
Previous Sonnet default. Still excellent but superseded by Sonnet 5 at lower intro pricing.
Best: Most tasks, code, writing, general use
POWER
Claude Opus 4.8
Previous Opus flagship — superseded by Opus 5
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctx128K outputSelf-verify
Previous Opus top tier, now superseded by Claude Opus 5 at the same price. 1M context, 128K output, autonomous self-verification. Same $5/$25 pricing as 4.7. Migrate to Opus 5 for near-Fable-5 intelligence at no extra cost.
Best: Complex coding, agents, long-horizon tasks
POWER
Claude Opus 4.7
Previous Opus SOTA — still elite
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctxxhigh reasoningSelf-verify
Previous Anthropic best. 1M context, autonomous self-verification. Beat GPT-5.4 on BrowseComp. Now superseded by Opus 4.8 at the same price.
Best: Complex coding, agents, long-horizon tasks
POWER
Claude Opus 4.6
Proven workhorse
In
$5.00
per 1M
Out
$25.00
per 1M
14.5h horizon200K ctxFast mode
Still one of the best. 14+ hour autonomous tasks. Reliable, consistent, now the value play vs 4.7/4.8.
Best: Hard problems, research, complex agents
LIMITED
Claude Mythos 5
Fable 5 without safety classifiers — limited
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputGlasswing only
Same capabilities as Fable 5 without the safety classifiers. Limited availability through Project Glasswing (cybersecurity initiative). Restored Jul 1 after US government suspension. Successor to Claude Mythos Preview.
Best: Cybersecurity, frontier research, if you can get access
X

xAI

Elon's AI. Grok has real-time X data access and absurdly low pricing.

NEW
Grok 4.20
Same price as 4.3, more features
In
$1.25
per 1M
Out
$2.50
per 1M
2M ctxReasoningMulti-agentVision
Same pricing as Grok 4.3 with multi-agent orchestration. Cached input at $0.125/1M. 2M context window for complex agent swarms.
Best: Complex multi-agent workflows
RETIRED
Grok 4 / 4.1 Fast
Retired May 15
In
$0.20
per 1M
Out
$0.50
per 1M
2M ctxRetired
Retired May 15, 2026. Migrated? Good. If not, move to Grok 4.3 ($1.25/$2.50) or Grok 4.20 ($1.25/$2.50).
Best: → Migrate to: Grok 4.3
RETIRED
Grok Code Fast 1
Retired May 15
In
$0.20
per 1M
Out
$1.50
per 1M
256K ctxRetired
Retired May 15, 2026. Migrate to Grok 4.3 for coding.
Best: → Migrate to: Grok 4.3
BUDGET
Grok 3 Mini
Older gen cheap
In
$0.30
per 1M
Out
$0.50
per 1M
131K ctxReasoning
Budget fallback if Grok 4's 2M context is overkill for your use case.
Best: Simple tasks, testing
POWER
Grok 4-0709
Premium tier
In
$3.00
per 1M
Out
$15.00
per 1M
256K ctxReasoningVision
Premium Grok. Smaller context but more reasoning power.
Best: Grok style with more smarts
M

Meta

Meta's first paid API. Muse Spark brings aggressive pricing and agentic capabilities from Meta Superintelligence Labs.

NEW
Muse Spark 1.1
Meta's first paid API — 25% cheaper
In
$1.25
per 1M
Out
$4.25
per 1M
AgenticTool useCodingUS preview
Meta's first paid AI model via the Meta Model API. Agentic model from Meta Superintelligence Labs (run by Alexandr Wang). ~25% cheaper than comparable OpenAI/Anthropic models. $20 free credits for new accounts. US preview only — no EU access yet.
Best: Agentic tasks, tool use, cost-sensitive apps
G

Google DeepMind

Gemini has quietly become excellent. Massive context, strong multimodal, and a generous free tier.

NEW
Gemini 3.1 Flash-Lite
Cheapest Gemini 3
In
$0.25
per 1M
Out
$1.50
per 1M
PreviewBudget
Cheapest way into Gemini 3.1. Preview tier with budget-friendly pricing.
Best: Budget Gemini 3 apps, prototyping
NEW
Gemini 3 Flash
New budget
In
$0.50
per 1M
Out
$3.00
per 1M
PreviewFast
Gemini 3 Flash preview. Balanced performance at budget pricing.
Best: Budget apps, prototyping
VALUE
Gemini 2.5 Flash
Best value
In
$0.30
per 1M
Out
$2.50
per 1M
1M ctxMultimodalFree tier
Cheapest way to process 1M context. Free tier available. Multimodal - images, video, audio.
Best: High-volume, multimodal, prototypes
NEW
Gemini 2.5 Flash-Lite
Ultra-cheap Flash
In
$0.10
per 1M
Out
$0.40
per 1M
1M ctxBudget
Flash-Lite tier for Gemini 2.5. Cheaper than standard Flash with 1M context support. Best for high-volume simple tasks.
Best: High-volume, simple tasks, cost-sensitive apps
NEW
Gemini 3.5 Flash
GA Flash — 1M context, agentic
In
$1.50
per 1M
Out
$9.00
per 1M
1M ctx65K outputAgenticComputer Use
Gemini 3.5 Flash is GA. Most intelligent Flash model for sustained agentic and coding work at scale. 1M context, 65K output, thinking, Computer Use, function calling. Free tier available.
Best: Agentic tasks, coding, production Flash workloads
Gemini 3 Pro
3rd gen flagship
In
$2.00
per 1M
Out
$12.00
per 1M
1M ctx
Third-generation Gemini Pro. Now stable — no Preview tag. Same / pricing as preview tier. Strong general-purpose flagship.
Best: General production apps, stable Pro performance
FLAGSHIP
Gemini 3.1 Pro Preview
New flagship
In
$2.00
per 1M
Out
$12.00
per 1M
4h horizonPreviewVideo
77.1% ARC-AGI-2. Price increased from $1.25/$10. Batch and Flex tiers at 50% off.
Best: Video analysis, complex reasoning
LONG
Gemini 2.5 Pro
Long outputs
In
$1.25
per 1M
Out
$10.00
per 1M
1M ctx64K output
Same price as 3.1 Pro but 64K max output vs 16K. Choose for long-form content generation.
Best: Long-form writing, large outputs
DEPRECATED
Gemini 2.0 Flash
Shuts down Jun 1
In
$0.15
per 1M
Out
$0.60
per 1M
1M ctx8K outputRetiring Jun 1
Deprecated — shuts down June 1, 2026. Migrate to Gemini 2.5 Flash or 3.1 Flash-Lite.
Best: Migrate away from this model

Open Source & Local

Open-weight models you can run yourself or call via cheap APIs. The frontier is no longer closed.

NEW
Kimi K2.6
88% cheaper than Opus
In
$0.60
per 1M
Out
$2.50
per 1M
256K ctxOpen weightMoE 1T/32B
Beats GPT-5.4 and Opus 4.6 on SWE-Bench Pro. 1T params, 32B active. 300 sub-agent orchestration. OpenAI-compatible API.
Best: Coding, agents, long-horizon tasks
CHEAPEST
Qwen 3.6 Plus
1M context, free tier
In
$0.10
per 1M
Out
$0.30
per 1M
1M ctxReasoningFree tier
Alibaba's latest. Mandatory chain-of-thought reasoning. Free tier available. Topped 6 coding benchmarks on release.
Best: Budget coding, massive context
NEW
Llama 4 Scout
10M context MoE
In
$0.15
per 1M
Out
$0.55
per 1M
10M ctxOpen weightMoE 109B
Longest context of any open model. 109B total, 17B active. Multimodal. Runs on 24GB VRAM.
Best: Massive context, multimodal, local
Llama 4 Maverick
Frontier coding MoE
In
$0.20
per 1M
Out
$0.80
per 1M
1M ctxOpen weightMoE 400B
Beats GPT-4o on coding. 400B total, 17B active. 128 experts. Frontier quality at MoE prices.
Best: Coding, complex reasoning
DeepSeek V3.2
Matches GPT-4o
In
$0.27
per 1M
Out
$1.10
per 1M
128K ctxOpen weightMoE 685B
94.2% MMLU matching GPT-4o. 685B MoE with 37B active. Best open model for general knowledge.
Best: General knowledge, research
FREE
Qwen3-Coder 8B
Local coding king
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
32K ctxLocal only8B dense
Runs on any 8GB GPU. 92 programming languages. 80-150 tok/s. Best local coding model under 10B. Set it up locally →
Best: Local coding, autocomplete
FREE
DeepSeek R1 Distill 14B
Local reasoning
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
Local onlyReasoning10GB VRAM
Chain-of-thought reasoning on 10GB VRAM. The sweet spot for local reasoning. 55 tok/s on modern GPUs. Run it offline →
Best: Local reasoning, budget hardware
💡 Did You Know?
Claude Opus 5 — near-Fable-5 at half the price
Anthropic launched Claude Opus 5 on Jul 24, 2026 — $5/$25 per 1M (same as Opus 4.8, half of Fable 5's $10/$50). It tops the Artificial Analysis Intelligence Index at 61 points, ahead of Fable 5 (60) and GPT-5.6 Sol (59). ARC-AGI-3: 30.2% — 4x GPT-5.6 Sol, 20x Opus 4.8. New default on Claude Max. Fast mode runs 2.5x faster at $10/$50 (API research preview). Batch API at $2.50/$12.50. Opus 5 is Anthropic's fourth model launch in under two months.
Gemini 3.5 Pro ships — 2M context
Google DeepMind launched Gemini 3.5 Pro on Jul 17 with a 2-million-token context window — double anything at the frontier. New Deep Think reasoning mode (gated to $250/mo Ultra). Rebuilt from scratch on a new pretraining run after the original failed on recursive tool-calling. $1.25/$10 — 4x cheaper input than GPT-5.6 Sol.
Kimi K3 — world's largest open weight
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model — the largest ever released. $3/$15 per 1M, 1M context, native multimodal. A deliberate shift away from ultra-low-price strategy toward enterprise performance competition. Full weights public Jul 27. Shook US chip stocks on release.
GPT-5.6 goes GA
GPT-5.6 launched July 9 in three tiers — Sol ($5/$30), Terra ($2.50/$15), Luna ($1/$6). Sol beats Fable 5 on Agents' Last Exam by 13 points at ~1/4 the cost. Luna outperforms Opus 4.8 at ~1/16 the cost. First frontier launch with US government coordination.
Grok 4.5's token efficiency
Grok 4.5 solves SWE Bench Pro tasks with 15,954 output tokens on average — 4.2x fewer than Opus 4.8 (67,020). At $2/$6 and 80 TPS, it delivers Opus-class results at a fraction of the cost.
Meta joins the API game
Muse Spark 1.1 is Meta's first ever paid API model — $1.25/$4.25, ~25% cheaper than OpenAI and Anthropic. Meta Superintelligence Labs, run by former Scale AI chief Alexandr Wang, built it for agentic tasks and tool use. $20 free credits for new accounts.
The price war
July 2026 is the AI price war. Claude Opus 5 at $5/$25 with near-Fable-5 intelligence, Gemini 3.5 Pro at $1.25/$10 with 2M context, GPT-5.6 Luna at $1/$6, Muse Spark at $1.25/$4.25, Grok 4.5 at $2/$6 — frontier compute is commodifying. The top three on the AA Intelligence Index (Opus 5 61, Fable 5 60, GPT-5.6 Sol 59) are separated by 2 points. Kimi K3 breaks the cheap-China pattern at $3/$15, betting on enterprise performance.