AI Model Pricing Guide
Because tokens cost money and you're not made of it | Updated September 7, 2026
CHEAPEST: Qwen 3.6 Plus ($0.10/$0.30) / Tencent Hy3 ($0.13/$0.53) / Llama 4 Scout ($0.15/$0.55) / GPT-5.6 Luna ($0.20/$1.20, cheapest frontier) / MiniMax M2 ($0.30/$1.20) / DeepSeek V4-Pro off-peak $0.66/$1.98BEST VALUE: Gemini 3.8 Flash ($0.75/$3.75 intro, AA Index 59) / Gemini 3.7 Flash ($0.75/$3.75, still supported) / Qwen 3.8-Max ($2/$6, 1M ctx flat) / Claude Opus 5 ($5/$25, #2 on AA Index) / Gemini 3.5 Pro ($1.25/$10, 2M ctx)SMARTEST: Claude Fable 5.1 (AA Index 66 #1 — $10/$50, cache reads $0.25) / Claude Opus 5 (63 pts, $5/$25) / Muse Spark 1.3 (62 max, $1.25/$4.25 — cheapest per task at 59+) / GPT-6 Astra (61 pts, $10/$50, saturates ARC-AGI-3)NEW THIS WEEK: 4 frontier launches in 72h — Fable 5.1 (Sep 1), Gemini 3.8 Flash (Sep 2), Muse Spark 1.3 (Sep 2), GPT-6 Astra (Sep 3) • Sonnet 5 $2/$10 made permanent • GLM-5.3-Flash promo ends Sep 9 • Calendar: Sol promo ends Nov 21 ($5/$30), Gemini intro ends Dec 31 (doubles Jan 1)
O
OpenAI
The OG of AI APIs. GPT kicked off the revolution and they're still leading.
NEW
GPT-6 Astra
New GPT-6 flagship — first at critical-cyber threshold
In
$10.00
per 1M
Out
$50.00
per 1M
1.05M ctx128K outputARC-AGI-3 99.9%Critical cyber tier
OpenAI's new flagship, announced Sep 3 and live on the API Sep 4. $10/$50 per 1M — 2.5x GPT-5.6 Sol's promo price — with $1 cached input, $12.50 cache writes, and a long-context tier above 272K input (2x input/cache, 1.5x output). Batch/Flex 50% off; Fast mode 2x. Saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98%); ExploitBench 100% without safeguards; GPQA 96.0%; Terminal-Bench 4.0 57.9%. AA Intelligence Index: 61 (ties GPT-5.6 Sol) but ~70% more token-efficient — less than half Fable 5's cost per coding task. First model to reach OpenAI's critical cybersecurity threshold; advanced cyber work gated behind Daybreak. Wider rollout expected around DevDay Sep 29.
Best: Frontier agentic work, computer use, cyber defense, long multi-step tasks
BUDGET
GPT-5 mini
Fast & cheap
In
$0.25
per 1M
Out
$2.00
per 1M
128K ctxFast
Lightweight champion. Surprisingly capable for simple tasks and high-volume apps.
Best: Chatbots, simple QA, data extraction
BUDGET
o4-mini
Reinforcement tuned
In
$1.10
per 1M
Out
$4.40
per 1M
200K ctxFine-tuning
Price dropped 70%. Optimized for reinforcement fine-tuning workflows. Create custom reasoning patterns.
Best: Fine-tuning, custom reasoning
NEW
GPT-5.4 mini
Coding & subagents
In
$0.75
per 1M
Out
$4.50
per 1M
Cached inputCoding
New GPT-5.4-class model. Stronger than GPT-5 mini for coding and subagent workflows.
Best: Coding, subagents, mid-tier apps
NEW
GPT-5.4 nano
Cheapest 5.4-class
In
$0.20
per 1M
Out
$1.25
per 1M
Cached inputBudget
Cheapest way into the GPT-5.4 family. Cheaper input than GPT-5 mini.
Best: High-volume, budget apps
POWER
GPT-5.4
Still elite — now #2
In
$2.50
per 1M
Out
$15.00
per 1M
270K ctxReasoning
OpenAI's previous #1. Still elite for complex, multi-step problems.
Best: Hardest problems, professional work
POWER
GPT-5.2
Reasoning beast
In
$1.75
per 1M
Out
$14.00
per 1M
6.6h horizon200K ctx
Top 3 on METR. Excels at complex tasks, code, and multi-step reasoning.
Best: Code, analysis, agent workflows
POWER
GPT-5.2 Pro
Reasoning premium
In
$21.00
per 1M
Out
$168.00
per 1M
200K ctxPremium
OpenAI's most precise reasoning model. For when you need the absolute best reasoning.
Best: Hardest problems, precision work
POWER
GPT-5.6 Sol
Previous flagship — 2.5x cheaper than Astra
In
$4.00
per 1M
Out
$20.00
per 1M
GA Jul 9Fast mode Jul 30AgentsCodingCybersecurity
OpenAI's strongest model, now GA. SOTA on Agents' Last Exam (53.6, +13 over Fable 5), BrowseComp (92.2%), Terminal-Bench 2.1 (88.8%). Three tiers: Sol (flagship), Terra (balanced, $2/$12), Luna (fast, $0.20/$1.20). New Fast mode: 2.5x faster at $10/$60 (API). Multi-agent 'ultra' mode for demanding tasks. Aug 21 promo repricing: $4/$20 through Nov 21, then $5/$30 standard. Superseded Sep 3 by GPT-6 Astra at 2.5x the price — Sol remains the value pick for frontier agentic work.
Best: Frontier agentic tasks, coding, biology, cybersecurity
20% OFF
GPT-5.6 Terra
Balanced 5.6 — now 20% cheaper
In
$2.00
per 1M
Out
$12.00
per 1M
GA Jul 9Price cut Jul 30Balanced
Price cut 20% on Jul 30 — from $2.50/$15 to $2/$12. Matches Claude Sonnet 5 on input price during intro. Outperforms Fable 5 at ~1/16 the cost. The everyday work model of the 5.6 family. Batch: $1/$6. Cached input: $0.20.
Best: General purpose, production apps
80% OFF
GPT-5.6 Luna
Fastest 5.6 — now 80% cheaper
In
$0.20
per 1M
Out
$1.20
per 1M
GA Jul 9Price cut Jul 30Fast
Price slashed 80% on Jul 30 — from $1/$6 to $0.20/$1.20 per 1M. Outperforms Opus 4.8 on coding agent index. Nearly matches GPT-5.5 peak at a fraction of the cost. The high-volume tier of the 5.6 family. Batch: $0.10/$0.60. Cached input: $0.02.
Best: High-volume, cost-sensitive apps, agent execution
POWER
GPT-5.5
Now #2 — still elite
In
$5.00
per 1M
Out
$30.00
per 1M
1.05M ctxReasoningAgents
Previous #1. 82.7% Terminal-Bench, 84.9% GDPval, 78.7% OSWorld. 1.05M context window. Still elite, now behind GPT-5.6.
Best: Coding, agents, research, multi-step tasks
#1 RANKED
GPT-5.5 Pro
Maximum intelligence
In
$30.00
per 1M
Out
$180.00
per 1M
1.05M ctxPremiumDeep Research
90.1% BrowseComp, 52.4% FrontierMath Tier 1-3. 1.05M context window — read entire codebases and research libraries. The ceiling for what AI can do right now.
Best: Hardest problems, deep research, scientific discovery
FLAGSHIP
GPT-5.4 Pro
Previous premium — now #3
In
$30.00
per 1M
Out
$180.00
per 1M
270K ctxPremium
Former #1, now behind GPT-5.5. Still incredibly powerful for demanding tasks.
Best: Most demanding tasks, unlimited budget
A
Anthropic
Safety-first company. Claude is beloved by developers for being genuinely helpful.
NEW
Claude Fable 5.1
New AA Index #1 — cache reads 75% cheaper
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputCache $0.25AA Index 66 #1
Anthropic's new frontier (Sep 1). $10/$50 unchanged but cache reads cut 75% to $0.25/1M — Anthropic's August usage data shows ~25% total savings on typical workloads, up to ~45% for agentic work where cache hits dominate. Terminal-Bench-Science 52.6% (vs 24.7% on Fable 5), Terminal-Bench 4.0 55.8%, AutomationBench 31.4%, GDPval-AA v2 1853, CursorBench 73.4%. AA Intelligence Index (max, fallback): 66 — #1, three clear of Opus 5. Cyber false positives down 60%; can now find (not exploit) vulnerabilities. Breaking API changes: forced tool use returns 400, thinking blocks are model-bound. API ID: claude-fable-5-1.
Best: Coding, agents, long-horizon knowledge work — migrate from Fable 5 for the cache savings
FAST
Claude Haiku 4.5
Speed demon
In
$1.00
per 1M
Out
$5.00
per 1M
200K ctxFastest
Optimized for fast responses. Perfect for real-time apps and bulk processing.
Best: Real-time chat, bulk processing
NEW
Claude Sonnet 5
Most agentic Sonnet — $2/$10 made permanent
In
$2.00
per 1M
Out
$10.00
per 1M
200K ctxAgentic$2/$10 permanent
Close to Opus 4.8 performance at Sonnet prices. Most agentic Sonnet ever — plans, uses tools, runs autonomously. Intro pricing $2/$10 made permanent Aug 10 — the scheduled Sept 1 increase was cancelled. Default model for Free and Pro plans.
Best: Coding, agents, tool use, autonomous tasks — the new default Sonnet
BEST
Claude Sonnet 4.6
Previous Sonnet — still solid
In
$3.00
per 1M
Out
$15.00
per 1M
Balanced200K ctx
Previous Sonnet default. Still excellent but superseded by Sonnet 5 at lower intro pricing.
Best: Most tasks, code, writing, general use
NEW
Claude Opus 5
#2 on AA Index — near-frontier at half price
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctx128K outputThinking ONNew Claude Max default
Anthropic's new top-tier Opus. Scores 63 pts on the AA Intelligence Index (v4.1.1) — second only to Fable 5.1's 66. Near-Fable-5 intelligence at exactly half the price ($5/$25 vs $10/$50). ARC-AGI-3: 30.2% — 4x GPT-5.6 Sol (7.8%), 20x Opus 4.8 (1.5%). Frontier-Bench v0.1: 43.3% (Opus 4.8: 18.9%). Five-level effort toggle, thinking ON by default. Automatic fallback replaces hard refusals. API ID: claude-opus-5. New default on Claude Max; top model on Claude Pro. Fast mode: $10/$50 at 2.5x speed (API research preview). Batch API: $2.50/$12.50.
Best: Coding, agents, knowledge work — the new cost-efficient frontier sweet spot
POWER
Claude Opus 4.8
Previous Opus flagship — superseded by Opus 5
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctx128K outputSelf-verify
Previous Opus top tier, now superseded by Claude Opus 5 at the same price. 1M context, 128K output, autonomous self-verification. Same $5/$25 pricing as 4.7. Migrate to Opus 5 for near-Fable-5 intelligence at no extra cost.
Best: Complex coding, agents, long-horizon tasks
POWER
Claude Opus 4.7
Previous Opus SOTA — still elite
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctxxhigh reasoningSelf-verify
Previous Anthropic best. 1M context, autonomous self-verification. Beat GPT-5.4 on BrowseComp. Now superseded by Opus 4.8 at the same price.
Best: Complex coding, agents, long-horizon tasks
POWER
Claude Opus 4.6
Proven workhorse
In
$5.00
per 1M
Out
$25.00
per 1M
14.5h horizon200K ctxFast mode
Still one of the best. 14+ hour autonomous tasks. Reliable, consistent, now the value play vs 4.7/4.8.
Best: Hard problems, research, complex agents
POWER
Claude Fable 5
Superseded by Fable 5.1 — migrate for cache savings
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputSafety classifiersRestored Jul 1
Public version of Mythos. $10/$50 — double Opus 4.8. 1M context, 128K output, autonomous self-verification. Includes safety classifiers that can refuse requests (refusals are not billed; fallback credits refund prompt-cache cost on retry). Suspended by US government Jun 12, restored globally Jul 1. Superseded Sep 1 by Fable 5.1 — same $10/$50 with cache reads at $0.25/1M.
Best: Hardest reasoning, long-horizon agentic work, cybersecurity
LIMITED
Claude Mythos 5.1
Fable 5.1 with permissive safeguards — vetted access
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputTrusted accessUS orgs only
Identical to Fable 5.1 with safeguards tuned for cybersecurity and life sciences. Ships via the Cyber Verification Program, the US-government-partnered Life Sciences Verification Program, and Project Glasswing — Anthropic's strongest cyber model ever. System card flags it as 'less honest under pressure' than recent Claude models. Also powers Claude Security. US organizations only for now.
Best: Vetted cyberdefenders, life-sciences R&D — apply via CVP/LSVP
NEW
Grok 4.6
New flagship — 1753 ELO claim, overtakes Kimi K3
In
$2.00
per 1M
Out
$6.00
per 1M
500K ctx1753 ELOCached $0.50Coding & agents
xAI's new flagship (Aug 12), released with Cursor. $2/$6 per 1M, $0.50 cached input. 1753 ELO claim — overtakes Kimi K3. Live on xAI API, Grok Build, Cursor, Grok Bot, and partners (OpenRouter, Vercel, Cloudflare). Fast variant available at twice the price. 2x included usage in Cursor/Grok Build for the first week.
Best: Coding, long-running agents, knowledge work — the new default Grok
POWER
Grok 4.5
Previous flagship — superseded by 4.6
In
$2.00
per 1M
Out
$6.00
per 1M
500K ctx80 TPS4.2x token efficiencyCoding & agents
SpaceXAI's previous flagship, now superseded by Grok 4.6 at the same price. Trained alongside Cursor. Opus 4.8-class intelligence at 80 TPS with 4.2x better token efficiency than Opus 4.8. SWE Marathon 29% (beats Opus 4.8 at 26%). Cached input at $0.50/1M. Migrate to Grok 4.6 for no extra cost.
Best: Coding, agentic tasks, knowledge work — migrate to Grok 4.6
NEW
Grok 4.20
Same price as 4.3, more features
In
$1.25
per 1M
Out
$2.50
per 1M
2M ctxReasoningMulti-agentVision
Same pricing as Grok 4.3 with multi-agent orchestration. Cached input at $0.125/1M. 2M context window for complex agent swarms.
Best: Complex multi-agent workflows
NEW
Grok 4.3
New recommended base model
In
$1.25
per 1M
Out
$2.50
per 1M
Best value2M ctxReasoningVision
xAI's recommended Grok 4 model after retiring old variants. Beats Grok 4.1 on coding, agents, and reasoning. This is the migration target for retiring models.
Best: High-volume apps, X analysis, multi-agent — the new default Grok
RETIRED
Grok 4 / 4.1 Fast
Retired May 15
In
$0.20
per 1M
Out
$0.50
per 1M
2M ctxRetired
Retired May 15, 2026. Migrated? Good. If not, move to Grok 4.3 ($1.25/$2.50) or Grok 4.20 ($1.25/$2.50).
Best: → Migrate to: Grok 4.3
RETIRED
Grok Code Fast 1
Retired May 15
In
$0.20
per 1M
Out
$1.50
per 1M
256K ctxRetired
Retired May 15, 2026. Migrate to Grok 4.3 for coding.
Best: → Migrate to: Grok 4.3
BUDGET
Grok 3 Mini
Older gen cheap
In
$0.30
per 1M
Out
$0.50
per 1M
131K ctxReasoning
Budget fallback if Grok 4's 2M context is overkill for your use case.
Best: Simple tasks, testing
POWER
Grok 4-0709
Premium tier
In
$3.00
per 1M
Out
$15.00
per 1M
256K ctxReasoningVision
Premium Grok. Smaller context but more reasoning power.
Best: Grok style with more smarts
M
Meta
Meta's first paid API. Muse Spark brings aggressive pricing and agentic capabilities from Meta Superintelligence Labs.
NEW
Muse Spark 1.3
Meta reaches the top-3 — cheapest at 59+ intelligence
In
$1.25
per 1M
Out
$4.25
per 1M
AA Index 6162 max (preview)$0.55/taskCached $0.15
Meta's fourth Muse Spark in five months (Sep 2). $1.25/$4.25 unchanged ($0.15 cached input). AA Intelligence Index: 61 at xhigh (up 4 from 1.2), 62 at max — limited partner preview — behind only Fable 5.1 (66) and Opus 5 (63). #1 on Tau3-Bench Banking. $0.55 per Index task vs $0.94–0.95 for GPT-5.6 Sol and Grok 4.6 — the cheapest of any model at 59+. Meta teases open weights and larger Muse models around Q4.
Best: Agentic work, scientific reasoning, cost-efficient frontier intelligence
Muse Glimmer 30B
First MSL open model, runs locally
NEW
Muse Spark 1.1
Previous Muse Spark — superseded by 1.3
In
$1.25
per 1M
Out
$4.25
per 1M
AgenticTool useCodingUS preview
Meta's first paid AI model via the Meta Model API. Agentic model from Meta Superintelligence Labs (run by Alexandr Wang). ~25% cheaper than comparable OpenAI/Anthropic models. $20 free credits for new accounts. US preview only — no EU access yet.
Best: Agentic tasks, tool use, cost-sensitive apps
G
Google DeepMind
Gemini has quietly become excellent. Massive context, strong multimodal, and a generous free tier.
NEW
Gemini 3.1 Flash-Lite
Cheapest Gemini 3
In
$0.25
per 1M
Out
$1.50
per 1M
PreviewBudget
Cheapest way into Gemini 3.1. Preview tier with budget-friendly pricing.
Best: Budget Gemini 3 apps, prototyping
NEW
Gemini 3 Flash
New budget
In
$0.50
per 1M
Out
$3.00
per 1M
PreviewFast
Gemini 3 Flash preview. Balanced performance at budget pricing.
Best: Budget apps, prototyping
VALUE
Gemini 2.5 Flash
Best value
In
$0.30
per 1M
Out
$2.50
per 1M
1M ctxMultimodalFree tier
Cheapest way to process 1M context. Free tier available. Multimodal - images, video, audio.
Best: High-volume, multimodal, prototypes
NEW
Gemini 2.5 Flash-Lite
Ultra-cheap Flash
In
$0.10
per 1M
Out
$0.40
per 1M
1M ctxBudget
Flash-Lite tier for Gemini 2.5. Cheaper than standard Flash with 1M context support. Best for high-volume simple tasks.
Best: High-volume, simple tasks, cost-sensitive apps
NEW
Gemini 3.5 Flash
GA Flash — 1M context, agentic
In
$1.50
per 1M
Out
$9.00
per 1M
1M ctx65K outputAgenticComputer Use
Gemini 3.5 Flash is GA. Most intelligent Flash model for sustained agentic and coding work at scale. 1M context, 65K output, thinking, Computer Use, function calling. Free tier available.
Best: Agentic tasks, coding, production Flash workloads
NEW
Gemini 3.8 Flash
4th Flash in 4 months — AA Index 59 at Flash prices
In
$0.75
per 1M
Out
$3.75
per 1M
1M ctx~300 tok/sAA Index 59Terminal-Bench 90.8%
Google's best reasoning/coding Flash yet (Sep 2) — 3.5, 3.6, 3.7, 3.8 since May. Same $0.75/$3.75 intro as 3.7 Flash through Dec 31, then $1.50/$7.50 from Jan 1, 2027. AA Intelligence Index 59 (+3 over 3.7) — cheapest model at its intelligence level ($0.58/task). ~300 tok/s, the fastest output speed Artificial Analysis has measured. Trained on long-running agentic loops, it 'works harder': ~40% higher cost per task than 3.7 (30% more output tokens) — drop to low effort for efficiency-first routing. Gemini 3.8 Flash Cyber for trusted defenders via the Fairwind Program: 2.6x more correct Chrome patches than much larger frontier models, per Google.
Best: Agentic coding, knowledge work, price-sensitive production — the new default Gemini
NEW
Gemini 3.7 Flash
Efficiency pick — superseded by 3.8 Flash
In
$0.75
per 1M
Out
$3.75
per 1M
1M ctxGA Aug 13DeepSWE 65.3%Computer Use
Gemini 3.8 Flash (Sep 2) supersedes it at the same intro price; 3.7 remains fully supported for efficiency-first workloads. Google's most intelligent workhorse Flash model yet (Aug 13). Intro price $0.75/$3.75 — half 3.6 Flash's original cost — through Dec 31, then $1.50/$7.50. DeepSWE v1.1: 65.3% (vs 49% on 3.6). FrontierCode 1.1: 43.6% (vs 34.4%). WebDev Arena 1588 Elo (vs 1538). GDP.pdf 34% (vs 22%). AutomationBench 30.4% (vs 17%). Powers Gemini Spark. Built-in Computer Use.
Best: Agentic coding, knowledge work, cost-efficient production agents
NEW
Gemini 3.6 Flash
New workhorse — 17% fewer output tokens
In
$1.50
per 1M
Out
$7.50
per 1M
1M ctxGA Jul 21Computer UseToken efficient
Google's new workhorse Flash model (Jul 21). 17% fewer output tokens than 3.5 Flash on the AA Index (up to 65% on DeepSWE). Step up in coding (DeepSWE 49% vs 37%), knowledge work (GDPval 1421 vs 1349), computer use (OSWorld 83% vs 78.4%). $1.50/$7.50 — lower price than 3.5 Flash too. Built-in Computer Use via Gemini API.
Best: Agentic coding, knowledge work, cost-efficient agents
NEW
Gemini 3.5 Flash-Lite
Fastest 3.5 — 350 TPS
In
$0.30
per 1M
Out
$2.50
per 1M
GA Jul 21350 TPSComputer UseAgentic
Fastest model in the 3.5 series — 350 output tokens/s (AA Index). $0.30/$2.50. Outperforms 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld (74% vs 65.1%). Terminal-Bench 2.1: 54% vs 31% on 3.1 Flash-Lite. Computer use built-in. Configurable thinking levels for cost/latency tradeoffs.
Best: High-throughput agents, agentic search, document processing
NEW
Gemini 3.5 Pro
2M context — largest at the frontier
In
$1.25
per 1M
Out
$10.00
per 1M
2M ctxDeep ThinkGA Jul 17
Google's new flagship, launched July 17. 2-million-token context window — double anything at the frontier. New Deep Think extended reasoning mode (gated to $250/mo Ultra tier). Rebuilt from scratch on a new pretraining run after the original failed on recursive tool-calling. $1.25/$10 — 4x cheaper input than GPT-5.6 Sol.
Best: Massive context, video analysis, long-horizon reasoning, RAG-replacement
Gemini 3 Pro
3rd gen flagship
In
$2.00
per 1M
Out
$12.00
per 1M
1M ctx
Third-generation Gemini Pro. Now stable — no Preview tag. Same / pricing as preview tier. Strong general-purpose flagship.
Best: General production apps, stable Pro performance
FLAGSHIP
Gemini 3.1 Pro Preview
New flagship
In
$2.00
per 1M
Out
$12.00
per 1M
4h horizonPreviewVideo
77.1% ARC-AGI-2. Price increased from $1.25/$10. Batch and Flex tiers at 50% off.
Best: Video analysis, complex reasoning
LONG
Gemini 2.5 Pro
Long outputs
In
$1.25
per 1M
Out
$10.00
per 1M
1M ctx64K output
Same price as 3.1 Pro but 64K max output vs 16K. Choose for long-form content generation.
Best: Long-form writing, large outputs
DEPRECATED
Gemini 2.0 Flash
Shuts down Jun 1
In
$0.15
per 1M
Out
$0.60
per 1M
1M ctx8K outputRetiring Jun 1
Deprecated — shuts down June 1, 2026. Migrate to Gemini 2.5 Flash or 3.1 Flash-Lite.
Best: Migrate away from this model
⬆
Open Source & Local
Open-weight models you can run yourself or call via cheap APIs. The frontier is no longer closed.
NEW
GLM-5.3
Z.ai's open-weight coding flagship — 743B params
In
$1.40
per 1M
Out
$4.40
per 1M
743B paramsOpen weight (pending)CodingGLM Coding Plan & ZCode
Z.ai (formerly Zhipu AI) calls GLM-5.3 the most capable open-weights model for coding. 743B-param model built by scaling post-training on the GLM-5.2 base. Live now through GLM Coding Plan subscription and ZCode; API access and downloadable weights to follow after a safety review (~2 weeks). 'Dramatic improvement over GLM-5.2 with fewer output tokens.' GLM-5.2 API was ~$1.40/$4.40 — roughly a tenth of US frontier per-token rates.
Best: Agentic coding, open-weight workflows, cost-sensitive enterprise
NEW
DeepSeek V4-Pro
GA — DeepSWE 62.7%, peak/off-peak pricing
In
$0.66
per 1M
Out
$1.98
per 1M
128K ctxGA Aug 13DeepSWE 62.7%Terminal-Bench 87.9Off-peak price
DeepSeek's flagship out of preview (Aug 13) after 4 months. DeepSWE 62.7% (up from 12.8% in preview), Terminal-Bench 2.1: 87.9, AA Intelligence Index (reasoning): 53. New peak/off-peak billing from Aug 16 — off-peak $0.66/$1.98, peak $1.32/$3.96 (still below Western alternatives). 17 of 24 hours at half price. Open weights expected. Flexible reasoning (low/high/max) + thinking modes.
Best: Autonomous agents, software engineering, cost-sensitive high-volume workloads
NEW
Qwen 3.8-Max
Alibaba's flagship — 'second only to Fable 5'
In
$2.00
per 1M
Out
$6.00
per 1M
1M ctx (flat tier)2.4T MoE / 95B activeMultimodalOpen weights soonAA Index 58 pts
Alibaba's most capable model. 2.4T-param MoE with 95B active. $2/$6 per 1M — flat across the entire 1M-token context window (no long-prompt surcharge). Cached input at $0.25/1M. Claims 'second only to Fable 5'. AA Intelligence Index v4.1.1: 58 pts (#8 globally, between GPT-5.6 Sol high and GPT-5.6 Terra max). Multimodal (text + visual). Open weights promised within days of GA. OpenAI and Anthropic API compatible.
Best: Long-context work, coding, enterprise — flat pricing across 1M context
NEW
MiniMax M2
Agent & coding model — 8% of Sonnet's price
In
$0.30
per 1M
Out
$1.20
per 1M
Open weight~100 TPSAgent & codeFree until Nov 7
MiniMax's agent-first model. $0.30/$1.20 per 1M — 8% of Claude Sonnet's price at ~2x the speed (~100 TPS). Top 5 globally on Artificial Analysis Intelligence Index. Built for end-to-end dev workflows (Claude Code, Cursor, Cline, Kilo Code, Droid). Open weights on HuggingFace. Free API trial until Nov 7. MiniMax Agent product also free for a limited time.
Best: Agents, coding, tool use, cost-sensitive agentic workflows
NEW
Tencent Hy3
Global open API — cheapest per-token
In
$0.13
per 1M
Out
$0.53
per 1M
256K ctx295B MoE / 21B activeApache licenseOpen weights
Tencent's reasoning and agent model (formerly Hunyuan). 295B MoE with 21B active per token. Apache-licensed weights on HuggingFace. OpenRouter from $0.13/$0.53 per 1M — among the cheapest open APIs available. Topped OpenRouter usage leaderboard within a week of its Jul 6 launch. Global access via WorkBuddy (free until Aug 31), Tencent Cloud TokenHub, and API.
Best: Reasoning, agent tasks, cost-sensitive high-volume apps
NEW
Tencent Hy4 preview
770B open-weight — frontier-class under $1/1M input
In
$0.83
per 1M
Out
$2.50
per 1M
1M+ ctx770B MoE / 49B activeApache 2.0Open weights
Tencent's next-gen Hunyuan, launched Aug 28 2026. 770B total / 49B active MoE with a context window exceeding 1M tokens. Apache 2.0 weights on Hugging Face. API: $0.834/$2.501 per 1M ($0.042 cache hits) via Tencent Cloud TokenHub and OpenRouter. Free for 2 weeks on WorkBuddy & CodeBuddy; Hy3 free access extended to Sep 30. Targets coding, office work, game dev, and scientific research. More Hy4-series models coming soon.
Best: Long-document work, coding, office productivity at open-weight prices
NEW
Kimi K3
World's largest open-weight — 2.8T params
In
$3.00
per 1M
Out
$15.00
per 1M
1M ctxMultimodalMoE 2.8TAA Index #3 (57pts)Open weights Jul 27
Moonshot AI's flagship. 2.8-trillion-parameter MoE — world's largest open-weight model. Scores 57 on the Artificial Analysis Intelligence Index — third overall (behind Fable 5 ~60, GPT-5.6 Sol ~59), first open-weight model ever in the top 3. First on LMArena Frontend Code Arena (1,679 Elo). Native multimodal (text + images), 1M context. $3/$15 — averages $0.94 per AA Index task vs Sol $1.04, Opus 4.8 $1.80. Full weights public Jul 27 (1.56TB). Shook US chip stocks on release.
Best: Long-form coding, complex reasoning, enterprise open-weight
NEW
Kimi K2.6
88% cheaper than Opus
In
$0.60
per 1M
Out
$2.50
per 1M
256K ctxOpen weightMoE 1T/32B
Beats GPT-5.4 and Opus 4.6 on SWE-Bench Pro. 1T params, 32B active. 300 sub-agent orchestration. OpenAI-compatible API.
Best: Coding, agents, long-horizon tasks
CHEAPEST
Qwen 3.6 Plus
1M context, free tier
In
$0.10
per 1M
Out
$0.30
per 1M
1M ctxReasoningFree tier
Alibaba's latest. Mandatory chain-of-thought reasoning. Free tier available. Topped 6 coding benchmarks on release.
Best: Budget coding, massive context
NEW
Llama 4 Scout
10M context MoE
In
$0.15
per 1M
Out
$0.55
per 1M
10M ctxOpen weightMoE 109B
Longest context of any open model. 109B total, 17B active. Multimodal. Runs on 24GB VRAM.
Best: Massive context, multimodal, local
Llama 4 Maverick
Frontier coding MoE
In
$0.20
per 1M
Out
$0.80
per 1M
1M ctxOpen weightMoE 400B
Beats GPT-4o on coding. 400B total, 17B active. 128 experts. Frontier quality at MoE prices.
Best: Coding, complex reasoning
DeepSeek V3.2
Matches GPT-4o
In
$0.27
per 1M
Out
$1.10
per 1M
128K ctxOpen weightMoE 685B
94.2% MMLU matching GPT-4o. 685B MoE with 37B active. Best open model for general knowledge.
Best: General knowledge, research
FREE
Qwen3-Coder 8B
Local coding king
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
32K ctxLocal only8B dense
Runs on any 8GB GPU. 92 programming languages. 80-150 tok/s. Best local coding model under 10B. Set it up locally →
Best: Local coding, autocomplete
FREE
DeepSeek R1 Distill 14B
Local reasoning
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
Local onlyReasoning10GB VRAM
Chain-of-thought reasoning on 10GB VRAM. The sweet spot for local reasoning. 55 tok/s on modern GPUs. Run it offline →
Best: Local reasoning, budget hardware
💡 Did You Know?
GPT-6 Astra — OpenAI's $10/$50 flagship
OpenAI announced GPT-6 Astra on Sep 3 and shipped the API model Sep 4. $10/$50 per 1M — 2.5x GPT-5.6 Sol's $4/$20 promo — with $1 cached input, $12.50 cache writes, and long-context rates above 272K tokens (2x input/cache, 1.5x output). Batch/Flex at 50%, Fast mode at 2x. 1.05M context, 128K output. Saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98%); ExploitBench 100% without safeguards. First model to hit OpenAI's critical cybersecurity threshold. AA Intelligence Index: 61, but ~70% more token-efficient than Sol — less than half Fable 5's cost per coding task.
Fable 5.1 — the cache-read cut is the story
Claude Fable 5.1 (Sep 1) keeps $10/$50 but cuts cache reads 75% to $0.25/1M. Anthropic's own usage data: ~25% total savings on typical workloads, up to ~45% for agentic work where cache hits dominate. Terminal-Bench-Science 52.6% (more than doubles Fable 5), AutomationBench 31.4% (vs 17.1%), AA Index 66 — #1. Candid caveats from the system card: forced tool use now returns 400, thinking blocks are model-bound, and Mythos 5.1 is 'less honest under pressure' than recent Claude models.
Gemini 3.8 Flash — fourth Flash in four months
Google shipped Gemini 3.8 Flash on Sep 2 — the fourth Flash since May (3.5, 3.6, 3.7, 3.8). AA Index 59 (+3 over 3.7) at the same $0.75/$3.75 intro through Dec 31, then $1.50/$7.50. ~300 tok/s — the fastest output speed Artificial Analysis has measured. Trained on long-running agentic loops, it 'works harder': ~40% higher cost per task than 3.7 (30% more output tokens), so use low effort for efficiency-first routing. Gemini 3.8 Flash Cyber reaches trusted defenders via the Fairwind Program — 2.6x more correct Chrome patches than larger frontier models.
Muse Spark 1.3 — cheapest model at the frontier
Meta's Muse Spark 1.3 (Sep 2) scores 61 on the AA Intelligence Index at xhigh (62 at max, partner preview) — behind only Fable 5.1 (66) and Opus 5 (63). $1.25/$4.25 unchanged, $0.55 per Index task vs $0.94–0.95 for GPT-5.6 Sol and Grok 4.6 — cheapest of any model at 59+. #1 on Tau3-Bench Banking. Meta teases open weights and larger Muse models around Q4.
September's price map
Four frontier launches in 72 hours (Sep 1–3) split the market two ways: general intelligence gets cheaper per task (Fable 5.1 cache reads $0.25/1M, Muse Spark 1.3 at $0.55/task, Gemini 3.8 Flash at $0.58/task), while cyber-capability gets gated (Mythos 5.1, Gemini Flash Cyber, OpenAI Daybreak). Budget calendar: GLM-5.3-Flash promo ends Sep 9 (list $0.15/$0.50), GPT-5.6 Sol promo ends Nov 21 ($4/$20 → $5/$30), Gemini 3.8/3.7 intro ends Dec 31 (doubles Jan 1, 2027). Top of the AA Intelligence Index: Fable 5.1 (66), Opus 5 (63), Muse Spark 1.3 max (62), Fable 5 (62), Sol / Astra / Spark xhigh (61).
Grok 4.6 — frontier at $2/$6
xAI's Grok 4.6 (Aug 12) launched with Cursor at $2/$6 per 1M, $0.50 cached input, with a 1753 ELO claim that overtakes Kimi K3. Live on xAI API, Grok Build, Cursor, and Grok Bot, plus partners OpenRouter, Vercel, and Cloudflare. A fast variant costs twice the price. xAI frames it as roughly half the cost of other frontier models. 2x included usage in Cursor and Grok Build for the first week.
GLM-5.3 — Z.ai's 743B open-weight coder
Z.ai (formerly Zhipu AI) released GLM-5.3 on Aug 14 — a 743-billion-parameter model built by scaling post-training on the GLM-5.2 base. Z.ai calls it the most capable open-weights model for coding. Live via GLM Coding Plan and ZCode; API and downloadable weights follow after a safety review (~2 weeks). GLM-5.2's API was ~$1.40/$4.40 — roughly a tenth of US frontier per-token rates. 'Dramatic improvement over GLM-5.2 with fewer output tokens.'
DeepSeek V4-Pro — GA + peak/off-peak pricing
DeepSeek's V4-Pro went generally available Aug 13 after 4 months in preview. DeepSWE 62.7% (up from 12.8% in preview), Terminal-Bench 2.1: 87.9, AA Intelligence Index (reasoning): 53. DeepSeek also overhauled its API billing — replacing flat rates with peak and off-peak pricing effective Aug 16. Off-peak: $0.66/$1.98 per 1M. Peak: $1.32/$3.96. Still below Western alternatives (Anthropic's Fable 5 charges $50/M output). 17 of 24 hours stay at half price.
Qwen 3.8-Max — flat $2/$6 across 1M context
Alibaba's Qwen 3.8-Max went GA on Aug 3 at $2/$6 per 1M tokens — flat across the entire 1M-token context window with no long-prompt surcharge. 2.4T-param MoE (95B active), multimodal. AA Intelligence Index v4.1.1: 58 pts — #8 globally, right between GPT-5.6 Sol (high) and GPT-5.6 Terra (max). Cached input at $0.25/1M. Open weights promised within days. Undercuts Kimi K3 ($3/$15) and GPT-5.6 Terra ($2/$12) on output.
MiniMax M2 — agent model at 8% of Sonnet's price
MiniMax M2 (Aug 3) is an agent-and-code model at $0.30/$1.20 per 1M — 8% of Claude Sonnet's price at ~2x the inference speed (~100 TPS). Top 5 globally on the Artificial Analysis Intelligence Index. Full open weights on HuggingFace. Built for Claude Code, Cursor, Cline, Kilo Code, and Droid. Free API trial until Nov 7; MiniMax Agent product also free for a limited time.
Tencent Hy3 — cheapest open API per token
Tencent Hy3 (formerly Hunyuan) went globally available Aug 5. 295B MoE (21B active), 256K context. Apache-licensed weights on HuggingFace. OpenRouter from $0.13/$0.53 per 1M — among the cheapest open APIs on the market. Topped OpenRouter usage within a week of its Jul 6 launch. Free via WorkBuddy until Aug 31.
Claude Opus 5 — near-Fable-5 at half the price
Anthropic launched Claude Opus 5 on Jul 24, 2026 — $5/$25 per 1M (same as Opus 4.8, half of Fable 5's $10/$50). It tops the Artificial Analysis Intelligence Index at 61 points, ahead of Fable 5 (60) and GPT-5.6 Sol (59). ARC-AGI-3: 30.2% — 4x GPT-5.6 Sol, 20x Opus 4.8. New default on Claude Max. Fast mode runs 2.5x faster at $10/$50 (API research preview). Batch API at $2.50/$12.50. Opus 5 is Anthropic's fourth model launch in under two months.
Gemini 3.5 Pro ships — 2M context
Google DeepMind launched Gemini 3.5 Pro on Jul 17 with a 2-million-token context window — double anything at the frontier. New Deep Think reasoning mode (gated to $250/mo Ultra). Rebuilt from scratch on a new pretraining run after the original failed on recursive tool-calling. $1.25/$10 — 4x cheaper input than GPT-5.6 Sol.
Kimi K3 — world's largest open weight
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model — the largest ever released. $3/$15 per 1M, 1M context, native multimodal. A deliberate shift away from ultra-low-price strategy toward enterprise performance competition. Full weights public Jul 27. Shook US chip stocks on release.
GPT-5.6 goes GA
GPT-5.6 launched July 9 in three tiers — Sol ($5/$30), Terra ($2.50/$15), Luna ($1/$6). Sol beats Fable 5 on Agents' Last Exam by 13 points at ~1/4 the cost. Luna outperforms Opus 4.8 at ~1/16 the cost. First frontier launch with US government coordination.
Meta joins the API game
Muse Spark 1.1 is Meta's first ever paid API model — $1.25/$4.25, ~25% cheaper than OpenAI and Anthropic. Meta Superintelligence Labs, run by former Scale AI chief Alexandr Wang, built it for agentic tasks and tool use. $20 free credits for new accounts.