In July we called GPT-5.6 Sol the new top of the frontier. Two and a half months later, that ranking is ancient history. OpenAI shipped GPT-6 Astra on September 3. Google shipped Gemini 3.8 Flash on September 2. And on September 22, Anthropic released Claude Opus 5.5 the same day OpenAI answered with two cheaper GPT-6 models, Sol and Luna. It was the most crowded month for frontier AI we've seen.
Our verdict: Claude Opus 5.5 is the best AI model right now. It's #1 on the main independent intelligence index, leads SWE-bench Pro, and costs 40% of GPT-6 Astra's price with no long-context surcharge. GPT-6 Astra still wins on science and automation. GPT-6 Sol is the best value workhorse. And Fable 5.1, Anthropic's former flagship, has been out-priced by its own sibling.
The Short Version
Best overall: Claude Opus 5.5 — #1 on Artificial Analysis (58), 89.9% SWE-bench Pro, $4/$20.
Best for science and automation: GPT-6 Astra — but $10/$50, and $20/$75 past 272K input tokens.
Best value workhorse: GPT-6 Sol at $2/$10.
Cheapest good-enough model: GPT-6 Luna at $0.10/$0.50 — see our budget model verdict.
The September 2026 Frontier Table
| Model | Released | Price per 1M tokens (in / out) | AA Intelligence Index | SWE-bench Pro |
|---|---|---|---|---|
| Claude Opus 5.5 | Sep 22 | $4 / $20 | 58 (#1) | 89.9% |
| GPT-6 Astra | Sep 3 | $10 / $50 ($20 / $75 over 272K) | 53 | — |
| Claude Fable 5.1 | Summer update | Premium tier | 53 | 81.2% |
| GPT-6 Sol | Sep 22 | $2 / $10 | Behind Opus 5.5 | — |
| Gemini 3.8 Flash | Sep 2 | $0.75 / $3.75 (promo) | Fast tier | — |
A note on benchmarks: vendors publish tables with different task sets, and only seven tests appear in both Anthropic's and OpenAI's launch tables. Opus 5.5 leads four, Astra leads two, and one isn't comparable. Independent scores run a few points below Anthropic's own. We weight the independent Artificial Analysis index most heavily, and treat everything else as directional.
Claude Opus 5.5: The New Default
Opus 5.5 was built for long-running agentic coding and knowledge work, and it shows. It tops SWE-bench Pro at 89.9% — ahead of Fable 5.1 at 81.2% and Mythos 5 at 80.3% — and holds the #1 spot on the Artificial Analysis Intelligence Index at 58. It has a 1M token context window and costs $4 input and $20 output per million tokens, with no surcharge for long prompts. Anthropic also cut Opus pricing about 20% versus the previous generation.
The catch: at maximum effort, Opus 5.5 thinks hard and burns a lot of tokens. Some independent testers put a hard agentic task at around $13, so per-token savings can disappear if you run everything on max. The practical move is to use default effort and only turn it up for the hardest problems.
GPT-6 Astra: Still the Science Pick
Astra launched first, on September 3, with a heavy emphasis on coding, research, computer use and multi-step professional tasks. It still leads on science-heavy reasoning and automation workflows. But at $10/$50 it costs 2.5x Opus 5.5, and its pricing jumps to $20/$75 for the entire request once input passes 272,000 tokens. For long-document work, that surcharge is the deciding factor.
GPT-6 Sol: OpenAI's Real Answer
The more important OpenAI launch is Sol, released the same day as Opus 5.5 at $2/$10 — roughly half the cost of the GPT-5.6 generation. It trails Opus 5.5 in independent tests, but it's the model most teams already on OpenAI should route everyday work to. If you're still paying GPT-5.6 Sol prices, switch.
What About Fable 5.1 and Gemini?
Fable 5.1 is still an excellent writer, and some teams will keep it for brand-voice work. But Opus 5.5 now matches it on most tasks for less money, which makes Fable hard to justify as a default. Gemini 3.8 Flash isn't a flagship; it's Google's fast, cheap tier and one of the better budget picks. We compare it with GPT-6 Luna and Haiku 4.5 in our cheapest AI model verdict.
How to Route Your Work Now
- Agentic coding and hard engineering: Claude Opus 5.5 at default effort, max effort only for the hardest tasks.
- Science research and computer-use automation: test GPT-6 Astra, and keep prompts under 272K tokens.
- Everyday production workloads on OpenAI: GPT-6 Sol.
- High-volume classification, extraction and chat: GPT-6 Luna or Gemini 3.8 Flash.
- Privacy-sensitive or self-hosted: the open-weight models in our non-cloud LLM guide.
The Verdict
Claude Opus 5.5 is the best AI model of September 2026: the top independent score, the top coding score and a price that undercuts its closest rivals by 60%. GPT-6 Astra is the specialist pick for science and automation. GPT-6 Sol is the best value for teams on OpenAI. Expect this table to change again within weeks — route by task, not by brand loyalty.
FAQ
What is the best AI model in September 2026?
Claude Opus 5.5 is the best overall model right now. Artificial Analysis ranks it #1 of 212 models on its Intelligence Index at 58, five points ahead of GPT-6 Astra and Claude Fable 5.1, and it leads SWE-bench Pro at 89.9%. GPT-6 Astra still wins some science and automation tasks.
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with a 1M token context window and no long-context surcharge. That is 40% of GPT-6 Astra's $10/$50 list price.
Is GPT-6 Astra worth the higher price?
For most teams, no. Astra costs $10/$50 per million tokens and jumps to $20/$75 once a request passes 272,000 input tokens. It is worth testing for science-heavy research and computer-use automation, where it still leads, but Opus 5.5 is cheaper and ahead on most shared benchmarks.
What happened to Claude Fable 5?
Fable 5 was updated to Fable 5.1, which scores 81.2% on SWE-bench Pro and 53 on the Artificial Analysis index. Claude Opus 5.5, released September 22, now matches or beats Fable 5.1 on most tasks at a lower price, so most teams should default to Opus 5.5.
What is GPT-6 Sol?
GPT-6 Sol is OpenAI's cheaper GPT-6 model, released September 22 at $2 per million input tokens and $10 per million output tokens. It roughly halves the cost of the GPT-5.6 generation and is OpenAI's best value workhorse, though independent tests place it behind Opus 5.5.