Verdict

Cheapest Good AI Model 2026: GPT-6 Luna vs Gemini 3.8 Flash vs Claude Haiku 4.5

The frontier fight gets the headlines, but the bigger September story for anyone paying an AI bill is at the bottom of the price list. OpenAI released GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output. That's a tenth of Claude Haiku 4.5. For the huge amount of AI work that's really classification, extraction, routing, summarizing and simple chat, the cost of running a capable model just fell off a cliff.

Our verdict: GPT-6 Luna is the cheapest good AI model in 2026 and the new default for high-volume work. Gemini 3.8 Flash is the step-up pick when you need better reasoning and agent behavior without frontier prices. Claude Haiku 4.5 is now the expensive option in the budget tier, worth keeping mainly if you're already built on Anthropic.

A glowing cyan microchip surrounded by gold coins dissolving into light particles
The price floor dropped: capable AI now costs cents per million tokens. Illustration: TechVerdict.
Compared GPT-6 Luna Gemini 3.8 Flash Claude Haiku 4.5

The Short Version

Cheapest good model: GPT-6 Luna, $0.10 / $0.50.

Best budget model for agents and coding: Gemini 3.8 Flash, $0.75 / $3.75 promo through Dec 31.

Budget pick for Anthropic shops: Claude Haiku 4.5, $1 / $5.

Price Comparison

ModelReleasedInput / 1MOutput / 1MCost of 1B output tokens
GPT-6 LunaSep 22, 2026$0.10$0.50$500
Gemini 3.8 FlashSep 2, 2026$0.75 (promo)$3.75 (promo)$3,750
Claude Haiku 4.52025$1.00$5.00$5,000

At a billion output tokens a month — realistic for a busy support bot or document pipeline — the gap between Luna and Haiku is $4,500 a month. That's the whole reason this comparison matters.

GPT-6 Luna: The New Default for Volume

Luna is OpenAI's smallest GPT-6 model and it's priced to take the entire high-volume market. OpenAI halved Luna's price versus the previous generation, and it inherits the mature OpenAI tooling: function calling, structured outputs and the broadest third-party ecosystem. Use it for tagging and classifying, extracting fields from documents, routing requests to bigger models, first-pass summaries and simple customer chat.

The catch: small models still fail on long multi-step reasoning and tricky edge cases. Put a quality check on anything customer-facing, and escalate hard requests to a bigger model.

Gemini 3.8 Flash: The Smart Middle

Google shipped Gemini 3.8 Flash on September 2, just three weeks after 3.7 Flash, and it's noticeably better at software engineering, agent workflows and multi-step reasoning than the model it replaced. At $0.75/$3.75 on promotional pricing through December 31, it costs more than Luna but meaningfully less than any frontier model. It's the pick for lightweight agents, coding helpers and anything that needs more thinking than a classifier. Google also released a Cyber variant that posted frontier-level results on autonomous vulnerability discovery.

The catch: promotional pricing ends December 31, so budget for a price change in 2027.

Claude Haiku 4.5: Good, Now Expensive

Haiku 4.5 is still a solid, fast, well-behaved model, and Anthropic's writing quality shows even at the small tier. But at $1/$5 it now costs ten times Luna and more than Gemini 3.8 Flash's promo price. If your stack already runs on Claude — Claude Code, Skills, Anthropic's SDK — the consistency may be worth the premium. If you're choosing fresh, start elsewhere.

When to Stop Being Cheap

Budget models are the right default for most requests, not all of them. When a wrong answer is expensive — hard coding tasks, long agent runs, legal or financial analysis, anything where quality is the product — route to a frontier model. Right now that's Claude Opus 5.5 or GPT-6 Sol, compared in our September frontier ranking. The best setups use a cheap model to triage and a strong model for the hard 10%.

The Verdict

GPT-6 Luna is the cheapest good AI model of 2026 and the new default for high-volume work at $0.10/$0.50. Gemini 3.8 Flash is the better budget model when you need reasoning and agents. Claude Haiku 4.5 makes sense mainly for teams already committed to Anthropic.

FAQ

What is the cheapest good AI model in 2026?

GPT-6 Luna, released September 22, 2026, at $0.10 per million input tokens and $0.50 per million output tokens. That is one tenth of Claude Haiku 4.5 and far below Gemini 3.8 Flash's promotional $0.75/$3.75.

Is Gemini 3.8 Flash better than GPT-6 Luna?

Gemini 3.8 Flash is the stronger model for software engineering, agent workflows and multi-step reasoning, and it is priced like a step-up tier at $0.75/$3.75 through December 31, 2026. GPT-6 Luna is the better pick when volume and cost matter more than peak quality.

Is Claude Haiku 4.5 still worth it?

Mostly for teams already built on Anthropic's API and tools. At $1/$5 per million tokens it now costs ten times GPT-6 Luna, and both newer budget models compete closely on quality. If you are choosing fresh, start with Luna or Gemini 3.8 Flash.

When should I pay for a frontier model instead?

When a wrong answer is expensive: hard agentic coding, long multi-step tasks, legal or financial analysis and anything customer-facing where quality is the product. For those, route to Claude Opus 5.5 or GPT-6 Sol and keep budget models for classification, extraction, routing and simple chat.

Until then, every verdict lives here.