We Compared How Top AI Models Changed Pricing in the Last 12 Months
By Nishrath

In the last 12 months, the pricing of AI models has turned messy.
What was once a simple "better model = higher price" formula has now become a situation of multiple model launches, cheaper variants, prices up and down…
We tracked seven major AI providers across the last 12 months to get a better sense of what's happening in the industry.
How much does it cost to access each provider’s best AI today vs. a year ago?Â

Providers are pushing their top AI models toward more premium pricing, with higher prices justified by stronger performance, reasoning, speed, and capability.Â
How has the cost changed in terms of mid-tier AI models

The mid-tier market has become the main battleground for AI pricing. Companies are creating cheaper versions of model families or lowering prices on older models to serve different customer segments.
Note: For this comparison, we selected the mid-tier models directly after the frontier plan, even though many more models are available below these tiers. Â
A closer look at each provider's pricing trendsÂ
OpenAI went deep on multi-tier pricing strategy
By Nov last year, OpenAI had already built some smaller models around their GPT-5 model.
This year, in September 2026, OpenAI has launched GPT-6 Astra which is $10 /1Mtok input tokens and $50 /1Mtok output tokens for short context.

If we compare the two top models side by side, Astra's price is 8× higher for input and 5× higher for output.
But to serve different markets, they have added several layers between their most capable model.
For example, Sol costs $4/$20, Terra costs $2/$12, and Luna costs $0.20/$1.20 per 1M tokens, creating several price points between Astra and the low-cost end of the lineup.
2. Claude is pushing its advanced models higher while keeping lower tiers competitive
In late 2025, Claude's top model, Opus 4.5, was priced around $5 /1Mtok input tokens and $25 /1Mtok output tokensÂ
By September 2026, Anthropic launched a new series, Claude Fable 5.1 now costs $10 /1Mtok for input and $50 /1Mtok for output, pushing the newest model price 2x higher.Â

Though the company has a similar goal of OpenAI to be affordable, instead of introducing several generations, it has been reducing the prices of its older models.Â
Sonnet 4.5 was priced at $3/$15 last year, while Sonnet 5 is now $2/$10, representing a 33% reduction in both input and output pricing.
3. Gemini keeps expanding the Flash familyÂ
Gemini launched two new Flash models this year, 3.6 and 3.7, which are currently available at the same price.Â

If we compare Gemini 3.7 Flash with last year's Gemini 2.5 Pro, the newer Flash model is 0.6x cheaper on input price and 0.375x cheaper on output price.Â
Gemini has also been pushing around use-case-specific releases. Gemini 3.7 Flash is positioned for more complex coding workflows, while 3.6 Flash was built around speed, search, grounding, and general agentic work. Google also added agentic video understanding and more advanced control over thinking in 3.7.Â
4. DeepSeek’s AI pricing took a turn from cheap to time-basedÂ
In September 2025, DeepSeek-V3.2-Exp entered the market at $0.28/$0.42 per 1M tokens for cache-miss input and output. The model itself was already a pricing move, with DeepSeek said to have reduced inference costs and allowed API prices to fall by more than 50%.
Those prices then barely moved. From September 2025 through March 2026, DeepSeek kept the same $0.28 input / $0.42 output rates.
In April 2026, DeepSeek launched V4-Flash at $0.14 for cache-miss input and $0.28 for output. That was a 50% cut in input and 33% cut in output compared with V3.2, while the context window jumped from 128K to 1M tokens.Â
In August 2026, the company introduced peak and off-peak pricing for the first time. Through this, the V4-Flash moved to $0.22/$0.66 during off-peak hours and $0.44/$1.32 during peak hours, meaning the off-peak rate alone was already 57% higher on input and over 2x higher on output than the original April launch.

5. Grok kept the flagship at $2/$6 while making older models cheaper
xAI kept its flagship API price at $2 per million input tokens and $6 per million output tokens even as it moved through three flagship generations in 2026.Â
But once it introduced newer models at the existing flagship price, it pushed its previous models into cheaper tiers.

Grok also got strategic with its context pricing. Early flagship models like Grok 4.20 charged a single flat rate of $2/$6 per million input/output tokens across their entire context window
Grok's newest flagships, Grok 4.6 and 4.5, constrained that to 500K tokens of context and the $2/$6 rate only holds below 200K tokens. Cross that threshold, and the price jumps to $4/$12, a straight 2x increase on both input and output.
6. Mistral cut its flagship price by 75%
In December 2025, Mistral Large 3 launched at $0.50 per million input tokens and $1.50 per million output tokens, down from Mistral Large 2’s $2/$6, marking a 75% cut in both input and output pricing, one of the steepest flagship price cuts in the market.Â
But Mistral did not follow that same direction across all of its newer models.Â

Some were introduced at much higher prices despite being newer generations. Mistral Medium 3.5, for example, costs $1.50 per million input tokens and $7.50 per million output tokens, making it 3x the input price and 5x the output price of Mistral Large 3, which costs $0.50/$1.50.Â
7. Kimi steadily increased their model price
Unlike Anthropic or DeepSeek, Kimi's pricing shows no dips or reversals. Each new model generation, from K2 through K3, has cost more than the last.Â
In July 2025, Kimi K2 launched at $0.60 per million input tokens and $2.50 per million output tokens
K2.6 then raised pricing to $0.95/$4 in April, a 58% increase in input pricing and a 33% increase in output pricing from K2.5.
In July 2026, Kimi K3 launched at $3/$15, making it more than 3x the input price and nearly 4x the output price of K2.6. Compared with K2, K3 is 5x more expensive on input and 6x more expensive on output.

The one price that actually went down: cache tokensÂ
Cached tokens make it easier for providers to lower compute per request while making their models more attractive for high-volume workloads.Â
Anthropic had already priced cache reads below regular input, but between November 2025 and September 2026, it pushed the number much further.Â
For Claude Opus 4.5 in November, a cache read cost $0.50/M against $5/M for regular input.Â
By September, Fable 5.1 and Mythos 5.1 had a cache-read price of just $0.25/M, while regular input had risen to $10/M.

So while regular input doubled from $5/M to $10/M, cache reads fell by 50%, from $0.50/M to $0.25/M, making reused context 75% cheaper, dropping from 10% of the input price to just 2.5%.
Key takeaways
Most providers raised frontier pricing significantly. Five of the seven providers increased the pricing of their flagship models, with OpenAI seeing the steepest jump at +$8.75 for input and +$40 for output with GPT-6 Astra.
Providers are introducing cheaper variants and multiple model generations to serve different use cases. OpenAI alone now has four price points across Astra, Sol, Terra, and Luna, with input pricing ranging from $0.20 to $10.
Time-based pricing introduced by DeepSeek is a unique strategy. Its August 2026 peak and off-peak pricing for the same model can vary by 2x depending on the time of day.
Context length has also become a pricing factor. Some providers charge more when a request goes beyond a certain context length, usually around 200K tokens.Â
Cache pricing comes with its own trade-offs. Some providers charge more to write data to the cache before giving you a discounted read rate. For workloads with low repetition, you could end up paying that extra cost without using the cache enough to make it worthwhile.
FAQs
Does a lower token price always mean a cheaper AI model to use?+
Not always. Your total cost also depends on input and output volume, caching, context length, and how many requests you need to complete a task.
Why do AI providers charge different prices for input and output tokens?+
Input tokens are the information you send to a model. Output tokens are the response it generates. Providers price them separately because generating output generally requires more compute.
How does context length affect AI model pricing?+
Some providers charge more when requests exceed a certain context length. This means a model's advertised price may not apply to very large requests.