Why This Matters
AI API pricing determines who can build what. When GPT-4 launched in March 2023 at $30 per million input tokens, only well-funded companies could afford to integrate frontier intelligence into their products. Today, equivalent or superior capability is available from multiple providers at $2-3 per million tokens — and lightweight models cost as little as $0.10. This 90%+ collapse in the cost of intelligence is reshaping every industry that depends on language, code, analysis, or reasoning.
Pricing also reveals competitive strategy. When a provider cuts prices aggressively, it signals confidence in cost structure, a push for market share, or pressure from competitors. When a new provider enters at a fraction of incumbents’ pricing — as DeepSeek did in early 2025 — it can force an industry-wide repricing overnight. Every price change is a strategic signal, and tracking them in real time reveals the competitive dynamics that press releases and blog posts obscure.
For developers, enterprises, and investors, understanding AI pricing is no longer optional. The difference between $3 per million tokens and $0.30 per million tokens can mean the difference between a viable product and a money-losing one. This tracker exists to make that landscape legible.
Current Landscape
The AI API pricing landscape in mid-2026 is defined by three tiers of competition, aggressive price cuts, and an emerging consensus that inference costs will continue falling for years.
Frontier models — the most capable available — now cost between $2 and $15 per million input tokens depending on the provider and model. This tier includes OpenAI’s GPT-4.1, Anthropic’s Claude Opus 4, Google’s Gemini 2.5 Pro, and xAI’s Grok 3. These models handle complex reasoning, long-context analysis, and tasks requiring deep domain knowledge.
Mid-tier models occupy the sweet spot between capability and cost, priced between $0.50 and $3 per million input tokens. This includes Anthropic’s Claude Sonnet 4, OpenAI’s GPT-4.1 Mini, Google’s Gemini 2.5 Flash, and Mistral’s Large model. For most production applications — customer support, content generation, code assistance, data extraction — these models deliver sufficient quality at a fraction of frontier pricing.
Lightweight models are priced below $0.25 per million input tokens and serve high-volume, latency-sensitive use cases. OpenAI’s GPT-4.1 Nano, Google’s Gemini Flash Lite, and numerous open-weights models served through providers like Together AI, Fireworks, and Groq fall into this category. These models power real-time autocomplete, classification, and routing tasks where cost per query must be measured in fractions of a cent.
The overall trend is unmistakable: per-token costs have fallen approximately 60% year over year since 2023, with no sign of deceleration. Multiple forces are driving this — hardware improvements with each NVIDIA generation, architectural innovations like Mixture of Experts (MoE), inference optimization through speculative decoding and quantization, and competitive pressure from open-weights models that can be self-hosted at marginal compute cost.
Pricing Comparison: Major Providers (May 2026)
Frontier / Flagship Models
| Provider | Model | Input (per 1M tokens) | Output (per 1M tokens) | Context Window |
|---|---|---|---|---|
| OpenAI | GPT-4.1 | $2.00 | $8.00 | 1M |
| Anthropic | Claude Opus 4 | $15.00 | $75.00 | 200K |
| Anthropic | Claude Sonnet 4 | $3.00 | $15.00 | 200K |
| Gemini 2.5 Pro | $1.25–$2.50 | $5.00–$10.00 | 1M | |
| xAI | Grok 3 | $3.00 | $15.00 | 128K |
| DeepSeek | DeepSeek-V3 | $0.27 | $1.10 | 128K |
Mid-Tier Models
| Provider | Model | Input (per 1M tokens) | Output (per 1M tokens) | Context Window |
|---|---|---|---|---|
| OpenAI | GPT-4.1 Mini | $0.40 | $1.60 | 1M |
| Anthropic | Claude Haiku 3.5 | $0.80 | $4.00 | 200K |
| Gemini 2.5 Flash | $0.15–$0.30 | $0.60–$1.20 | 1M | |
| Mistral | Mistral Large | $2.00 | $6.00 | 128K |
| Mistral | Mistral Small | $0.10 | $0.30 | 128K |
Lightweight / Nano Models
| Provider | Model | Input (per 1M tokens) | Output (per 1M tokens) | Context Window |
|---|---|---|---|---|
| OpenAI | GPT-4.1 Nano | $0.10 | $0.40 | 1M |
| Gemini 2.0 Flash Lite | $0.075 | $0.30 | 1M | |
| DeepSeek | DeepSeek-V3 (via API) | $0.27 | $1.10 | 128K |
Prices reflect standard API rates as of May 2026. Volume discounts, committed-use pricing, and batch processing rates can reduce costs by 25-75% at scale.
Key Players
OpenAI has historically set the pricing benchmark that others react to. The company’s strategy has shifted from premium pricing (GPT-4 at $30/$60 per million tokens in 2023) to aggressive competition across all tiers. The launch of GPT-4.1 in April 2025 at $2/$8 represented a dramatic reduction from GPT-4 Turbo pricing, and the Nano tier at $0.10 input signals OpenAI’s intent to compete for high-volume workloads. OpenAI’s pricing power comes from scale — with the largest developer ecosystem and the most inference volume, they can amortize infrastructure costs across more queries than any competitor.
Anthropic occupies a distinctive pricing position. Claude Opus 4 at $15/$75 is the most expensive major API model, reflecting Anthropic’s positioning as the premium option for enterprises that prioritize safety, reliability, and quality over cost. Claude Sonnet 4 at $3/$15 is more competitively priced and handles the majority of production volume. Anthropic’s pricing strategy bets that enterprises will pay a premium for models with demonstrably lower hallucination rates and better instruction following — a bet that appears to be paying off given the company’s reported revenue growth.
Google is the most aggressive price cutter among the frontier labs. Gemini 2.5 Flash at $0.15-$0.30 per million input tokens offers near-frontier capability at a price point that undercuts most competitors by 5-10x. Google can afford this because it controls its own TPU silicon, data centers, and energy infrastructure end-to-end. The pricing signals a strategy of winning through volume and ecosystem lock-in rather than per-query margins.
DeepSeek disrupted the market in January 2025 when it released DeepSeek-V3 with benchmark performance rivaling GPT-4o at less than one-tenth the price. The company’s API pricing — $0.27 per million input tokens for a model competitive with offerings priced at $2-5 — forced every major provider to reassess their pricing structure. DeepSeek’s cost advantage stems from architectural efficiency (Mixture of Experts), training on less expensive hardware (pre-H100 GPUs), and China’s lower labor costs for the engineering team.
Mistral AI has carved out a position in the European market and among privacy-conscious customers. Pricing spans from Mistral Small at $0.10/$0.30 (competitive with the cheapest offerings) to Mistral Large at $2/$6. Mistral’s pitch combines competitive pricing with EU data sovereignty — a compelling combination for European enterprises navigating GDPR and the EU AI Act.
Open-weights inference providers — Together AI, Fireworks AI, Groq, and others — serve self-hostable models like Llama 4, Qwen, and Mistral variants at prices that often undercut the original providers. Groq’s custom LPU chips deliver Llama-class models at sub-$0.10 per million tokens with sub-100ms latency. These providers create a pricing floor that prevents any proprietary model from charging more than a modest premium over the open-weights alternative.
What We’re Tracking
JustSaid monitors pricing across the AI API ecosystem through automated field-level change detection. The system works as follows.
Pricing page monitoring. Every major provider’s pricing page is checked multiple times daily. When a price changes, the pipeline captures the exact before-and-after values, the timestamp, and the surrounding context. This creates a complete historical record of every price change across every provider.
New model launches. When a provider releases a new model, its pricing is immediately captured and compared against existing tiers. This reveals whether a launch represents a new capability tier, a replacement for an existing model, or a competitive response to another provider’s move.
Volume and committed-use pricing. Enterprise pricing tiers, committed-use discounts, and batch processing rates are tracked alongside standard API rates. These often change independently of headline pricing and can reveal whether a provider is gaining or losing large-scale customers.
Cross-provider comparisons. Normalized comparisons — adjusting for context window length, output quality on standardized benchmarks, and latency — provide an apples-to-apples view of cost-per-capability across providers. A model that costs twice as much but produces 3x fewer errors may be cheaper in effective terms.
Historical Price Trajectory
The speed of AI API price deflation is unprecedented in the history of software:
| Period | Benchmark Model | Input Price (per 1M) | Equivalent Today |
|---|---|---|---|
| Mar 2023 | GPT-4 (8K) | $30.00 | $2.00 (GPT-4.1) |
| Nov 2023 | GPT-4 Turbo | $10.00 | — |
| May 2024 | GPT-4o | $5.00 | — |
| Jan 2025 | DeepSeek-V3 | $0.27 | $0.27 |
| Apr 2025 | GPT-4.1 | $2.00 | $2.00 |
| May 2026 | Gemini 2.5 Flash | $0.15 | $0.15 |
From GPT-4’s launch to Gemini 2.5 Flash, the cost of a capable AI API call has fallen by approximately 99.5% in just over three years. For comparison, cloud compute costs (as measured by AWS EC2 pricing) fell roughly 90% over a decade from 2006 to 2016. AI pricing deflation is running at roughly 10x the speed of cloud computing’s already-remarkable cost curve.
Recent Developments
May 2026: Anthropic launches Claude Opus 4 and Sonnet 4. Opus 4 debuted at $15/$75, maintaining Anthropic’s premium positioning. Sonnet 4 at $3/$15 is roughly flat versus the prior Sonnet generation, reflecting Anthropic’s view that capability improvements justify holding price rather than cutting.
April 2026: Google cuts Gemini 2.5 Pro pricing. Google reduced Gemini 2.5 Pro input pricing by approximately 50% for prompts under 200K tokens, bringing it to $1.25 per million. The move widened Google’s price advantage over OpenAI and Anthropic’s flagship offerings.
Q1 2026: The “race to the bottom” in lightweight models. OpenAI, Google, and multiple inference providers engaged in aggressive price cutting for their smallest models. GPT-4.1 Nano at $0.10 and Gemini Flash Lite at $0.075 set new floors. The pricing pressure reflects that lightweight model inference is becoming a commodity — differentiation increasingly happens at the frontier tier.
January 2025: The DeepSeek shock. DeepSeek-V3’s release at $0.27 per million input tokens — for a model competitive with GPT-4o — forced the entire industry to reckon with the possibility that frontier-class AI could be delivered at a fraction of prevailing prices. Within weeks, multiple providers announced price cuts or accelerated plans for lower-cost model tiers.
Outlook
AI API pricing will continue to fall, but the rate of decline and the competitive dynamics will evolve in the second half of 2026 and beyond.
Inference cost floors are approaching. At some point, the cost of running a model through a GPU becomes the binding constraint, and further price cuts require either hardware breakthroughs or accepting losses. For lightweight models, pricing is already approaching the marginal cost of compute. Further reductions will require custom silicon (like Google’s TPUs or Groq’s LPUs) that reduce cost per token at the hardware level.
Frontier pricing will stabilize at a premium. The most capable models — those that can handle complex reasoning, long-context analysis, and agentic workflows — will continue to command a 10-50x premium over lightweight models. Enterprises will pay this premium because the cost of model errors (hallucinations, missed insights, failed tasks) exceeds the cost of the API call.
Usage-based pricing will face pressure from subscription models. OpenAI’s ChatGPT Pro ($200/month unlimited), Anthropic’s Claude Pro, and Google’s Gemini Advanced are all flat-rate subscription products. As these subscriptions grow, they create an alternative pricing model that could erode API revenue — developers may build on top of consumer products rather than paying per-token API rates.
Open-weights models will continue to set the pricing floor. Every open-weights release that matches proprietary model quality forces API providers to justify their premium. The gap between open and closed model capability is narrowing (see our Open Weights Tracker), and as it does, the premium that closed-model providers can charge narrows with it.
Frequently Asked Questions
Why do output tokens cost more than input tokens? Output tokens require the model to generate new content auto-regressively — each token depends on all preceding tokens, requiring a full forward pass through the model. Input tokens can be processed in parallel. The computational cost of generation is typically 2-5x higher than the cost of processing input, and pricing reflects this difference.
Is DeepSeek’s pricing sustainable or is it subsidized? DeepSeek’s pricing appears to reflect genuine cost advantages rather than below-cost subsidization. The company uses Mixture of Experts architectures that activate only a fraction of model parameters per query, reducing compute costs. Its training was conducted on older, less expensive GPU generations. And China’s lower engineering labor costs reduce operational overhead. Multiple providers serving open-weights models at similar price points validate that this cost structure is achievable.
How should I choose between a cheap model and an expensive one? The right model depends on your error tolerance and the cost of failures. For classification, routing, and simple extraction tasks, lightweight models at $0.10-$0.30 per million tokens are typically sufficient. For tasks where accuracy matters — legal analysis, medical information, financial reasoning — the 5-10% quality improvement from a frontier model can justify a 10-50x price premium. Run evaluations on your specific use case rather than relying on generic benchmarks.
Will AI API pricing ever reach zero? Effectively, yes — for basic capabilities. Google already offers a generous free tier for Gemini, and multiple providers offer free access to lightweight models with rate limits. But frontier capability will always have a cost because the hardware required to run the largest models has real and significant power and capital costs. The trend is toward free basic AI and premium pricing for the most capable models — similar to how basic cloud storage became free while high-performance compute remained paid.