This page tracks whether the AI investment boom is sustainable โ using five data series that the professionals watch but rarely explain in plain English. How much are the labs spending vs earning? How fast are token prices falling? Is DeepSeek actually as cheap as claimed? Which model leads on benchmarks right now? And what does the profitability trajectory actually look like?
Not investment advice. Figures compiled from public company disclosures, Epoch AI, Bloomberg, The Information and benchmark leaderboards. Updated monthly. All costs in USD.
Lab Spending vs Revenue โ Are They Going to Be Profitable?
The AI labs are growing revenue at extraordinary speed. But training and infrastructure costs are growing at a similar pace. The key question is whether revenue growth outpaces cost growth fast enough to reach sustainable profitability before investor patience runs out.
The Path to Profitability
The AI lab business model works as follows: inference (running the model for customers) earns roughly 30% gross margin at current pricing. That sounds healthy โ but every dollar of gross profit is being reinvested into training the next model generation, which is more expensive than the last. The result is strong revenue growth but near-zero GAAP profitability at the company level.
OpenAI's revenue trajectory is genuinely impressive โ from $2bn ARR in 2023 to $12.7bn in Q1 2026 to a $40bn run-rate by late 2026. That is 150% growth in one year. But training GPT-4 cost an estimated $100m-$500m; training GPT-5-era models costs materially more. The next generation will cost more still.
The revenue breakdown is approximately: 65% from ChatGPT subscriptions (Plus $20/month, Pro $200/month, Team $25-30/user, Enterprise custom), 25% from API access (token-based), 10% from partnerships (Microsoft, Apple etc.).
| Date | Lab | Revenue Run-Rate | Est. Cost Base | Operating Margin | Valuation |
|---|---|---|---|---|---|
| Q1 2023 | OpenAI | ~$1bn ARR | ~$3bn | Deeply negative | $29bn |
| Q4 2024 | OpenAI | ~$4bn ARR | ~$8bn | Negative | $157bn |
| Q4 2024 | Anthropic | ~$1.3bn ARR | ~$3bn | Deeply negative | $61bn |
| Q4 2024 | xAI / Grok | ~$0.3bn ARR | ~$1bn | Pre-revenue scale | $50bn |
| Q1 2026 | OpenAI | $12.7bn ARR | ~$18bn | ~30% gross margin, ~0% EBIT | $300bn |
| Q1 2026 | Anthropic | ~$2.5bn ARR (est) | ~$5bn | Negative โ heavy training spend | $61bn |
| Q1 2026 | xAI / Grok | ~$1bn ARR (est) | ~$2bn | Negative | $80bn |
| Sep 2026 (est) | OpenAI | ~$40bn run-rate | ~$41bn | Approaching break-even | $300bn+ |
| Sep 2026 (est) | Anthropic | ~$4bn run-rate | ~$7bn | Still negative | $61bn |
| Sep 2026 (est) | xAI / Grok | ~$2bn run-rate | ~$3bn | Negative | $80bn+ |
Token Cost Collapse โ 99% Cheaper in Three Years
One of the most significant and underreported stories in AI is how fast the cost of using these models has fallen. In March 2023, GPT-4 cost $60 per million output tokens. By September 2026, comparable quality costs under $1 per million output tokens โ a fall of over 98% in three years. This is faster than Moore's Law and is reshaping who can afford to build AI products.
What a Token Actually Costs in Practice
A token is roughly 0.75 words, or 4 characters. A typical ChatGPT conversation uses approximately 2,000-5,000 tokens. A typical API call for a business task (summarise a document, analyse some data, generate a report) uses 5,000-50,000 tokens. Here is what that costs at current prices:
| Task | Approx Tokens | GPT-5.6 Sol (premium) | Claude Sonnet 5 / Gemini 3.1 | Grok 4 (mid) | DeepSeek V4-Flash (cheapest) |
|---|---|---|---|---|---|
| Single ChatGPT conversation | 3,000 | $0.10 | $0.035 | $0.022 | $0.0009 |
| Summarise a 10-page report | 15,000 | $0.50 | $0.17 | $0.11 | $0.005 |
| Analyse a full contract (50 pages) | 75,000 | $2.48 | $0.85 | $0.56 | $0.025 |
| Process 1,000 customer emails | 500,000 | $16.50 | $5.70 | $3.75 | $0.17 |
| Full codebase review (large app) | 2,000,000 | $66.00 | $22.80 | $15.00 | $0.70 |
Prices per million tokens: GPT-5.6 Sol $5/$30, Claude Sonnet 5 $3/$15 (blended ~$3.80/M), Grok 4 $2/$10 (blended ~$2.50/M), DeepSeek V4-Flash $0.14/$0.28. Actual costs depend on input/output ratio and caching. September 2026.
DeepSeek โ Is It Really That Cheap, and What's the Catch?
DeepSeek โ a Chinese AI lab spun out of quantitative hedge fund High-Flyer โ caused a global market shock in January 2025 when it released DeepSeek R1: a reasoning model that matched GPT-4o on most benchmarks, cost $0.55 per million input tokens (vs $2.50 for GPT-4o), and was released as open source under MIT licence. NVIDIA's share price fell 17% in a single day.
By April 2026, DeepSeek released V4-Pro and V4-Flash โ with V4-Flash benchmarking above GPT-4o on coding tasks at $0.14/$0.28 per million tokens. That is approximately 107 times cheaper on output tokens than GPT-5.5. The question everyone is asking: how is this possible, and is the advantage real?
How DeepSeek Achieves Lower Costs โ The Real Explanation
Mixture of Experts (MoE) architecture: DeepSeek V4-Pro has 1.6 trillion total parameters but only activates approximately 49 billion for any given token. This means the compute cost per inference call is equivalent to a 49B model, not a 1.6T model โ giving frontier-scale knowledge at a fraction of the inference cost.
Training efficiency: DeepSeek claimed V3 was trained for approximately $5.6 million in compute costs โ compared to estimates of $100m-$500m+ for comparable Western models. Independent researchers believe this is plausible because DeepSeek used aggressive hardware optimisation, knew exactly what techniques to use from published research, and had access to cheaper Chinese compute infrastructure. However the $294,000 training cost figure cited for earlier versions is likely misleading โ it counts only the final training run, not the full research and development cost.
The catch โ data and safety: DeepSeek models are trained in China, subject to Chinese regulations. They will not discuss Tiananmen Square, Taiwan, Xinjiang or other politically sensitive topics. For applications where political neutrality matters, this is a genuine limitation. For code generation, data analysis, and most business tasks, it rarely matters.
The catch โ export controls: DeepSeek used NVIDIA H100 GPUs in earlier training runs. Since then, US export controls have tightened further. Whether DeepSeek can continue training frontier-scale models under current restrictions is genuinely uncertain โ this is the biggest risk to its long-run competitive position.
Users, Adoption and What People Are Paying
The user figures for AI tools are staggering by any historical comparison for technology adoption โ but the monetisation rate per user is still low relative to the infrastructure investment being made.
The comparison to Google is instructive. Google earns approximately $120 per user per year through AI-enhanced advertising. OpenAI earns approximately $12 per total user per year ($120 per paying user ร ~10% conversion). Google's monetisation model โ where AI makes existing search advertising more effective โ may prove more durable than the subscription model, where users pay explicitly for AI access.
Enterprise adoption is growing faster than consumer and commanding higher prices. ChatGPT Enterprise (custom pricing, typically $60-100 per user/month for large deployments) is driving a growing share of revenue. API usage โ where developers pay per token and build AI into their own products โ accounts for approximately 25% of OpenAI revenue but is under more competitive pressure from DeepSeek and open-source alternatives.
Model Capability Leaderboard โ September 2026
The benchmark landscape has evolved significantly. MMLU (general knowledge) is now approaching saturation โ top models score 88-94% and the differences are within noise. The more discriminating benchmarks are GPQA Diamond (graduate-level reasoning), SWE-Bench Verified (real software engineering tasks), and AIME 2025 (olympiad mathematics). These better separate genuine capability differences.
| # | Model | GPQA Diamond | SWE-Bench | MMLU-Pro | Input $/M | Released |
|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra Closed | 96.0% | 93% | 94% | $10+ | Aug 2026 |
| 2 | Claude Opus 5 Closed | 95% | 91% | 93% | $15 | May 2026 |
| 3 | Gemini 3.1 Pro Closed | 94.3% | 80.6% | 92% | $2 | Feb 2026 |
| 4 | GPT-5.5 Closed | 92% | 88.7% | 91% | $5 | Apr 2026 |
| 5 | DeepSeek V4-Pro OpenCN | 89% | 76.4% | 88% | $0.45 | Apr 2026 |
| 6 | Grok 4 Closed | 88% | 74% | 88% | $2 | Mar 2026 |
| 7 | Claude Sonnet 5 Closed | 87% | 72% | 87% | $3 | Mar 2026 |
| 8 | Qwen3-Max OpenCN | 85% | 70% | 86% | $0.40 | Apr 2026 |
| 9 | Kimi K2 OpenCN | 83% | 68% | 84% | $0.60 | Jun 2026 |
| 10 | DeepSeek V4-Flash OpenCN | 79% | 67% | 82% | $0.14 | Apr 2026 |
| 11 | Llama 4 Scout 109B Open | 76% | 62% | 80% | Self-host | Jan 2026 |
Sources: LLM Stats leaderboard, LocalAI Master, PrecisionAI Academy LLM Leaderboard, vendor model cards. Scores are rounded. GPQA Diamond measures graduate-level reasoning; SWE-Bench Verified measures real software engineering tasks on GitHub issues; MMLU-Pro measures breadth across 57 academic subjects. GPT-6 Astra available September 2026 to enterprise Trusted Access Program only. Prices as of September 2026.
HOW TO READ THIS PAGE
New data rows are added monthly. The signal badges at the top summarise the three original warning signals: revenue gap (is the money coming in fast enough?), data centre debt (are the lenders getting nervous?), and chipmaker capex (is TSMC being disciplined?). Sections 2-5 on token costs, DeepSeek, users and the leaderboard are updated monthly as new pricing and benchmark data is published.
The overall picture as of September 2026: AI revenues are growing fast but not yet faster than costs; token prices are falling at extraordinary speed which is good for users but compresses lab margins; DeepSeek's price advantage is real and structurally significant; model capability is advancing faster than most observers expected; and the question of long-run profitability for the labs remains genuinely open.