AI INVESTMENT TRACKER โ€” SPENDING, REVENUE, TOKEN COSTS AND WHO'S WINNING

Lab Spend vs Revenue ยท Token Price Collapse ยท DeepSeek's Real Advantage ยท Model Capability Leaderboard ยท Updated Monthly
โš  WATCH
Signal 1 โ€” Revenue Gap
Narrowing slowly
โš  WATCH
Signal 2 โ€” Data Centre Debt
Elevated stress
โœ“ OK
Signal 3 โ€” Chipmaker Capex
Disciplined so far
Sep 2026
Last Updated

This page tracks whether the AI investment boom is sustainable โ€” using five data series that the professionals watch but rarely explain in plain English. How much are the labs spending vs earning? How fast are token prices falling? Is DeepSeek actually as cheap as claimed? Which model leads on benchmarks right now? And what does the profitability trajectory actually look like?

Not investment advice. Figures compiled from public company disclosures, Epoch AI, Bloomberg, The Information and benchmark leaderboards. Updated monthly. All costs in USD.

01 โ€” THE BIG QUESTION

Lab Spending vs Revenue โ€” Are They Going to Be Profitable?

The AI labs are growing revenue at extraordinary speed. But training and infrastructure costs are growing at a similar pace. The key question is whether revenue growth outpaces cost growth fast enough to reach sustainable profitability before investor patience runs out.

$12.7bn
OpenAI ARR (Q1 2026)
$40bn
OpenAI Revenue Run-Rate (late 2026)
~$41bn
OpenAI Total Cost Base (est. 2026)
~0%
OpenAI Operating Margin (GAAP)
~30%
Gross Margin on Inference
Consumed
That Gross Profit โ€” by Next Model Training

The Path to Profitability

The AI lab business model works as follows: inference (running the model for customers) earns roughly 30% gross margin at current pricing. That sounds healthy โ€” but every dollar of gross profit is being reinvested into training the next model generation, which is more expensive than the last. The result is strong revenue growth but near-zero GAAP profitability at the company level.

OpenAI's revenue trajectory is genuinely impressive โ€” from $2bn ARR in 2023 to $12.7bn in Q1 2026 to a $40bn run-rate by late 2026. That is 150% growth in one year. But training GPT-4 cost an estimated $100m-$500m; training GPT-5-era models costs materially more. The next generation will cost more still.

The revenue breakdown is approximately: 65% from ChatGPT subscriptions (Plus $20/month, Pro $200/month, Team $25-30/user, Enterprise custom), 25% from API access (token-based), 10% from partnerships (Microsoft, Apple etc.).

โš ๏ธ The honest profitability assessment: On current trajectories, OpenAI could reach GAAP profitability in 2027-2028 โ€” but only if training cost growth slows as efficiency improvements kick in (a bet on algorithmic progress) and revenue continues growing at 100%+ annually (a bet on continued adoption). Both are plausible but neither is certain. The most profitable layer in AI right now is not the labs โ€” it is NVIDIA (selling the GPUs) and the cloud providers (renting the compute). Labs remain revenue-rich but not yet reliably profitable.
DateLabRevenue Run-RateEst. Cost BaseOperating MarginValuation
Q1 2023OpenAI~$1bn ARR~$3bnDeeply negative$29bn
Q4 2024OpenAI~$4bn ARR~$8bnNegative$157bn
Q4 2024Anthropic~$1.3bn ARR~$3bnDeeply negative$61bn
Q4 2024xAI / Grok~$0.3bn ARR~$1bnPre-revenue scale$50bn
Q1 2026OpenAI$12.7bn ARR~$18bn~30% gross margin, ~0% EBIT$300bn
Q1 2026Anthropic~$2.5bn ARR (est)~$5bnNegative โ€” heavy training spend$61bn
Q1 2026xAI / Grok~$1bn ARR (est)~$2bnNegative$80bn
Sep 2026 (est)OpenAI~$40bn run-rate~$41bnApproaching break-even$300bn+
Sep 2026 (est)Anthropic~$4bn run-rate~$7bnStill negative$61bn
Sep 2026 (est)xAI / Grok~$2bn run-rate~$3bnNegative$80bn+
02 โ€” THE TOKEN PRICE STORY

Token Cost Collapse โ€” 99% Cheaper in Three Years

One of the most significant and underreported stories in AI is how fast the cost of using these models has fallen. In March 2023, GPT-4 cost $60 per million output tokens. By September 2026, comparable quality costs under $1 per million output tokens โ€” a fall of over 98% in three years. This is faster than Moore's Law and is reshaping who can afford to build AI products.

๐Ÿ“‹ Why prices are falling so fast: Three forces simultaneously. First, hardware efficiency โ€” each generation of GPU does more per dollar. Second, software efficiency โ€” better training methods (MoE architecture, distillation, quantization) mean smaller models do what larger ones used to. Third, competition โ€” DeepSeek, Gemini Flash, Mistral and open-source models are forcing OpenAI and Anthropic to cut prices to remain competitive. The result is a classic technology cost curve โ€” where falling prices drive volume, which maintains revenue even as per-unit cost collapses.

What a Token Actually Costs in Practice

A token is roughly 0.75 words, or 4 characters. A typical ChatGPT conversation uses approximately 2,000-5,000 tokens. A typical API call for a business task (summarise a document, analyse some data, generate a report) uses 5,000-50,000 tokens. Here is what that costs at current prices:

TaskApprox TokensGPT-5.6 Sol (premium)Claude Sonnet 5 / Gemini 3.1Grok 4 (mid)DeepSeek V4-Flash (cheapest)
Single ChatGPT conversation3,000$0.10$0.035$0.022$0.0009
Summarise a 10-page report15,000$0.50$0.17$0.11$0.005
Analyse a full contract (50 pages)75,000$2.48$0.85$0.56$0.025
Process 1,000 customer emails500,000$16.50$5.70$3.75$0.17
Full codebase review (large app)2,000,000$66.00$22.80$15.00$0.70

Prices per million tokens: GPT-5.6 Sol $5/$30, Claude Sonnet 5 $3/$15 (blended ~$3.80/M), Grok 4 $2/$10 (blended ~$2.50/M), DeepSeek V4-Flash $0.14/$0.28. Actual costs depend on input/output ratio and caching. September 2026.

03 โ€” THE DEEPSEEK QUESTION

DeepSeek โ€” Is It Really That Cheap, and What's the Catch?

DeepSeek โ€” a Chinese AI lab spun out of quantitative hedge fund High-Flyer โ€” caused a global market shock in January 2025 when it released DeepSeek R1: a reasoning model that matched GPT-4o on most benchmarks, cost $0.55 per million input tokens (vs $2.50 for GPT-4o), and was released as open source under MIT licence. NVIDIA's share price fell 17% in a single day.

By April 2026, DeepSeek released V4-Pro and V4-Flash โ€” with V4-Flash benchmarking above GPT-4o on coding tasks at $0.14/$0.28 per million tokens. That is approximately 107 times cheaper on output tokens than GPT-5.5. The question everyone is asking: how is this possible, and is the advantage real?

$0.14/M
DeepSeek V4-Flash Input Price
$0.28/M
DeepSeek V4-Flash Output Price
$5.00/M
GPT-5.6 Sol Input (36ร— more expensive)
$15.00/M
Claude Opus 5 Input (107ร— more expensive)
$3.00/M
Claude Sonnet 5 Input (21ร— more expensive)
$2.00/M
Grok 4 Input (14ร— more expensive)
$294k
DeepSeek V3 Claimed Training Cost (vs $100m+ GPT-4)
MoE
Key Architecture โ€” Only 49B of 1.6T Params Active

How DeepSeek Achieves Lower Costs โ€” The Real Explanation

Mixture of Experts (MoE) architecture: DeepSeek V4-Pro has 1.6 trillion total parameters but only activates approximately 49 billion for any given token. This means the compute cost per inference call is equivalent to a 49B model, not a 1.6T model โ€” giving frontier-scale knowledge at a fraction of the inference cost.

Training efficiency: DeepSeek claimed V3 was trained for approximately $5.6 million in compute costs โ€” compared to estimates of $100m-$500m+ for comparable Western models. Independent researchers believe this is plausible because DeepSeek used aggressive hardware optimisation, knew exactly what techniques to use from published research, and had access to cheaper Chinese compute infrastructure. However the $294,000 training cost figure cited for earlier versions is likely misleading โ€” it counts only the final training run, not the full research and development cost.

The catch โ€” data and safety: DeepSeek models are trained in China, subject to Chinese regulations. They will not discuss Tiananmen Square, Taiwan, Xinjiang or other politically sensitive topics. For applications where political neutrality matters, this is a genuine limitation. For code generation, data analysis, and most business tasks, it rarely matters.

The catch โ€” export controls: DeepSeek used NVIDIA H100 GPUs in earlier training runs. Since then, US export controls have tightened further. Whether DeepSeek can continue training frontier-scale models under current restrictions is genuinely uncertain โ€” this is the biggest risk to its long-run competitive position.

๐Ÿšจ The geopolitical dimension: DeepSeek's competitive pricing is not just a business story โ€” it is a strategic one. A capable, cheap, open-source Chinese AI model reduces dependence on US AI infrastructure globally. Countries and companies that might otherwise pay OpenAI or Anthropic API rates can self-host DeepSeek for near-zero marginal cost. This directly challenges the assumption that US labs will monetise AI globally at premium pricing. It is one reason US AI companies have accelerated their own open-source releases (Meta's Llama series, OpenAI's GPT-OSS) โ€” the competitive pressure is real.
04 โ€” WHO IS ACTUALLY USING THIS

Users, Adoption and What People Are Paying

The user figures for AI tools are staggering by any historical comparison for technology adoption โ€” but the monetisation rate per user is still low relative to the infrastructure investment being made.

200m+
ChatGPT Weekly Active Users
$20/mo
ChatGPT Plus (most common paid tier)
~5%
Estimated Conversion to Paid (ChatGPT)
$120/yr
Revenue per Paid User (Plus tier)
$12/yr
Revenue per Total User (incl free)
$120/yr
Google earns per user via AI-enhanced ads

The comparison to Google is instructive. Google earns approximately $120 per user per year through AI-enhanced advertising. OpenAI earns approximately $12 per total user per year ($120 per paying user ร— ~10% conversion). Google's monetisation model โ€” where AI makes existing search advertising more effective โ€” may prove more durable than the subscription model, where users pay explicitly for AI access.

Enterprise adoption is growing faster than consumer and commanding higher prices. ChatGPT Enterprise (custom pricing, typically $60-100 per user/month for large deployments) is driving a growing share of revenue. API usage โ€” where developers pay per token and build AI into their own products โ€” accounts for approximately 25% of OpenAI revenue but is under more competitive pressure from DeepSeek and open-source alternatives.

05 โ€” WHO LEADS RIGHT NOW

Model Capability Leaderboard โ€” September 2026

The benchmark landscape has evolved significantly. MMLU (general knowledge) is now approaching saturation โ€” top models score 88-94% and the differences are within noise. The more discriminating benchmarks are GPQA Diamond (graduate-level reasoning), SWE-Bench Verified (real software engineering tasks), and AIME 2025 (olympiad mathematics). These better separate genuine capability differences.

Key caveat: Benchmark scores do not reliably predict real-world performance on your specific task. A model scoring 94% on GPQA Diamond may underperform a 89% model on your particular use case. The only reliable test is to benchmark on your own data. These numbers tell you about the direction of capability progress, not a precise ranking for every application.
#ModelGPQA DiamondSWE-BenchMMLU-ProInput $/MReleased
1 GPT-6 Astra Closed 96.0%93%94% $10+ Aug 2026
2 Claude Opus 5 Closed 95%91%93% $15 May 2026
3 Gemini 3.1 Pro Closed 94.3%80.6%92% $2 Feb 2026
4 GPT-5.5 Closed 92%88.7%91% $5 Apr 2026
5 DeepSeek V4-Pro OpenCN 89%76.4%88% $0.45 Apr 2026
6 Grok 4 Closed 88%74%88% $2 Mar 2026
7 Claude Sonnet 5 Closed 87%72%87% $3 Mar 2026
8 Qwen3-Max OpenCN 85%70%86% $0.40 Apr 2026
9 Kimi K2 OpenCN 83%68%84% $0.60 Jun 2026
10 DeepSeek V4-Flash OpenCN 79%67%82% $0.14 Apr 2026
11 Llama 4 Scout 109B Open 76%62%80% Self-host Jan 2026

Sources: LLM Stats leaderboard, LocalAI Master, PrecisionAI Academy LLM Leaderboard, vendor model cards. Scores are rounded. GPQA Diamond measures graduate-level reasoning; SWE-Bench Verified measures real software engineering tasks on GitHub issues; MMLU-Pro measures breadth across 57 academic subjects. GPT-6 Astra available September 2026 to enterprise Trusted Access Program only. Prices as of September 2026.

๐Ÿ“‹ The capability trajectory in plain terms: Frontier model capability is advancing roughly every 6-12 months at a pace that a year ago would have seemed impossible. GPQA Diamond scores have moved from ~60% (GPT-4, 2023) to ~96% (GPT-6 Astra, 2026) โ€” approaching or exceeding human expert performance on graduate-level science. The open-source frontier (DeepSeek, Qwen, Kimi) now sits 5-10 percentage points behind the closed frontier โ€” a gap that has narrowed dramatically from ~20pp in 2024. The implication: the performance advantage of expensive closed models is shrinking while their price premium remains large. This is the core competitive pressure on OpenAI and Anthropic's pricing power.

HOW TO READ THIS PAGE

New data rows are added monthly. The signal badges at the top summarise the three original warning signals: revenue gap (is the money coming in fast enough?), data centre debt (are the lenders getting nervous?), and chipmaker capex (is TSMC being disciplined?). Sections 2-5 on token costs, DeepSeek, users and the leaderboard are updated monthly as new pricing and benchmark data is published.

The overall picture as of September 2026: AI revenues are growing fast but not yet faster than costs; token prices are falling at extraordinary speed which is good for users but compresses lab margins; DeepSeek's price advantage is real and structurally significant; model capability is advancing faster than most observers expected; and the question of long-run profitability for the labs remains genuinely open.

Disclaimer: This page is for general information and is not investment advice. Figures are estimates compiled from public reporting and may be revised. AI company financials are mostly not publicly audited โ€” treat all revenue and cost figures as estimates. Model benchmark scores reflect performance at a specific date on specific tasks and do not reliably predict real-world performance. Always do your own research or speak to a qualified financial adviser before making investment decisions.
Independent analysis ยท Not investment advice ยท Disclaimer