


Author: Li Gang, Tencent Research Institute
Recently, some media outlets reported that Microsoft has revoked internal licenses for Claude Code. Claude Code is an AI programming tool developed by Anthropic. Within just six months of its internal release at Microsoft, it became one of the most popular development assistants. However, this was accompanied by a dramatic surge in token consumption and skyrocketing costs, with output quality failing to meet expectations. After weighing multiple factors, Microsoft hit the brakes, steering employees toward its own Copilot CLI.
The phenomenon where token consumption is disproportionate to actual output is also prevalent among other platform companies. Uber exhausted its entire 2026 AI programming tool budget in just four months; some Amazon employees consumed tokens without meaningful results; Meta quietly removed its internal Tokenmaxxing leaderboard, no longer encouraging token consumption without output. Everyone is embracing AI, but nobody has found the right approach yet; companies are all emphasizing being "AI-native," but (for now) they see no returns, only increasingly long bills. I call this "token inefficiency."
Token inefficiency is the result of a combination of multiple factors: poor internal corporate governance, limited returns on token usage, and inherent issues in Agent architecture design (such as repetitive skill calls, internal friction in long-horizon tasks, and the coordination costs of multi-agent systems). In the future, these issues may gradually ease with tighter internal controls and ongoing technological optimization of consumption. However, turning the net benefit of tokens positive will require not only optimizing token costs from the supply side but also solving the problem of how to generate real value from token consumption across a wide range of industry scenarios from the demand side.
Over the past two years, as mainstream large language models have rapidly iterated, development companies have adopted different product portfolio strategies based on their market positioning, leading to changes in API call prices ($ per million tokens). While model performance has improved significantly, good things aren't cheap—the call prices for similar-tier products have also quietly increased, becoming a major driver of higher token consumption costs for downstream users.
(I) The Tiered Strategy of the Leader
Anthropic was the first proprietary model provider to recognize that coding is the core monetization scenario for tokens. The primary paying users of large models are developers and enterprise technical teams, who are less price-sensitive and more focused on coding efficiency and quality. By seizing the early-mover advantage in the coding business scenario, Anthropic could achieve token premium pricing. Therefore, its R&D efforts have been focused on coding. After establishing its coding capability advantage, starting with the launch of the Claude 3 series in early 2024, it pioneered a flagship-mid-range-lightweight product portfolio within the industry, achieving tiered pricing for same-generation models while targeting both high-end and mass markets.
1. The Opus series positioned as the coding benchmark, anchoring the high-end market at $15/$75 (input/output price per million tokens, same below);
2. The Sonnet series ($3/$15) offers a high cost-performance ratio for daily coding and office tasks;
3. The Haiku series ($1/$5) targets lightweight, fast interactive scenarios with an affordable price point.
This fine-grained tiering allows Anthropic to maximize profit extraction at every price point while protecting market share. This pricing strategy gives the technology leader Anthropic more competitive leeway and operational flexibility.
For example, after sensing that the performance gap with competitors was rapidly closing, it significantly cut prices with the release of Opus 4.5 to squeeze competitors' market space.
Another example: with the release of the next-generation model Mythos Preview ($25/$125), a new ultra-high-end tier was placed above Opus, raising the price of the flagship product and reversing the previous trend of price declines for high-end products. The subsequently released Fable 5, which shares the same underlying architecture, has some features restricted for safety reasons and is priced at $10/$50 (still double the Opus series) for a broader market.
This creates a three-dimensional pricing strategy: pricing not only by performance but also by the degree of safety constraints, forming tiers of capability, risk, and price, thereby reclaiming the premium market. The effectiveness of this positioning strategy was fully validated between 2025 and 2026. Anthropic's Annual Recurring Revenue (ARR) surged from approximately $1 billion at the end of 2024 to around $450 billion by May 2026.
More importantly, this strategy fully protects the market premium of the product leader, leveraging performance advantages to escape the trap of price wars and complete the value cycle where "good things aren't cheap."
(II) The Price Game of the Chasers
In contrast, OpenAI and Google chose a more diversified path early in the commercialization of large models, differing from Anthropic.
1. In 2024, OpenAI invested significant resources in multimodal projects like Sora;
2. Google built an ecosystem strategy around Gemini covering multiple product lines such as Search, Cloud, and Workspace.
While these investments expanded their technological footprint, the dispersion of resources meant their performance in office and coding scenarios was relatively less prominent. By the time they realized coding was the primary battleground for monetizing model capabilities and turned back to catch up, they had lost their first-mover advantage. OpenAI's pivot was very decisive.
1. On one hand, it refocused on coding and Agent capabilities, cutting resource-intensive projects like Sora;
2. On the other hand, it followed Anthropic in establishing its own tiered product matrix for close one-on-one competition, while deliberately widening the price gap between flagship and lightweight models—high prices for flagships to maintain the leading model brand, low prices for lightweight models to grab market share.
GPT 5.5's pricing ($5/$30) aligns with Opus 4.7/4.8 ($5/$25), establishing an equivalent high-end price anchor as Claude Opus. Sub-tier models like GPT 5.4 mini ($0.75/$4.50) and nano ($0.20/$1.25) are priced significantly lower than their counterpart Claude Haiku 4.5 ($1.00/$5.00), trading price for market share.
Google sits at the core of the Android ecosystem, already having a complete business loop. It needs to manage more complex relationships and acts more cautiously. Gemini must simultaneously serve Google Cloud's enterprise customers, Workspace's productivity users, and the consumer experience of Search products. Even recognizing the importance of coding, it cannot decisively focus all resources on coding and office tasks and must still pursue a multimodal, diversified path. Google also followed Anthropic, starting from the Gemini 1.5 generation, dividing its products into flagship Pro series and lightweight Flash series, but with a slower iteration speed and lower price positioning.
1. In early 2024, the flagship Gemini 1.5 Pro priced output tokens at just $5 per million for short prompts (<128k), one-third the cost of the contemporary GPT-4o and one-fifteenth of Opus 3;
2. By February 2026, the output token price for Gemini 3.1 Pro had risen to $12 per million, significantly lower than GPT 5.4's $15 and Opus 4.6/4.7's $25.
Furthermore, Google took a reverse approach by adding an ultra-lightweight product line, Flash-Lite, under its lightweight Flash series, pushing call prices down to the same level as open-source models—a classic trade of price for volume. The delayed official release of the eagerly awaited Gemini 3.5 Pro also reflects Google's internal struggles in balancing performance, safety, and ecosystem compatibility. The pricing strategy for the next-generation flagship model is also under intense market scrutiny.Figure 1: Pricing Trends for Flagship Models Pricing for Claude series, GPT-4o/4.1/5.4 from official pages; GPT-5.5 series, Gemini 3.5 Flash from OpenAI/Google platforms and third-party aggregators; GLM series based on Z.ai platform overseas, subject to exchange rates and dual-track pricing. Charting: Codebuddy
(III) Sub-tier/Lightweight and Open-source/Semi-open-source Models Quietly Raise Prices Amidst Surging Demand
Flagship models compete on performance, while sub-tier/lightweight models compete on price—this is the expected norm for market competition. Faced with intense competition, the general expectation is a continuous decline in the market price midpoint. However, the reality is quite the opposite: the price midpoint for the economy token market, composed of sub-tier/lightweight and open-source/semi-open-source models, has quietly shifted upward over the past two years, and it is in this upward movement that the true floor price of the token market has been raised. On the surface, this appears to be a brutal red ocean.
Inexpensive sub-tier/lightweight models like Sonnet, mini, and Flash are the affordable options from mainstream proprietary model providers targeting the mass market, primarily aiming to capture market share.
Simultaneously, open-source or semi-open-source models like DeepSeek, Qwen, and GLM have rapidly risen, generally adopting a strategy of flagship positioning at sub-tier/lightweight prices, continuously exerting pricing pressure on the sub-tier/lightweight proprietary model market. By the end of 2024, DeepSeek V3 entered the market at approximately $0.27/$1.10, far below comparable proprietary models. The later R1 offered reasoning-enhanced capabilities at $0.55/$2.19, directly compressing the pricing headroom for GPT-4.1 mini and Claude Haiku. GLM-4 Plus offers capabilities close to GPT-4 at just $0.69/$0.35, strongly attracting price-sensitive developers.
Competing on price seems normal for this tier. However, on the other hand, the launch of each new generation of sub-tier/lightweight and open-source/semi-open-source models has been accompanied by a rise in the price floor.
1. For example, Haiku 3.5, launched in October 2024, was priced at $0.80/$4.00 for input/output;
2. A year later, Haiku 4.5 saw a 20% price increase to $1.00/$5.00.