Futures
Access hundreds of perpetual contracts
CFD
Gold
One platform for global traditional assets
Options
Hot
Trade European-style vanilla options
Unified Account
Maximize your capital efficiency
Demo Trading
Introduction to Futures Trading
Learn the basics of futures trading
Futures Events
Join events to earn rewards
Demo Trading
Use virtual funds to practice risk-free trading
CFD
Stock CFD Derivatives
US Stocks
Access real US stocks and ETFs
HK Stocks
Trade quality Hong Kong-listed stocks
Korean Stocks
SK Hynix
Real Korean stocks and top assets
Stock Futures
High leverage, 24/7 trading
Tokenized Stocks
Backed by real stock assets
IPO Access
Unlock full access to global stock IPOs
GUSD
3.8%
Mint GUSD for Treasury RWA yields
Stocks Activities
Trade Popular Stocks and Unlock Generous Airdrops
Launch
CandyDrop
Collect candies to earn airdrops
Launchpool
Quick staking, earn potential new tokens
HODLer Airdrop
Hold GT and get massive airdrops for free
IPO Access
Unlock full access to global stock IPOs
Alpha Points
Trade on-chain assets and earn airdrops
Futures Points
Earn futures points and claim airdrop rewards
Promotions
AI
Gate AI
Your all-in-one conversational AI partner
Gate AI Bot
Use Gate AI directly in your social App
GateClaw
Gate Blue Lobster, ready to go
Gate for AI Agent
AI infrastructure, Gate MCP, Skills, and CLI
Gate Skills Hub
10K+ Skills
From office tasks to trading, the all-in-one skill hub makes AI even more useful.
After talking with 15 Silicon Valley AI companies, I’ve come to a new understanding of the open-source vs. closed-source debate.
Author: Chenzi Tina; Source: Silicon Valley Non-consensus
From Token Optimization to Model Routing, my biggest few takeaways from the past month.
First, three conclusions:
The hottest term in Silicon Valley has shifted from Token Maxing to Value Maxing: enterprises have started systematically managing AI costs for the first time.
The token share of open-source models is rising faster than most people expect: on a leading Coding Agent platform, it went from 0.1% to 15% in 7 months. But token share and revenue share are two different things.
About a month ago, Gavin Baker provided a framework: frontier models capture 90% of the value, while open-source models carry 80% of the tokens. Testing with the latest data, the realization of that half of tokens is happening faster than expected: in early June, the token volume of China models on OpenRouter already surpassed that of US models.
Yesterday, Google released its Q2 earnings report, which is a ready-made footnote for this article—Google Cloud revenue grew year over year by 82% to $24.8 billion, and backlog rose to $51.4 billion. Even more noteworthy is this figure: throughput of Google’s first-party model APIs has reached about 22 billion tokens per minute, up from 16 billion in the previous quarter.
After half a year of talking about Token Optimization, the total token volume is still accelerating—that’s the real-time demonstration of Jevons Paradox that will be discussed later. But after-hours market reaction was a drop: capital expenditures doubled to $44.9 billion, and for the first time since listing, quarterly free cash flow turned negative. Growth isn’t in doubt, but the market has started pricing the “bill.”
Last week in Menlo Park, I attended a closed-door meeting and talked with a dozen-plus AI Native companies—Cursor, Fireworks, Cognition, Sierra, and others—over two days. Plus, over this period, at Stanford and in repeated one-on-one interviews with founders, we kept circling the same questions; I’m increasingly certain of one thing:
AI competition is shifting from Capability to Economics.
For the past two years, the question has been who can train the strongest model; today, enterprises are asking: which tasks are worth calling the strongest models, and which tasks are fine with cheaper ones.
The two corresponding keywords are Token Optimization and Model Routing. The former means enterprises have started systematically managing AI costs for the first time; the latter means enterprises no longer rely on a single model—instead, they dynamically allocate traffic among closed-source, open-source, and even self-developed models based on the task.
Silicon Valley’s discussion of open-source models has moved from “is it smart enough” to “can it handle more production traffic.” And behind this, what truly matters is the revenue growth curve over the next few years for Frontier Labs like OpenAI and Anthropic.
Growth hasn’t slowed, but the logic has changed
Let me correct an easily misunderstood take: enterprises optimizing tokens doesn’t mean AI investment is cooling down.
On the contrary, the growth rate of these AI Native companies is still astonishing. Three examples are enough: Fireworks’ ARR was about $400 million at the start of the year; it has now surpassed $1 billion, and it just closed a round at a valuation of $17.5 billion. Cognition is about $500 million, with year-over-year growth approaching seven times. Higgsfield in video generation: revenue was zero 14 months ago; now it’s $500 million.
There are three reasons behind the support for growth: model capability has made a clear jump in the past half year, and enterprises truly see ROI; the product has evolved from Copilot to Agent, and tokens consumed per task are significantly higher; more and more companies have moved to a subscription + usage hybrid pricing model, and token consumption directly converts into revenue.
AI spending hasn’t slowed—it has entered a phase of more fine-grained operations: enterprises are still willing to spend, but they’re no longer willing to pay for low-value tokens.
Token Optimization: every customer conversation is about cost
Token Optimization is the hottest topic these past two days. Cursor put it plainly: now, every customer conversation is about this.
Enterprises are starting to redesign workflows—analyzing which tasks truly require the strongest models, and which tasks can be pushed down to cheaper models.
But not everyone is hitting the brakes. ClickHouse is an interesting counterexample: they pay Claude $20,000 per month in January this year, rising to $1 million per month in June—up 50x, equivalent to 5% of their own revenue. Even so, the company still encourages employees to use it freely, with an expectation to keep rising to 10%, because they believe the productivity gains from AI far outweigh that cost.
That said, across the board, ClickHouse is a minority. Most enterprises are already installing brake systems. Some companies attribute the spark of this optimization wave to software vendors collectively switching to token billing—for example, Microsoft on June 1 changed GitHub Copilot to token-based pricing. The moment cost becomes clearly visible, control begins.
This reminds me of Cloud Optimization from 2022 to 2024. Back then, enterprises didn’t give up on cloud; they were optimizing resource utilization. But it’s worth remembering that the bulls at the time also said, “once optimized, we’ll spend more.” The result was AWS, Snowflake, Datadog, and others experiencing four consecutive quarters of slowing growth. Today, the thing being optimized has shifted from cloud resources to tokens—will the script replay? One of the most worth watching questions in the second half of the year.
Model Routing: enterprises aren’t buying a single model anymore—they’re buying a system
If Token Optimization is the goal, Model Routing is the method.
We’ve long passed the stage of “give users a dropdown menu of models.” Cognition can now split a single inference task into multiple subtasks, automatically assigning them to different models based on cost, reasoning ability, latency, and accuracy. Users don’t even need to know who is being called behind the scenes.
What enterprises are actually procuring isn’t just one model, but a system that can intelligently schedule multiple models.
The question last year was: Who builds the best model?
The question this year is: Who routes to the best model?
How fast is this migration? Two data points: customer AI company Decagon disclosed that 30% of token requests are already routed to open-source models. On the Factory platform, open-source model token share was only 0.1% in December last year, 1% in February, and now is 10% to 15%—up over a hundredfold in 7 months. Even their founder believes that in 12 months, it could reach 90%.
I’m skeptical about the 90%, but the direction and slope are clear. Also worth noting: routing to open-source models isn’t only about being cheaper. Several companies mentioned another motivation: they don’t want to be locked in by any single model vendor.
The dumbbell structure of the model market: either Great or Cheap
After talking with these companies, my view on the model landscape can be condensed into one word: a dumbbell.
The model market will ultimately form a dumbbell structure.
Either Great or Cheap.
The middle ground will become increasingly hard to survive.
Both ends will undergo consolidation. On the cheap end, new faces keep appearing over the past half year—Kimi K3, GLM 5.2, Qwen taking turns for the spotlight—but the “floor price” is still pinned by DeepSeek: it remains the largest single open-model supplier on OpenRouter, exceeding any US vendor. Some investors call it the market’s “price-performance floor.”
DeepSeek sets a floor for value-for-money across the market: any model priced higher than it must be significantly better, otherwise users would rather use DeepSeek and try multiple times. If a model can only tie it, or just be slightly better, it will be hard to live.
And on the strong end, the real question is whether cheaper models can cross the “good enough” line.
The bear case is: once cheap models are good enough, most users will be satisfied, and the premium commanded by Frontier Labs will collapse.
I disagree. I believe demand for frontier intelligence will never be fully satisfied—every time model capability steps up a level, it unlocks a batch of tasks that previously couldn’t be done. Both ends of the dumbbell will have durable demand.
A month ago, Gavin Baker gave a framework: frontier takes 90% of the value, open-source carries 80% of the tokens. At the time, I wondered whether this judgment would become outdated quickly—after all, Kimi K3 and DeepSeek V4 are later developments. Testing with the latest data, the half about tokens realizing value is happening faster than he said; the half about value is still standing for now.
OpenRouter routing data shows that token volume from China models officially surpassed US models by early June; by mid-July, Asian models made up about 60% of all platform tokens, and their share has tripled since the start of the year.
Looking over a longer one-year window, it’s even more striking: the combined token share of Google, OpenAI, and Anthropic fell from about 70% in June last year to about 30% in June this year. DeepSeek alone is above 16%, already the number one supplier on the platform; after V4’s release, its share doubled within half a year, and agentic workloads are the main driver.
But the other half of the framework also holds: Anthropic accounts for only about 12–13% of tokens, yet its revenue capture is far beyond what a pure per-usage computation would suggest—within programming-related spending, the Claude series has long held more than 60%. The market is splitting into two lanes: the commodity lane runs volume, and the premium lane charges money.
Of course, there should be some discount: OpenRouter’s sample is biased toward developers and price-sensitive traffic, and can’t fully represent the enterprise market. But as a directional indicator, it’s clear enough.
A concrete example is Harvey. Last week I chatted with their technical lead: they use proprietary legal data, do RL and SFT on open-source models in cooperation with third parties, and add a router—resulting in better performance and lower cost than using Opus 4.8 directly. This could be what the future looks like: frontier models are still used heavily, but more and more token traffic is going to task-specific open-source models.
Price tags are losing their reference value: what should really be calculated is Fully Loaded Cost
Another change worth watching is that model price tags (Input/Output Token Price) are becoming less and less meaningful as reference. What enterprises truly care about isn’t the price per million tokens, but the total cost to complete the same task.
Cache hit rate, inference stability, retry count, and token efficiency all affect the final cost. Sometimes, a model that looks four times cheaper can save much less in real production.
The frontline feedback from Silicon Valley also confirms this. If a cheap model has a low first-pass rate and needs 2–3 more rounds of retries, the token price gap is quickly eaten up; retries also generate cached tokens that further erode the cost advantage on paper. DeepSeek pushes the floor price the lowest, relying on extremely low prices; but what frontline users complain about the most is exactly output instability and inconsistent results for the same task across multiple runs. An even cheaper but more failure-prone model may end up not being cheap in production.
So the one line I want to leave here is: what you should compare isn’t the per-million-token price tag, but the fully loaded cost per successful task—the complete cost of each successful task.
A similar arithmetic applies to the on-prem versus cloud debate. Someone on Twitter has seriously run the numbers: GLM 5.2 API price is $1.40 per million input tokens and $4.40 per million output tokens. If you buy hardware and run locally, the threshold is about $20,000, and output is about 20 tok/s. Even if you run at full capacity 24 hours a day for a full year, it takes about 5.5 years to break even. By total cost, the cloud’s token efficiency still wins handily. In a world where compute is scarce and hardware iterates so quickly, this calculation will only tilt further toward cloud.
Finally, there’s availability—which is a constraint I believe is underestimated. Models like GLM are currently severely constrained by chips, mainly provided through inference platforms like Together and Fireworks, and have not been fully rolled out across major public clouds (though AWS and Azure have already started putting up some China models). Even Kimi once paused new user subscriptions due to insufficient compute. The threat from China models is worth tracking, but it’s too early to place heavy bets.
Next, I’ll keep tracking with two questions:
When can China models achieve large-scale, stable supply on public clouds?
Will their gap versus OpenAI and Anthropic converge or stabilize?
The answers to these two questions determine the real impact, not the emotions on Twitter.
What does this mean for investors?
Frontier Labs
What truly needs rethinking isn’t whether OpenAI and Anthropic will be replaced by open source; it’s whether the market needs to re-evaluate their ARR growth assumptions. A blunt take: in Anthropic’s high growth in the first half, some portion came from “getting odd jobs done with the most expensive models”—running non-frontier tasks on Opus-class models.
In the second half, those tasks will either be cut by Token Optimization or routed to cheaper models. Expensive models will still handle the highest-value tasks, but revenue growth will depend more on high-value reasoning rather than synchronous expansion of all tokens.
Semiconductors
I’m more optimistic here. People worried that Token Optimization will suppress GPU demand ignore Jevons Paradox: unit cost declines never reduce total consumption; they only expand total demand—this is a pattern the tech industry has consistently followed. Cheaper models will take AI from early adopters to the mainstream. And don’t forget: China has insufficient compute—these open-source models ultimately need to be hosted in places with abundant compute.
Industry focus will gradually shift from Training to Inference, but GPUs, HBM, high-speed networks, and advanced packaging remain the most deterministic directions within the entire AI infrastructure stack. One detail: giant open-source models like Kimi K3 can only fit on racks with the highest HBM density—so the larger the open-source models, the happier storage vendors will be.
Cloud Infrastructure
In cloud, the future competition won’t be just “how many GPUs you have,” but “who can help customers complete inference at lower cost.” Inference efficiency, model deployment capability, scheduling capability, and TCO will become increasingly important. The deployment of China models is actually increasing—not reducing—the demand for compute in the cloud. Even Databricks has publicly confirmed this.
AI Software
The biggest opportunity is shifting from the model layer to the application layer and the scheduling layer. Models are gradually becoming commodity, and platforms that can help enterprises automatically pick models, optimize costs, and improve ROI—Routing, Inference Infra, Agent orchestration—see rising value.
This is also why companies like Fireworks are increasingly worth paying attention to: they’re not competing on models, but on the efficiency of enterprise AI systems. The risk factors are also clear: if Frontier Labs faces a more competitive model market and chooses to accelerate vertical integration into the software layer, Infra software’s days won’t be good (there’s a long article later analyzing competition and opportunities at the Infra software layer).
Written at the end
Bear case, two points: cheap models cross the “good enough” line, diverting most tasks; and Token Optimization replays the 2022–2024 Cloud Optimization script—back then, the bulls also said “we’ll spend more after optimization,” and the result was a slowdown for four straight quarters.
Bull case, two points: Jevons Paradox—cheaper means a larger market. As Factory’s founder puts it, “this cake has rocket boosters strapped all around”; and China lacks compute—these open-source models are constrained by chip supply, so in the short term they can’t carry the traffic they want.
I’m personally more Bull-leaning. But views need to be falsifiable, so over the next two quarters I’ll focus on four things:
Revenue signals from Anthropic and OpenAI. If the first-half growth really has an “inflated by ‘work with the most expensive models’” component, the impact should first show up in the sequential change curve. Right now on X, third-party data providers’ estimates are already spreading—Anthropic’s ARR is still at innovation highs, but the growth-rate slope is starting to flatten; in the same tracking set, OpenAI is accelerating instead. The debate over this curve is arriving faster than I expected.
The availability of China models. Whether they can move from Together and Fireworks to large-scale, stable supply on public clouds—this is the prerequisite for moving from token share to revenue share.
The direction of the quality gap. Compared with benchmarks, I trust builders’ retry rate more.
Will fully loaded cost become procurement language? If so, many “cheap by 5x” narratives based on price tags need to be recalculated.
AI competition isn’t over—it’s just the function that’s changing.
What we optimized in the past was Intelligence; what we’ll optimize in the future is Intelligence per Dollar.
Come back in 12 months to see the answer.
— Chenzi Tina
Written the night Google released its financial report, Palo Alto