Futures
Access hundreds of perpetual contracts
CFD
Gold
One platform for global traditional assets
Options
Hot
Trade European-style vanilla options
Unified Account
Maximize your capital efficiency
Demo Trading
Introduction to Futures Trading
Learn the basics of futures trading
Futures Events
Join events to earn rewards
Demo Trading
Use virtual funds to practice risk-free trading
CFD
Stock CFD Derivatives
US Stocks
Access real US stocks and ETFs
HK Stocks
Trade quality Hong Kong-listed stocks
Korean Stocks
SK Hynix
Real Korean stocks and top assets
Stock Futures
High leverage, 24/7 trading
Tokenized Stocks
Backed by real stock assets
IPO Access
Unlock full access to global stock IPOs
GUSD
3.8%
Mint GUSD for Treasury RWA yields
Stocks Activities
Trade Popular Stocks and Unlock Generous Airdrops
Launch
CandyDrop
Collect candies to earn airdrops
Launchpool
Quick staking, earn potential new tokens
HODLer Airdrop
Hold GT and get massive airdrops for free
Pre-IPOs
Unlock full access to global stock IPOs
Alpha Points
Trade on-chain assets and earn airdrops
Futures Points
Earn futures points and claim airdrop rewards
Promotions
AI
Gate AI
Your all-in-one conversational AI partner
Gate AI Bot
Use Gate AI directly in your social App
GateClaw
Gate Blue Lobster, ready to go
Gate for AI Agent
AI infrastructure, Gate MCP, Skills, and CLI
Gate Skills Hub
10K+ Skills
From office tasks to trading, the all-in-one skill hub makes AI even more useful.
The cheaper the Kimi K3 chips are, the more scarce they are? Efficiency surges, and hardware demand surges
Kimi K3’s price cut will not reduce chip demand; instead, it may actually drive higher overall hardware consumption.
(Background: the local hardware threshold for running Kimi K3? Starting from 64 GPUs, 45kW of power—players setting up their own rigs are out of luck.)
(Additional context: on the transaction scale, it’s $10 billion! Reportedly, Anthropic is considering leasing compute resources from Meta to ease the GPU shortage crisis.)
Table of contents
Toggle
Kimi K3 is bringing AI investors back to a core question: if models get cheaper and more efficient, does that mean chip and memory demand will fall? The answers from Citi and Bank of America Securities both lean toward “no”—improved efficiency may unlock more usage, which in turn increases total resource consumption.
According to a “follow-the-market” trading desk, the Kimi K3 released by Moonshot has 2.8 trillion parameters, targeting a 1 million-token context window and long-cycle agent task workloads. Citi semiconductor analyst Peter Lee believes that even if Kimi K3 is widely used, general memory requirements such as server DDR5 and eSSD will still rise.
In research dated July 17, Bank of America Securities semiconductor analyst Vivek Arya also said that U.S. leading frontier AI labs are more likely to increase compute investment rather than reduce it. If Chinese open-source models continue to close the gap, leaders such as OpenAI, Anthropic, and Google will need larger training runs, heavier inference, and faster product iteration to maintain differentiation.
This means the market’s pricing logic for Kimi K3 cannot be judged only by the lower cost of a single inference. More importantly, does a lower-priced model lead to more calls, longer task chains, and more token generation? If the answer is yes, GPUs, HBM, DDR5, eSSD, high-speed networking, and inference systems may all continue to benefit.
Kimi K3 debuts: efficiency surge rebounds to boost hardware demand
Citi views Kimi K3 as a potential example of “Jensen’s paradox” in the AI industry chain. The core of this paradox is that when technological efficiency reduces the unit cost of use, usage may increase dramatically—ultimately resulting in total resource consumption not falling, but rising.
Kimi K3’s appeal lies in the combination of low cost and high performance. According to publicly listed pricing, the input price for cached hits is $0.3 per million tokens, while the output price is $15 per million tokens. Its full weights plan will be disclosed on July 27, with the goal of supporting long-cycle agent work.
Citi’s assessment is that lowering the model’s price will not automatically reduce hardware demand. On the contrary, lower costs may increase developers’ and enterprises’ willingness to call the model, driving more AI agent deployments. Agent tasks are not one-time Q&A; they continuously generate tokens during ongoing tasks, including reading and processing. As call counts and task lengths rise, the reduced unit cost is converted back into higher total resource consumption.
Therefore, Kimi K3’s impact on the semiconductor supply chain is not primarily about whether “each call is cheaper,” but about whether “the total token volume increases.” This is exactly why Citi is bullish on demand for server DDR5 and eSSD.
Jensen’s paradox: low pricing triggers higher token consumption
Kimi K3 adopts three key technologies: Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. Kimi Delta Attention is used to reduce the cost of a 1 million-token context window. Attention Residuals is used to selectively retrieve representations between different depths of the model. Stable LatentMoE increases sparsity, activating only 16 of 896 experts per token.
Moonshot’s technical materials state that this architecture enables Kimi K3’s expansion efficiency to reach 2.5 times that of K2. High sparsity and long-context capability are major foundations for lowering execution costs.
However, Citi emphasizes that this does not mean inference-side resource pressure disappears. Kimi K3 is still not a lightweight deployment solution; its execution requires multi-node clusters, with super-node configurations of more than 64 GPUs. More importantly, long-context and agent task workloads will increase KV Cache usage, thereby raising the memory burden on the inference side.
KV Cache requirements are directly related to server DDR5 and eSSD. DDR5 handles high-frequency data access, while eSSD benefits from larger cache and data storage needs. For memory vendors, the key variable is not whether a single model uses less compute, but how many times the model is called after prices drop, how many tokens are generated per call, and how long the agent task chain becomes.
KV cache skyrockets: long context squeezes memory
Bank of America Securities’ conclusion echoes Citi’s, but its focus is more on GPUs and AI infrastructure. Vivek Arya believes that stronger Chinese open-source models do not necessarily mean U.S. AI giants will cut compute investment. Instead, as the model gap narrows, the cost for leaders to maintain their advantage may increase.
Bank of America Securities states that if open-source models continue to close in, OpenAI, Anthropic, and Google must rely on larger-scale training, more reinforcement learning and synthetic data cycles, heavier test-time inference, and faster product release cadences to hold onto differentiation.
This logic does not hinge on which model leads in the short term. Big-model rankings may rotate quickly, but what enterprises truly purchase is stable, low-latency, high-availability AI output—and lower “cost per unit of effective output.” As model capabilities converge, competitive pressure will flow down to the underlying infrastructure, including GPUs, HBM, high-speed networking, and inference systems.
Therefore, Bank of America Securities believes that the emergence of models like Kimi K3 should not be simply understood as “models becoming more efficient and requiring fewer chips.” A more realistic path is that competition among models intensifies, and leaders continue adding compute to keep their distance.
Competitive chill: U.S. AI giants continue to ramp up
Kimi K3 uses a MoE architecture, meaning the total parameter count and the parameters actually activated per inference can be separated. This change reduces some computation pressure, but also shifts the infrastructure bottleneck away from pure compute and toward video memory access, expert routing, response latency, and interconnect capability.
Bank of America Securities emphasizes that MoE inference requires the system to move data more efficiently, schedule expert modules, and connect to larger compute clusters. Nvidia has proposed that modern MoE inference requires a larger GPU domain. Its calculations show that, in terms of performance per watt on leading open-source models, GB300 NVL72 can reach up to 25 times that of the Hopper platform.
CoreWeave’s tests around Kimi K2.6 also illustrate that open-source MoE models still need highly optimized Nvidia GB300/GB200 NVL72 infrastructure to maintain leadership in speed and cost-effectiveness, even if only part of the parameters are activated per call.
This means model weights may gradually become commoditized, but the GPUs, high-bandwidth memory, network interconnects, and inference systems required to run the models will not lose value. Instead, as model call volume grows, these areas may become increasingly concentrated battlegrounds for competition.
Open-source buying spree: OpenRouter data shows growth
Bank of America Securities also cites OpenRouter data showing that model call volume continues to rise rapidly. As a third-party API platform that connects multiple model labs, token usage on OpenRouter keeps increasing, and the token usage for Chinese AI lab models has already surpassed that of non-Chinese lab models.
This data cannot directly represent all enterprise and consumer scenarios, but it at least shows that developers’ choices within open model ecosystems are changing. As model prices fall and capabilities improve, call volume may keep expanding rather than being offset by the efficiency gains of a single model.
Paid usage by enterprises is also increasing. Ramp data shows that as of June 2026, about 55% of U.S. enterprises have paid subscriptions to AI models—platforms or tools—higher than the 21% estimate from the U.S. Census Bureau’s BTOS survey. By model, Anthropic’s enterprise adoption rate is 42.4%, and OpenAI’s is 39.5%.
Conclusion: Kimi K3 as a hardware catalyst
However, AI spending remains highly concentrated. The top 1% of corporate users spend about $4,833 per employee per month on AI; the top 10% spend $516, and the overall median is only $11. Subscription rates for the technology and media industry reach 79.8%, and adoption by large enterprises is 65.5%, higher than 61.3% for mid-sized enterprises and 48.7% for small enterprises.
This set of data supports one assessment: AI usage is still in the diffusion process. If low-priced models lower the entry barrier, future incremental demand may come from more enterprises, more developers, and more agent applications.
Citi and Bank of America Securities share the conclusion that the significance of Kimi K3 is not whether a single model reduces unit inference costs, but whether lower costs unlock a larger-scale workload.
For Citi, the most direct beneficiary direction is server DDR5 and eSSD. Long context, KV Cache, and agent tasks will increase memory access and storage needs. For Bank of America Securities, even more significant are GPUs, HBM, high-speed networking, and inference systems, because model competition will force leaders to keep investing in infrastructure.
Bank of America Securities also mentioned that Kimi K3’s involvement with “45-nanometer open-source EDA” should not be interpreted simply as commercial EDA being replaced. Instead, it indicates that chip design still cannot do without EDA tools. At advanced process nodes, commercial EDA vendors such as Cadence and Synopsys remain in key positions.
The risk is another scenario: if models compress and inference optimization and hardware efficiency improvements outpace the growth of new workloads, AI infrastructure expansion may cool down in phases. But under the frameworks of these two investment banks, Kimi K3 looks more like a catalyst that pushes usage higher rather than a signal that weakens chip demand. After efficiency improvements, what the market should focus on is instead more tokens, more inference, and more consumption of underlying hardware.