Futures
Access hundreds of perpetual contracts
CFD
Gold
One platform for global traditional assets
Options
Hot
Trade European-style vanilla options
Unified Account
Maximize your capital efficiency
Demo Trading
Introduction to Futures Trading
Learn the basics of futures trading
Futures Events
Join events to earn rewards
Demo Trading
Use virtual funds to practice risk-free trading
CFD
Stock CFD Derivatives
US Stocks
Access real US stocks and ETFs
HK Stocks
Trade quality Hong Kong-listed stocks
Korean Stocks
SK Hynix
Real Korean stocks and top assets
Stock Futures
High leverage, 24/7 trading
Tokenized Stocks
Backed by real stock assets
IPO Access
Unlock full access to global stock IPOs
GUSD
3.8%
Mint GUSD for Treasury RWA yields
Stocks Activities
Trade Popular Stocks and Unlock Generous Airdrops
Launch
CandyDrop
Collect candies to earn airdrops
Launchpool
Quick staking, earn potential new tokens
HODLer Airdrop
Hold GT and get massive airdrops for free
Pre-IPOs
Unlock full access to global stock IPOs
Alpha Points
Trade on-chain assets and earn airdrops
Futures Points
Earn futures points and claim airdrop rewards
Promotions
AI
Gate AI
Your all-in-one conversational AI partner
Gate AI Bot
Use Gate AI directly in your social App
GateClaw
Gate Blue Lobster, ready to go
Gate for AI Agent
AI infrastructure, Gate MCP, Skills, and CLI
Gate Skills Hub
10K+ Skills
From office tasks to trading, the all-in-one skill hub makes AI even more useful.
Is self-custody AI models cheaper? Cline’s tests show Kimi K3 has a cost advantage; netizens push back: it’s more expensive than Fable 5 when priced by task
As the compute costs of AI models continue to rise, whether companies should choose “self-hosting” is becoming a focus. AI development team Cline yesterday (20th) released the latest hands-on test, saying that self-hosting the open-weight model Kimi K3 can save up to 25% in costs, and predicting it will become a standard approach for companies to balance pricing and data sovereignty. However, the community has raised different views, arguing that if you calculate by “cost per task,” Kimi is actually more expensive than its competitor Fable 5.
(Background: Hugging Face got smashed by AI agents! In the end, the forensic model relied on China’s open-source GLM 5.2)
(Background addition: EU clamps down in August: AI that doesn’t “self-report identity” faces up to a 3% fine of global revenue; Taiwan is also within scope)
Table of contents
Toggle
In the arms race of large language models (LLMs), besides the performance of the model itself, the API call cost has also become a major consideration for enterprises adopting AI. On July 20, 2026 (Taipei time), AI development team Cline published attention-grabbing real-world test data on the community platform X, examining how much money enterprises could save by using a self-hosted open-weight model versus directly using a commercial API.
Self-hosting Kimi K3 hands-on test: up to 25% cost savings
In their post, the Cline team said that based on the pricing of input and output tokens, Kimi’s costs are already about 3 to 12 times cheaper than competitor Fable. But what they were more curious about was this: if a company chooses to self-host the model, how much further could it cut costs?
To do so, Cline took the real traffic from its live production environment, simulated it on a server setup with 500k200 GPU nodes, and conducted the test while carefully measuring key metrics such as time to first token (TTFT), per-stream speed, and cache hit rate. The results showed that adopting a self-hosted model can bring about 10% cost savings; if you pair it with “time-of-day autoscaling,” the savings can reach 25% or more. However, Cline also specifically reminded that optimizations at this level are generally only applicable to large enterprises spending more than $500k per year on AI compute.
Balancing price and data sovereignty—will self-built nodes become standard?
Based on this hands-on data, Cline made bold predictions about the future of AI infrastructure. They believe that as enterprise adoption of AI continues to rise and token consumption expands rapidly, “self-hosting open-weight models” will inevitably become the industry standard.
The team emphasized that this is not just to gain an advantage on price. More importantly, enterprises can keep sensitive data within their own servers, thereby gaining full “data sovereignty.” Cline said that the release of the Kimi K3 model takes this future vision one step further, giving the market more broadly available and more competitive options.
Community counters: if you look at “cost per task,” Kimi is actually more expensive
Despite the potential shown by Cline’s self-hosted architecture, the report has sparked different voices in the community. Some developers believe that simply comparing token prices cannot truly reflect the cost-effectiveness of an AI model in real-world applications.
A well-known community user, @morganlinton, sharply replied in the comments, saying: “If you look at cost per task, Kimi K3 is actually more expensive than Fable 5. We can’t just look at the price of input and output tokens.” They added an example: completing the same specific task might cost Fable 5 only $12.18, but using Kimi K3 would require $27.43.
This viewpoint points to a key blind spot: when a model has weaker reasoning logic and instruction-following ability, it often needs to consume more tokens by repeatedly trying, which can make Kimi K3, in some complex scenarios, end up being a relatively expensive model. This debate over compute costs also highlights that when evaluating AI solutions, enterprises must strike a delicate balance between unit price and real task efficiency.