Post

Xiaomi just published something labs usually keep buried: the actual RL bill.


MiMo-V2.6 Pro and Flash both trained on 753k samples and stopped at step 30.
> Pro cost $2.62M.
> Flash cost $854k.
That extra 3.1x spend bought 6.89 points on DeepSWE.
Then Flash still edged Pro on AutomationBench.
You can already see the same product shape showing up everywhere now: one model for the absolute hard stuff, another sitting just underneath it that’s cheap enough to run all day.
ZAI, Deepseek, Qwen and Xiaomi now all have this in common.
Once the smaller model clears the reliability bar for most work, the bigger model starts looking more like something you route into selectively rather than the default for everything.
That probably ends up being the real model stack: cheap intelligence underneath, expensive intelligence pulled in only when the task actually earns it.
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
DEEPSEEKDEEPSEEK-1.58%

  • 1

Add a comment
Add a comment

Comment
GateUser-67882efc
4 hours ago
Ape In 🚀
0
GateUser-67882efc
4 hours ago
Buy to Generate 💎
0View Original
GateUser-67882efc
4 hours ago
Bull Run 🐂
0
GateUser-67882efc
4 hours ago
HODL Tight 💪
0
haw81
5 hours ago
First Review
Kosalana kosalana washa hawraman washa washington
0