Futures
Access hundreds of perpetual contracts
CFD
Gold
One platform for global traditional assets
Options
Hot
Trade European-style vanilla options
Unified Account
Maximize your capital efficiency
Demo Trading
Introduction to Futures Trading
Learn the basics of futures trading
Futures Events
Join events to earn rewards
Demo Trading
Use virtual funds to practice risk-free trading
CFD
Stock CFD Derivatives
US Stocks
Access real US stocks and ETFs
HK Stocks
Trade quality Hong Kong-listed stocks
Korean Stocks
SK Hynix
Real Korean stocks and top assets
Stock Futures
High leverage, 24/7 trading
Tokenized Stocks
Backed by real stock assets
IPO Access
Unlock full access to global stock IPOs
GUSD
3.8%
Mint GUSD for Treasury RWA yields
Stocks Activities
Trade Popular Stocks and Unlock Generous Airdrops
Launch
CandyDrop
Collect candies to earn airdrops
Launchpool
Quick staking, earn potential new tokens
HODLer Airdrop
Hold GT and get massive airdrops for free
IPO Access
Unlock full access to global stock IPOs
Alpha Points
Trade on-chain assets and earn airdrops
Futures Points
Earn futures points and claim airdrop rewards
Promotions
AI
Gate AI
Your all-in-one conversational AI partner
Gate AI Bot
Use Gate AI directly in your social App
GateClaw
Gate Blue Lobster, ready to go
Gate for AI Agent
AI infrastructure, Gate MCP, Skills, and CLI
Gate Skills Hub
10K+ Skills
From office tasks to trading, the all-in-one skill hub makes AI even more useful.
Deng Yu won the Fields Medal: the problems he solved have stalled AI
Author: Meng Xing, Partner at Wuyuan Capital
Today the Fields Medal is flooding the internet; you should have seen Wang Hong and Deng Yu’s names everywhere. Interestingly, these days Wang Hong’s videos are being widely circulated, while Deng Yu has been much quieter—almost no exposure.
Over these days, I’ve been thinking again and again, and the one who keeps coming up is Deng Yu. Not because this is the story of a Chinese mathematician—but because the 125-year-old “old problem” that he chewed through together with Zaher Hani and Ma Xiao just happens to hit a blind spot in today’s AI development: increasingly critical, yet almost everyone avoids it. It’s even more worth discussing than “yet another Fields Medal.”
This problem is called Hilbert’s Sixth Problem.
The first time I heard about this concept was from Liu Ziming, a newly appointed assistant professor at Tsinghua Institute for Advanced Study. He said his vision was to solve Hilbert’s Sixth Problem. At the time, I didn’t understand what kind of problem it was that made it so important, and why the first five problems were skipped.
Later, after I learned more, I found that textbooks usually summarize it as “establishing physics in an axiomatic way.” That’s not wrong, but it’s too grand—so grand that you can hardly feel any connection between it and today’s AI entrepreneurship or AI investment. More specifically, what it wants to do is: starting from the fundamental laws describing a single microscopic particle, to derive—strictly, without skipping a single step—the macroscopic equations that describe gases and fluids.
I’d rather translate it into one plain saying:
Knowing how a molecule moves doesn’t mean you know how a clump of gas flows.
Knowing how a protein folds doesn’t mean you know how a cell will react when it encounters a drug.
Knowing how a consumer fills out a questionnaire doesn’t mean you know which direction the market will ultimately go.
Even knowing how parameters in a small model update doesn’t mean we truly understand generalization, grokking, and emergent abilities (emergence) in large models.
These questions look like they belong to physics, biology, social science, and machine learning—seemingly unrelated. But underneath, they share the same structure:
That is the true meaning of Hilbert’s Sixth Problem in the AI era. It’s also very likely the same issue that a large number of today’s so-called “world models” still haven’t truly solved.
A bunch of balls—why does it become a stream of water?
First, what exactly did Deng Yu’s team do?
Imagine a box filled with countless tiny hard balls. Each ball follows the most basic Newtonian mechanics: if it doesn’t collide, it moves in a straight line at constant velocity; if it collides, it changes direction according to conservation of momentum and conservation of energy. That’s it.
If the box only had ten balls, it would be easy—you’d number each ball, record its position and velocity, and compute step by step.
But in a real gas, a clump of air contains about 10^23 molecules. Even if you had infinite compute power to track every molecule’s position and velocity precisely, this kind of description would be almost useless for “understanding a breeze.” Because when we look at a breeze, we don’t care whether the 7.84e10 trillion molecules are moving left or right. What we really want to know is: how dense is the gas here? Which direction does it flow overall? What’s the temperature? What’s the pressure?
So, when facing the same physical world, humanity developed three completely different description languages:
At the bottom is microscopic: Newton’s hard-sphere dynamics, tracking every particle’s position and velocity.
In the middle is mesoscopic: the Boltzmann equation. It no longer numbers particles; instead, it counts how many particles are moving at a certain velocity near a certain position.
At the top is macroscopic: the Euler equations and Navier–Stokes–Fourier equations. At this level, you don’t need the full velocity distribution anymore; you only keep a few macroscopic quantities: density, average velocity, and temperature.
From the perspective of information content, this is an extremely aggressive compression: compressing the enormous state space of 10^23 particles into just a handful of continuously varying fields in space.
But the hardest part isn’t “doing averaging.”
Averaging is easy. The real difficulty is proving it:
In other words—if I only keep density, velocity, and temperature, can the density, velocity, and temperature at the next second really be determined solely from those macroscopic variables? Or once I take one step, do I still have to go back to the microscopic layer and re-check the full history of every molecule?
In physics and mathematics, this question has a dedicated name: closure. A macroscopic theory counts as truly independent only if it achieves closure; otherwise, it’s not a theory—just a pretty interface draped over microscopic simulation.
Remember this word first. It will keep coming up later.
The hard part isn’t collisions—it’s the “memory” left by collisions
When moving from Newtonian particles to the Boltzmann equation, the biggest obstacle is correlations between particles.
If particles were independent from start to finish, everything would be easy. The problem is that they collide.
When A and B collide, their states are no longer independent. Then B collides with C, and C collides with D. After a while, A may meet D again through another path. At that moment, A and D had no direct initial relationship, yet they’re already quietly connected by a collision chain:
It’s like throwing pebbles into water: ripples spread outward in circles. At first it’s only a small relationship between a pair of particles, but with time it can form a highly complex web of relations: collision chains branch; branches reconverge; and particles that have collided will collide again.
The Boltzmann equation can greatly simplify the problem by relying on a key approximation—often called “molecular chaos”:
This doesn’t mean they truly have no history. Of course they do—there were past collisions, and they may have indirectly affected each other through several turns. What Boltzmann truly needs is: as the number of particles grows, their size shrinks, and the gas becomes sufficiently dilute, those entangled historical correlations’ influence on the outcome of the current collision will ultimately become negligible.
In 1975, Oscar Lanford proved this rigorously for the first time.
But this “proof” doesn’t come from running a huge simulation to see whether a bunch of balls moving around “looks like” a gas. Instead, he proved a true limiting theorem: when the number of particles goes to infinity and the particles themselves go to zero size, the statistical distribution of the Newtonian hard-sphere system converges to a solution of the Boltzmann equation.
However, Lanford’s result has a fatal limitation: it only holds for a very short time—roughly a small fraction of the average collision time.
The issue isn’t that real gases stop obeying the Boltzmann equation after that time, nor that particles only collide finitely many times. The issue is that—his proof method becomes uncontrolled.
To predict a particle’s present, you have to trace back who it collided with. And to figure out the collision partner, you must trace back further to who that partner collided with before. The further you go back in time, the number of possible collision histories grows explosively. In a short time window, the event of “so many complicated collisions happening” is itself a low-probability event, and that low probability can still suppress the combinatorial explosion of collision histories. But once you extend the time horizon, the number of combinations of histories expands exponentially, and the traditional proof method can no longer sum these terms.
Anyone who has ever done a long rollout will find this dilemma very familiar: a model may be very accurate for predictions within one step or ten steps, but that doesn’t mean that after rolling it out 10,000 steps, errors, branching, and correlations can still be controlled. Lanford proved the direction is correct—but he only walked a short stretch from the bridgehead.
Deng Yu’s team’s 2024 work breaks through exactly this time bottleneck.
They proved: as long as the Boltzmann equation itself has sufficiently regular solutions during that time interval, then the strict derivation from the Newtonian hard-sphere system to the Boltzmann equation can be extended to any given finite time—not stuck in Lanford’s small slice.
But we need to interpret “arbitrarily long” carefully. It doesn’t mean they unconditionally prove it up to the end of the universe. The precise meaning is: you pick how long a finite time interval you want; as long as the Boltzmann equation does not develop singularities, blow up, or lose regularity during that interval, then the microscopic particle system will converge along the entire interval to follow it.
In 2025, the second paper moved another step forward.
From the Boltzmann equation to fluid equations like the Euler and Navier–Stokes–Fourier equations, mathematicians have accumulated a lot of prior work. The real bottleneck is: the first bridge from Newtonian particles to Boltzmann equation previously only worked for a very short time, so the two halves of the theory could never connect. The 2025 paper extends the 2024 result to two-dimensional and three-dimensional periodic spaces, then combines it with existing fluid-limit theories. For the first time, it connects the full chain:
So the most straightforward understanding is: in 2024, they made the “microscopic → mesoscopic” bridge longer; in 2025, they completed the entire “microscopic → mesoscopic → macroscopic” connection.
Of course, it’s important to clarify the boundaries. What they solved is the classic version of Hilbert’s Sixth Problem: deriving fluid equations from Newtonian mechanics through the Boltzmann theory—not axiomatizing all of physics. It has clear applicability conditions: dilute gases, hard-sphere models, specific scaling, and existence of regular solutions. It’s far from complex liquids, long-range interactions, chemical reactions, and realistic boundary conditions.
And I think the fact that they “explain the boundaries clearly” is extremely important. That is precisely what distinguishes mathematics and much of the AI narrative: a truly trustworthy model should not only tell you when it works, but also be honest about when it does not work. We’ll come back to this later.
The essence of world models isn’t remembering more—it’s knowing how to forget
Why Deng Yu’s work makes me feel it’s highly relevant to AI is because it puts a fact we often underestimate right on the table:
Going from particles to fluids isn’t just scaling up the particle model by 1 trillion times. Once you enter the macroscopic scale, even the variables used to describe the world change: at the microscopic level you use positions and velocities; at the mesoscopic level you use probability distributions; at the macroscopic level you use density, flow velocity, and temperature.
Crossing to the next scale isn’t an increase in quantity—it’s a replacement of language.
When people talk about world models today, they often unconsciously interpret them as a bigger and bigger end-to-end simulator: more data, larger parameters, longer videos, farther rollouts, more agents—almost as if once the bottom-level simulation is detailed enough, higher-level laws will automatically emerge, and we’ll also automatically know how to use it.
Hilbert’s Sixth Problem tells us that step has never been free.
A truly useful world model, besides being able to predict, must also find the “sufficient state” at a certain scale: a set of states that is simple enough that you don’t need to record every detail; yet complete enough that it can independently determine what will happen next. Density, velocity, and temperature are exactly such a set of sufficient states for the fluid world. They discard almost all molecular-level details, yet still describe a breeze, a stream of water, and even a shock wave in air.
So the real sophistication of a world model may very well be the opposite of “remembering more”—it is:
The word “safe” is crucial. A model doesn’t only have to compress information; it must also know whether the variables that were compressed away might come back through feedback, correlations, or long-term evolution to change the outcome. In the end, this brings us right back to closure.
We can sharpen this difficulty with a few pointed questions:
Which low-level differences are actually irrelevant? Which relationships must be preserved? Which microscopic perturbations will be averaged out? Which microscopic correlations will instead be amplified into macroscopic structures? And most critically—are the high-level variables we keep sufficient to independently determine the future?
If the answer is “yes,” then we get a genuine macroscopic world model. If each step still forces us to go back and check the complete state of every molecule, then what we have isn’t a theory—it’s just a microscopic simulator with a layer of skin.
The next sections are where I want to focus: the “cross-scale” problem appears exactly the same way in several of today’s hottest AI directions.
AI for Drug Discovery: from binding a target to curing a person—how many layers in between?
Drug discovery is the most typical example.
Today, AI can already do quite well on many single-point problems: predicting protein structures, predicting molecular and target binding, generating candidate compounds, estimating how a genetic perturbation affects cellular expression, and predicting how a certain cell reacts to a drug.
All of these are important. But whether a drug is ultimately valuable isn’t determined at the molecular level—it’s determined in the patient. In between is an extremely long chain of scales:
A molecule binding tightly to a protein doesn’t necessarily mean it will change the fate of a cell as expected. Changing the fate of a cell doesn’t necessarily mean it will repair tissue. Even if it’s effective in the target organ, it doesn’t mean it won’t cause toxicity in another organ.
As you move up each layer, the system generates new variables and new feedback loops. Knocking out a gene might trigger compensation through another pathway; inhibiting a protein might cause the cell to switch to a different metabolic route entirely; killing part of the tumor cells might actually create selective pressure for the remaining cells. Biological systems are rarely linear.
So the hardest part of drug discovery is often not making predictions more accurate for one layer, but rather: how to reliably transfer an intervention at one scale up to a higher scale.
The first time I encountered this concept was from my country’s first cell dynamics academician, Professor Feng Xiqiao. This is exactly what the virtual cell direction is really made of today. The goal of AI Virtual Cell isn’t simply reconstructing a cell atlas; it’s learning the cell’s dynamic responses under different conditions and perturbations. Researchers also explicitly say that it needs cross-measurement and cross-scale unified representation, and the model should be validated using perturbation predictions.
But note this: even if we truly have a very good virtual cell, we’re still far from virtual patient. Because in an organ model, a cell once again becomes the underlying “particle.” Cells communicate with each other, compete, migrate, and change each other’s states. No matter how good a cell model is, if you stack one million of these models together, you won’t automatically get a trustworthy organ model.
Life isn’t a one-time micro-to-macro event; it’s an entire staircase of scales marching upward. Every time you cross a level, it becomes a brand-new Hilbert’s Sixth Problem.
Agent Society: simulating a person and simulating a society are two different capabilities
Another highly typical direction is Agent Society.
Including companies such as Aaru and Simile, along with a long list of academic work, are trying to simulate consumers, voters, enterprises, policies, and social behaviors using lots of generative agents. Aaru defines itself as a multi-agent population simulation system. Simile proposes expanding step by step—from individuals, long-term journeys, and interactions—to the entire market and even society. In academia, researchers can already conduct deep interviews to build an agent for over 1,000 real participants and test how well these agents reproduce the participants’ attitudes and behaviors.
The first step in this kind of work is to make a single agent sufficiently like the person it represents—whether it can reproduce a person’s preferences, experiences, speaking style, and decision-making habits. If you get this step right, it’s already valuable.
But to go from “one person” to “a group of people,” there’s a massive logic jump hidden in between:
This sounds self-evident. But it’s almost the same kind of error as “knowing the Newtonian motion of every particle automatically tells you the fluid motion.”
Assume a model can answer “will you buy this product?” extremely accurately. If you sum up 100,000 such answers, you might at best get a faster, cheaper market research report. It’s still not a market model.
Because the market is never just the sum of 100,000 independent answers. It also depends on: whether consumers change their minds after seeing others buy; whether competitors cut prices; how platform algorithms allocate traffic; how the influence of KOLs and friends spreads; how scarcity changes attractiveness; and—after the company takes actions based on the prediction, whether the original prediction still holds.
A influences B, B influences C, and C in turn influences A. These feedback loops and correlations are exactly the most lethal parts of a social system. And in between, you have endless “collisions” and “re-collisions” just like in gas.
So the truly hard problem in Agent Society is never whether “a single agent looks like a person.” It is:
When evaluating companies like this, I think you need to break it into at least four layers—and these four layers are progressive:
First layer: individual effectiveness—does the agent resemble the person it represents?
Second layer: interaction effectiveness—do interactions between two agents resemble interactions between two real people in real situations?
Third layer: group effectiveness—the distributions of public opinion, prices, norms, and behaviors that emerge after many interactions—do they match real markets and society?
Fourth layer: intervention effectiveness—when we change prices, policies, products, or communication structures, can the model predict how the real world will change?
The one with enormous commercial value is the fourth layer. But most benchmarks stop at the first layer, at best the second layer.
This also exposes a very common misalignment in AI investment:
Grokking and Emergence: AI itself also has its own Hilbert’s Sixth Problem
I used to chat with Liu Ziming about grokking. At the time, he described his research problem as “a version of Hilbert’s Sixth Problem for AI,” and I thought that analogy was quite precise.
Grokking is an extremely counterintuitive phenomenon: a model might memorize the training set very early—yet perform poorly on the test set. If you keep training for a long, long time, it can suddenly shift from “memorization” to “understanding,” and its generalization ability jumps up.
Liu Ziming and their team’s approach is to first understand the internal mechanisms in very small, very simple models—such as the dynamics of representation learning and phase diagrams—then study how those mechanisms extend upward as the scale of data, the scale of parameters, the amount of regularization, and the training conditions change, eventually reaching larger models.
This is obviously not the same as particles to fluids, but structurally it’s highly similar. Because neural networks also have different scales:
At the microscopic level: how each parameter changes, how each gradient is updated, and how each neuron and feature forms.
At a higher level: whether the model forms structured representations, whether it transitions from memorization to generalization, whether it suddenly groks, and whether after crossing a certain threshold in scale it produces an emergent ability.
Today, we’re used to describing macroscopic outcomes with scaling laws: more parameters, more data, more compute, and loss drops according to some pattern. But scaling laws mainly tell us “how the results change,” not “what exactly happens inside.”
So AI’s own Hilbert’s Sixth Problem might be stated like this:
This is very important because there’s a highly common methodology in AI research: first discover a phenomenon in toy models, then provide a beautiful explanation, and then assume it still holds in models with billions or hundreds of billions of parameters. But moving from small to large models is itself a scale transition—it needs to be proven, not just justified by analogy.
The reverse is also true: observing a beautiful empirical curve in large models doesn’t mean we already understand the mechanism behind it. Like seeing fluid flowing doesn’t mean you’ve derived Navier–Stokes from Newtonian mechanics.
Weather, robots, and enterprise organizations all hit the same wall
The same problem shows up in many other directions.
Weather and turbulence. We can’t solve every tiny vortex analytically, so we must compress processes smaller than the grid scale into a closure model. Machine learning is being used to learn how these subgrid processes feed back into the large-scale flow field. But a closure that fits beautifully on training data, when put back into long-term physical simulations, may quickly become unstable, even breaking energy conservation, symmetry, and generalization under extreme conditions.
Autonomous driving and robotics. Predicting the next-frame video doesn’t mean you understand a long-horizon interactive physical world. From pixels to objects, from objects to scenes, from scenes to other participants’ intentions, and from single-vehicle behavior to city traffic—this is also a sequence of step-by-step scale jumps. A single-vehicle model that’s accurate at every step may produce entirely different macroscopic outcomes when placed into a system made of lots of human drivers, pedestrians, and other autonomous vehicles.
AI in enterprises. We often hear an ROI derivation: if each software engineer’s efficiency increases by 30%, then company R&D efficiency increases by 30%. But organizational output is never a simple sum of individual output. Writing more code can increase review workload; requirements coming out faster can worsen priority chaos; easier communication can mean more meetings; and localized efficiency improvements can be completely eaten up by organizational coordination costs. This is the same kind of error as “smarter agents must form a smarter agent society.” From individual capability to system output, there’s an entire gap filled with organizational structure, workflows, incentives, information flows, and responsibility mechanisms.
Seeing all this, the core judgment is worth pinning down again:
Scaling up may produce new regularities—but it won’t automatically tell you what those regularities are, nor will it automatically prove they’re reliable.
Every world model has its own Hilbert’s Sixth Problem
Hilbert’s Sixth Problem matters not just because it exists for 125 years, and not just because someone finally took a big step forward. It matters because it proposes an extremely deep worldview:
At different scales, the important objects differ, the variables used to describe them differ, the effective laws differ, and even the effective boundaries differ. Microscopic laws don’t automatically become macroscopic laws just because you scale up the system. Between them, you need a real theoretical leap: new representations, new assumptions, plus knowing which details can be thrown away and which correlations must be kept.
Today, we’re using AI to simulate more and more “worlds”: molecules, cells, materials, weather, robots, human behavior, markets and society, and even AI models themselves. And in every direction, if it goes to the end, it will likely hit the same question:
If we can’t prove it, what we may have is only a very strong local predictor, not a truly world model.
We may be able to predict a protein extremely accurately yet still not know why a patient recovers; we may be able to simulate a consumer very accurately yet still not know why trends form; we may clearly grok a small model yet still not know why intelligence emerges at larger scales.
Bigger models are not the same as bigger worlds.
A true world model isn’t about stuffing every detail into an infinitely large neural network until it copies the entire universe. It needs, at every scale, to know: what must be remembered, what can be forgotten, what disappears in averaging, what is amplified in interactions, what can emerge upward—and why.
Being able to translate across scales is where the truly difficult—and truly valuable—step of world models begins.