Original title: Our Crypto AI Thesis (Part II): Decentralised Compute is King
Original author: Teng Yan, encryption researcher
Original Translation: Deep Tide TechFlow
Good morning! It's finally here.
Our entire paper is quite rich in content. In order to make it easier for everyone to understand (while also avoiding exceeding the size limit of email service providers), I have decided to divide it into several parts and gradually share it over the next month. Now, let's get started!
A huge missed opportunity that I have never been able to forget.
This matter has been weighing on my mind, because it is a clear opportunity that anyone in the follow market can see, but I missed it and didn't invest a penny.
No, this is not the next Solana killer, nor is it a memecoin with a funny hat-wearing dog.
It is...... NVIDIA。
NVDA year-to-date stock performance. Source: Google
Within just one year, NVIDIA's Market Cap soared from 1 trillion US dollars to 3 trillion US dollars, with the stock price tripling, even outperforming Bitcoin during the same period.
Of course, part of it is driven by the AI boom. But more importantly, this rise has a solid foundation in reality. NVIDIA's revenue in fiscal year 2024 reached 600 billion US dollars, a 126% rise from 2023. Behind this astonishing rise is the global tech giants competing to purchase GPUs to seize the initiative in the general artificial intelligence (AGI) arms race.
Why did I miss it?
In the past two years, my attention has been completely focused on the Crypto Assets field, and I haven't followed the developments in the AI field. This is a huge mistake that I regret to this day.
But this time, I won't make the same mistake again.
Today's Crypto AI gives me a sense of déjà vu.
We are on the verge of an innovation explosion. It bears striking similarities to the California Gold Rush of the mid-19th century—industries and cities rise overnight, infrastructure develops rapidly, and adventurous individuals reap huge rewards.
Like early NVIDIA, Crypto AI will be so obvious in hindsight.
Crypto AI: Unlimited Investment Opportunities
· Many people still see it as a 'castle in the air'.
· Crypto AI is currently in its early stages, and it may take another 1-2 years before reaching its hype peak.
· This field has at least $230 billion rise potential.
The core of Crypto AI is to combine artificial intelligence with encryption infrastructure. This makes it more likely to develop along the exponential rise trajectory of AI, rather than following the broader encryption market. Therefore, to stay ahead, you need to follow the latest AI research on Arxiv and communicate with the founders who believe they are building the next big thing.
Four Core Areas of Crypto AI
In the second part of my thesis, I will focus on analyzing the four most promising subfields in Crypto AI:
1.Decentralization calculation: model training, inference and GPU trading market
Data Network
Verifiable AI
AI agents running on-chain
This article is the result of several weeks of in-depth research and communication with the founders and teams in the Crypto AI field. It is not a detailed analysis of each area, but a high-level roadmap designed to inspire your curiosity, help you optimize your research direction, and guide your investment decisions.
I imagine the ecosystem of Decentralization AI as a hierarchical structure: starting from Decentralization computing and open data networks on one end, which provide the foundation for training Decentralization AI models.
All inputs and outputs of inference are verified through cryptography, economic incentives, and evaluation networks. These verified results flow to on-chain autonomous AI agents, as well as consumer-grade and enterprise-grade AI applications that users can trust.
Coordinated networking connects the entire ecosystem, enabling seamless communication and collaboration.
In this vision, any team engaged in AI development can access one or more levels in the ecosystem according to their own needs. Whether it is using Decentralization computing for model training or ensuring high-quality output through evaluation networks, this ecosystem provides a diverse range of choices.
Thanks to the composability of blockchain, I believe we are moving towards a modular future. Each layer will be highly specialized, and protocols will be optimized for specific functions rather than using integrated solutions.
In recent years, a large number of start-ups have sprung up at every layer of the Decentralization AI technology stack, presenting a "Cambrian" explosion. Most of these businesses are only 1-3 years old. This shows that we are still in the early stages of the industry.
In the most comprehensive and up-to-date Crypto AI startup ecosystem map I've seen, maintained by Casey and her team at topology.vc. This is an indispensable resource for anyone looking to track developments in this field.
As I delve into the various subfields of Crypto AI, I always wonder: how big are the opportunities here? I'm not following small markets, but rather huge opportunities that can expand to 100s of billions of dollars.
1. Market Size
When evaluating the market size, I will ask myself: Is this subfield creating a completely new market, or is it disrupting an existing market?
Taking Decentralization as an example, this is a typical disruptive field. We can estimate its potential by looking at the existing cloud computing market. Currently, the size of the cloud computing market is about 680 billion dollars, and it is expected to reach 25 trillion dollars by 2032.
In contrast, newer markets like AI intelligent agents are more difficult to quantify. Due to a lack of historical data, we can only estimate their capabilities through intuitive judgment and reasonable speculation about their problem-solving abilities. However, it is important to be cautious, as sometimes products that seem to be in a completely new market may actually just be solutions to existing problems.
2. Timing
Timing is the key to success. Although technology usually improves and becomes cheaper over time, the progress rate varies greatly in different fields.
In a certain subfield, how mature is the technology? Is it mature enough for large-scale application? Or is it still in the research stage and will take several years before actual application? Timing determines whether a field is worth investing in immediately to follow or whether it should be observed temporarily.
Taking Fully Homomorphic Encryption (FHE) as an example: its potential is undeniable, but the current technical performance is still too slow to achieve large-scale applications. It may take several more years before we see it entering the mainstream market. Therefore, I will prioritize following areas where the technology is already close to large-scale applications and focus my time and energy on those emerging opportunities.
If these subdomains were plotted on a 'Market Size vs. Timing' chart, it might look like this layout. It should be noted that this is just a conceptual sketch, not a strict guideline. There is also complexity within each domain - for example, different methods (such as zkML and opML) are at different stages of technical maturity in verifiable inference.
Nevertheless, I firmly believe that the future scale of AI will be exceptionally massive. Even the 'niche' areas that seem small today could potentially evolve into significant markets in the future.
At the same time, we also need to recognize that technological progress is not always linear - it often advances in leaps and bounds. When new technological breakthroughs occur, my views on market opportunities and scale will also adjust accordingly.
Based on the above framework, we will now break down the various sub-domains of Crypto AI one by one, exploring their development potential and investment opportunities.
· Decentralization computing is the core pillar of the entire Decentralization AI.
· The GPU market, Decentralization training, and Decentralization inference are closely related to each other and develop in a coordinated manner.
· The supply side mainly comes from GPU devices of small and medium-sized data centers and ordinary consumers.
· The demand side is currently relatively small, but it is gradually rising, mainly including price-sensitive users, users who do not have high requirements for latency, and some small-scale AI startups.
The biggest challenge facing the current Web3 GPU market is how to make these networks truly efficient.
· Coordinating the use of GPUs in the Decentralization network requires advanced engineering and robust network architecture design.
Currently, some Crypto AI teams are building Decentralization GPU networks to leverage globally underutilized computing resource pools to address the current situation of GPU demand far exceeding supply.
The core value of these GPU markets can be summarized in the following three points:
Calculating costs can be up to 90% lower than AWS. This low cost comes from two aspects: eliminating intermediaries and opening up the supply side. These markets allow users to access computing resources with the lowest global marginal cost.
No need to bind long-term contracts, no need for identity verification (KYC), no need to wait for approval.
Anti-censorship capability
In order to solve the supply side problem of the market, these markets obtain computing resources from the following sources:
· Enterprise-grade GPUs: Such as A100 and H100 high-performance GPUs, these devices usually come from small and medium-sized data centers (which are difficult to find enough customers when operating independently) or from BTCMiner who wants to diversify their sources of income. In addition, some teams are taking advantage of government-funded large-scale infrastructure projects, which construct a large number of data centers as part of technological development. These suppliers are often incentivized to continuously connect GPUs to the network to help offset the depreciation costs of the devices.
· Consumer-grade GPUs: Millions of gamers and home users connect their computers to the network and earn rewards through Token.
Currently, the demand side of Decentralization computing mainly includes the following types of users:
1. For users who are price-sensitive and do not have high latency requirements: For example, budget-limited researchers, independent AI developers, etc. They focus more on the cost rather than real-time processing capability. Due to budget constraints, they often find it difficult to afford the high fees of traditional cloud service giants such as AWS or Azure. Precise marketing targeting this group is crucial.
2. Small AI Startups: These companies require flexible and scalable computing resources but do not want to enter into long-term contracts with large cloud service providers. Attracting this group requires strengthening business cooperation as they are actively seeking alternative solutions beyond traditional cloud computing.
3. Crypto AI Startups: These companies are developing Decentralization AI products, but they need to rely on these Decentralization networks if they do not have their own computing resources.
4. Cloud Gaming: Although not directly related to AI, the demand for GPU resources in cloud gaming is rising rapidly.
One key point to remember is: developers always prioritize cost and reliability.
Many startups will see the scale of the GPU supply network as a sign of success, but in reality, this is just a "vanity metric".
The real bottleneck lies in the demand side, not the supply side. The key measure of success is not how many GPUs are in the network, but the utilization rate of GPUs and the actual number of GPUs rented.
Token incentive mechanisms are very effective in initiating the supply side and can quickly attract resources to join the network. However, they cannot directly solve the problem of insufficient demand. The real test is whether the product can be polished to a good enough state to stimulate potential demand.
As Haseeb Qureshi (from Dragonfly) said, this is the key.
Let the computing network truly start running
Currently, the biggest challenge facing the Web3 distributed GPU market is actually how to make these networks run efficiently.
This is not a simple matter.
Coordinating GPUs in a distributed network is an extremely complex task that involves multiple technical challenges, such as resource allocation, dynamic workload scaling, load balancing of nodes and GPUs, latency management, data transmission, fault tolerance, and handling diverse hardware devices distributed across the globe. These issues stack up to form a massive engineering challenge.
To solve these problems, it requires solid engineering and technical capabilities, as well as a robust and well-designed network architecture.
To better understand this, we can refer to Google's Kubernetes system. Kubernetes is widely regarded as the gold standard in container orchestration, automating tasks such as load balancing and scaling in distributed environments, which are similar to the challenges faced by distributed GPU networks. It is worth noting that Kubernetes is based on more than a decade of Google's distributed computing experience, and even so, it took several years of continuous iteration to perfect it.
Currently, some GPU computing markets that have been launched can handle small workloads, but once they try to scale to larger sizes, problems will be exposed. This may be because their architectural design has fundamental flaws.
Trustworthiness Issues: Challenges and Opportunities
Another important problem that Decentralization computing networks need to solve is how to ensure the credibility of Nodes, that is, how to validate whether each Node truly provides the computing power it claims. Currently, this verification process mostly relies on the reputation system of the network, and sometimes the computing providers are ranked based on reputation scores. Blockchain technology has a natural advantage in this field because it can achieve trustless verification mechanisms. Some startups, such as Gensyn and Spheron, are exploring how to solve this problem through trustless methods.
Currently, many Web3 teams are still working hard to address these challenges, which also means that there are still vast opportunities in this field.
So, how big is the market for the Decentralization computing network?
Currently, it may only account for a tiny fraction of the global cloud computing market (estimated at 6.8 trillion to 25 trillion dollars). However, as long as the cost of Decentralization computing is lower than that of traditional cloud service providers, there will definitely be demand, even if there is some additional friction in user experience.
I believe that in the short to medium term, the cost of Decentralization computing will remain relatively low. This is mainly due to two factors: first, Token subsidies, and second, the unlocking of supply from non-price-sensitive users. For example, if I can rent out my gaming laptop to earn extra income, whether it's $20 or $50 a month, I would be satisfied.
The true rise potential of the decentralized computing network, as well as the significant expansion of its market size, will depend on several key factors:
1. Feasibility of Decentralized AI Model Training: When the Decentralization network can support the training of AI models, it will bring huge market demand.
2. The explosion of inference demand: With the surge in AI inference demand, existing data centers may not be able to meet this demand. In fact, this trend has already begun to emerge. Jensen Huang of NVIDIA stated that the inference demand will rise by "a billion times".
3. Introduction of protocol (SLAs) for service level: Currently, Decentralization computing mainly provides services in a best-effort manner, and users may face uncertainties in service quality (such as normal operation time). With SLAs, these networks can provide standardized reliability and performance metrics, thereby breaking the key barriers adopted by enterprises and making Decentralization computing a viable alternative to traditional cloud computing.
Decentralization, permissionless computing is the foundation layer of the Decentralization AI ecosystem and one of its most important infrastructures.
Despite the continuous expansion of hardware Supply Chain such as GPU, I believe that we are still in the dawn of the 'human intelligence era'. In the future, the demand for computing power will be endless.
Please follow the key turning point that may trigger a revaluation of the GPU market - this turning point may come soon.
· The competition in the pure GPU market is very fierce, not only between Decentralization platforms, but also facing the strong rise of Web2 AI emerging cloud platforms (such as Vast.ai and Lambda).
Small nodes (such as 4 H100 GPUs) have limited use and are not in great demand in the market. However, it is almost impossible to find suppliers who sell large clusters because their demand is still very high.
Will the computing resources of the Decentralizationprotocol be consolidated by a dominant entity or continue to be distributed across multiple markets? I am more inclined towards the former and believe that the end result will exhibit a power law distribution, as consolidation often improves the efficiency of infrastructure. Of course, this process takes time, and during this period, market dispersion and chaos will continue.
Developers prefer to focus on building applications rather than spending time dealing with deployment and configuration issues. Therefore, the computing market needs to simplify these complexities and minimize the friction for users to access computing resources as much as possible.
Summary
· If the Scaling Laws hold, then it will become physically impractical to train next-generation cutting-edge AI models in a single data center in the future.
Training AI models requires a large amount of data transfer between GPUs, and the low interconnection speed of distributed GPU networks is often the biggest technical barrier.
· Researchers are exploring a variety of solutions and have made some breakthroughs (such as Open DiLoCo and DisTrO). These technological innovations will have a cumulative effect, accelerating the development of Decentralization training.
The future of Decentralization training may be more focused on small, specialized models designed for specific domains, rather than cutting-edge models for AGI.
· With the popularity of models like o1 from OpenAI, the demand for inference will experience an explosive rise, which also creates a huge opportunity for Decentralization inference networks.
Imagine this: a huge, world-changing AI model that is not developed by a secret top laboratory, but by millions of ordinary people working together. The GPUs of gamers are no longer just used to render cool graphics in Call of Duty, but to support a larger goal - an Open Source, collectively owned AI model with no centralized gatekeepers.
In such a future, basic-scale AI models are no longer exclusive to top laboratories, but the achievements of the entire population.
But back to reality, most heavyweight AI training is still concentrated in centralized data centers, and this trend may not change in the near future.
Companies like OpenAI are constantly expanding their massive GPU cluster. Elon Musk recently revealed that xAI is about to complete a data center with a GPU total equivalent to 200,000 H100s.
But the problem isn't just the number of GPUs. In its 2022 PaLM paper, Google proposed a key metric - Model FLOPS Utilization (MFU) - to measure the actual utilization of the maximum computing power of GPUs. Surprisingly, this utilization rate is usually only 35-40%.
Why is it so low? Despite the rapid advancement of GPU performance with Moore's Law, the improvement of network, memory, and storage devices lags far behind, creating significant bottlenecks. As a result, the GPU is often idle, waiting for data transfer to complete.
Currently, the fundamental reason for the highly centralized AI training is only one - efficiency.
Training large-scale models relies on the following key technologies:
· Data Parallelism: Splitting the dataset across multiple GPUs for parallel processing to accelerate the training process.
· Model parallelism: Distribute different parts of the model to multiple GPUs to overcome memory limitations.
These technologies require frequent data exchange between GPUs, so the interconnection speed (i.e., the rate of data transmission in the network) is crucial.
When the cost of training along AI models can be as high as 10 billion U.S. dollars, every bit of efficiency improvement is crucial.
Centralized data centers, with their high-speed interconnect technology, enable fast data transmission between GPUs, significantly reducing costs within training time. This is something that Decentralization networks currently cannot match... at least not yet.
If you talk to practitioners in the AI field, many of them may say that Decentralization training is not feasible.
In the architecture of Decentralization, the GPU clusters are not located in the same physical location, which leads to slower data transmission speed between them, becoming a major bottleneck. The training process requires GPU to synchronize and exchange data at each step. The farther the distance, the higher the latency. And higher latency means slower training speed and increased cost.
A training task that only takes a few days to complete in a centralized data center may take two weeks in a Decentralization environment and cost more. This is clearly not feasible.
However, this situation is changing.
It is exciting that the research on distributed training is rapidly rising. Researchers are exploring from multiple directions simultaneously, and the recent emergence of a large number of research results and papers is a clear proof. These technological advancements will have a cumulative effect and accelerate the development of decentralized training.
In addition, testing in actual production environments is also crucial as it helps us break through existing technological boundaries.
Currently, some Decentralization training techniques are able to handle small-scale models in low-speed interconnection environments. And cutting-edge research is working to extend these methods to larger-scale models.
For example, Prime Intellect's Open DiCoLo paper proposes a practical method: by dividing the GPU into 'islands', each island completes 500 local calculations before synchronization, reducing the bandwidth requirement to 1/500 of the original. This technology was initially researched by Google DeepMind for small models and has now been successfully scaled up to train a model with 10 billion parameters, and has recently been fully Open Source.
· The DisTrO framework from Nous Research further breaks through, reducing communication needs between GPUs by up to 10,000 times through optimizer technology, while successfully training a model with 1.2 billion parameters.
· This momentum is still continuing. Nous recently announced that they have completed pre-training a 150 billion parameter model, and its loss curve and convergence speed even surpass the performance of traditional centralized training.
(Tweet details)
In addition, methods like SWARM Parallelism and DTFMHE are also exploring how to train large-scale AI models on different types of devices, even if the speed and connection conditions of these devices are different.
Another challenge is how to manage diverse GPU hardware, especially consumer-grade GPUs commonly found in the Decentralization network, which often have limited memory. This issue is gradually being addressed through model parallelism technology (distributing different layers of the model to multiple devices).
Currently, the model scale of the Decentralization training method still lags far behind the cutting-edge models (according to reports, the parameter size of GPT-4 is close to one trillion, which is 100 times that of the 10 billion parameter model of Prime Intellect). To achieve true scalability, we need to make significant breakthroughs in model architecture design, network infrastructure, and task allocation strategies.
But we can boldly imagine: in the future, Decentralization training may be able to aggregate more GPU computing power than the largest centralized data centers.
Pluralis Research (a team that is very worth following in the field of Decentralization) believes that this is not only possible, but inevitable. Centralized data centers are limited by physical conditions, such as space and power supply, while Decentralization networks can leverage virtually unlimited resources globally.
Even NVIDIA's Jensen Huang mentioned that asynchronous Decentralization training may be the key to unleashing the potential of AI expansion. In addition, distributed training networks also have stronger fault tolerance.
Therefore, in one possible future, the most powerful AI model in the world will be trained in a Decentralization manner.
This vision is exciting, but I still have reservations at the moment. We need more compelling evidence to prove that training super large-scale models in Decentralization is technically and economically feasible.
I believe that the best application scenario for Decentralization training may lie in smaller, specialized Open Source models designed for specific use cases, rather than competing with large-scale, AGI-targeted cutting-edge models. Some architectures, especially non-Transformer models, have proven to be very suitable for the environment of Decentralization.
In addition, the Token incentive mechanism will also be an important part of the future. Once Decentralization training becomes feasible on a large scale, Token can effectively incentivize and reward contributors, thereby driving the development of these networks.
Despite the long road ahead, the current progress is inspiring. The breakthrough in Decentralization training will not only benefit the Decentralization network, but also bring new possibilities for large tech companies and top AI labs...
Currently, most of the computing resources for AI are focused on training large-scale models. There is an arms race between top AI labs to develop the strongest foundational models, with the ultimate goal of achieving AGI.
But I believe that this concentrated investment in computing resources for training will gradually shift to inference in the next few years. As AI technology is increasingly integrated into the applications we use in our daily lives - from healthcare to the entertainment industry - the computing resources needed to support inference will become enormous.
This trend is not without reason. Inference-time Compute Scaling has become a hot topic in the AI field. OpenAI recently released a preview/mini version of its latest model o1 (code name: Strawberry), which has a notable feature: it takes time to think. Specifically, it analyzes the steps it needs to take to answer a question and gradually completes these steps.
This model is designed for more complex tasks that require planning, such as solving crossword puzzles, and can handle problems that require Depth reasoning. Although it generates responses slower, the results are more detailed and thoughtful. However, this design also comes with high operational costs, with inference costs 25 times that of GPT-4.
From this trend, it can be seen that the next leap in AI performance will not only rely on training larger models, but also on the expansion of computational power in the inference phase.
If you want to learn more, several studies have already proven:
· By expanding inference calculations through repeated sampling, significant performance improvements can be achieved in many tasks.
The inference phase also follows a kind of exponential expansion law (Scaling Law).
Once powerful AI models are trained, their inference tasks (i.e. the practical application phase) can be offloaded to a Decentralization computing network. This approach is very attractive for the following reasons:
· The resource requirements for inference are much lower than for training. After training, the model can be compressed and optimized using techniques such as quantization, pruning, or distillation. The model can even be split using tensor parallelism or pipeline parallelism, allowing it to run on ordinary consumer devices. Inference does not require the use of high-end GPUs.
· This trend has already begun to emerge. For example, Exo Labs has found a way to run a Llama3 model with 450 billion parameters on consumer-grade hardware like MacBook and Mac Mini. By distributing inference tasks across multiple devices, even large-scale computing requirements can be efficiently and cost-effectively completed.
· Better user experience: Deploying computing power closer to users can significantly drop latency, which is crucial for real-time applications such as games, augmented reality (AR), or autonomous vehicles - in these scenarios, every millisecond of latency can make a difference in user experience.
We can think of Decentralization reasoning as AI's CDN (Content Delivery Network). Traditional CDNs quickly deliver website content by connecting to nearby servers, while Decentralization reasoning uses local computing resources to generate AI responses at extremely fast speeds. In this way, AI applications can become more efficient, respond faster, and be more reliable.
This trend has begun to emerge. The performance of Apple's latest M4 Pro chip is already close to NVIDIA's RTX 3070 Ti - a high-performance GPU that was once exclusive to hardcore gamers. Nowadays, the hardware we use in our daily lives is becoming more and more capable of handling complex AI workloads.
To make Decentralization inference network truly successful, it is necessary to provide participants with sufficient economic incentives. The computing Nodes in the network need to be rewarded reasonably for their contributions to computing power, while the system also needs to ensure fairness and efficiency in reward distribution. In addition, geographical diversity is also crucial. It can not only reduce the latency of inference tasks but also enhance the network's fault tolerance, thereby enhancing overall stability.
So, what is the best way to build a Decentralization network? The answer is Crypto Assets.
Token is a powerful tool that unifies the interests of all participants, ensuring that everyone is working towards the same goal: expanding the network and increasing the value of Token.
In addition, TOKEN can greatly accelerate the growth of the network. They help solve the classic 'chicken and egg' problem that many networks face in their early development. By rewarding early adopters, TOKEN can drive more people to participate in network construction from the beginning.
The success of BTC and Ethereum has proven the effectiveness of this mechanism - they have gathered the largest computing power pool on Earth.
The decentralized reasoning network will be the next successor. Through the geographical diversity, these networks can reduce latency, improve fault tolerance, and bring AI services closer to users. And with the incentive mechanism driven by Crypto Assets, the expansion speed and efficiency of the decentralized network will far exceed traditional networks.
Teng Yan
In the upcoming series of articles, we will delve into data networks and explore how they help overcome the data bottlenecks faced by AI.
This article is for educational purposes only and does not constitute any financial advice. This is not an endorsement of asset trading or financial decision-making. Please be sure to conduct your own research and exercise caution when making investment choices.
Original link
: