IOSG: From Hot Storage to Cold Memory—Decentralized Storage Amid the Hot Storage Boom in the AI Era

Authors: 0xjacobzhao, IOSG

Recently, Changxin Memory—the “top homegrown memory company”—officially listed on the ChiNext board, and its astonishing 500% surge has electrified the market. Although the storage sector as a whole has still been disturbed by the aftershocks of a recent pullback, AI storage has continued to be疯狂重新定价 by capital amid the current wave of technology narratives. Meanwhile, in the Web3 space, decentralized storage has fallen into long-term stagnation and disappointment. Why, despite both being branded with the name “storage,” has market performance been as different as ice and fire? The root cause lies in a fundamental split in the underlying value functions.

The AI-era revaluation of storage is, at its core, a frenzy centered on “hot data efficiency.” It is aimed at maximizing compute utilization and commercial monetization. Decentralized storage, in contrast, adheres to the value proposition of “cold data trustworthiness.” It defends data fairness, resistance to censorship, and humanity’s long-term memory. The former is an efficiency system for hot data; the latter is a trusted system for cold data. The current capital markets are undoubtedly坚定站在 “efficiency” side, but human civilization ultimately still needs an indelible memory foundation. The long-term value of trustworthy cold storage has never disappeared—it has simply been hiding in the darker side of market cycles, waiting for the era to reprice it.

Why storage is becoming an AI industry-chain focus again

In the traditional IT era, storage was a “capacity business.” Enterprise CIOs cared about unit-capacity cost, drive reliability, disaster-recovery backup plans, archiving strategies, and device refresh cycles lasting 3–5 years. Storage was viewed as a follow-on accessory purchased alongside servers.

This storage boom is not a traditional cycle recovery, but a re-pricing of data flow capabilities driven by AI. In the era of large models, the logic of storage has mutated from “capacity first” to “efficiency above all.” It is locked in on extreme indicators such as maximizing GPU “feeding” rates, checkpoint writing, and RAG with ultra-low latency. This marks storage’s value rising from “the final parking place for data” to “a high-speed channel for data entering computation.”

The evolution of resource bottlenecks in AI infrastructure is essentially a battle to “fill the wood barrel effect” (wooden-barrel principle). Compute’s true utilization is not a simple linear accumulation of single assets, but a harsh multiplier effect: compute true utilization = GPU × HBM × DRAM × SSD × network × file system. Any weak link will cause overall compute utilization to collapse. In the AI era, storage is for the first time transforming from a “cost center” into an “efficiency engine.” This is the fundamental logic behind storage being repriced.

AI storage architecture panorama: from the HBM bandwidth “organs” down to the data lake base

AI storage is by no means a pile-up of single hardware components, but a tightly coupled, layered scheduling system. In this system, industrial value and capital focus are highly concentrated on HBM, enterprise-grade SSDs, SSD controllers, NVMe/CXL protocols, and high-performance storage systems. To clearly break down where the value flows, we divide the AI storage architecture from top to bottom into four core layers:

  • Compute-near memory layer (bandwidth core): with HBM as the absolute main force, supported by DRAM and CXL memory pooling technology. This layer directly matches GPU/CPU packaging or the bus, aiming to break the “memory wall”—it is the first gate that determines whether compute can be fully released.

  • High-speed persistent storage layer (IO hub): the core logic is enterprise SSD = NAND flash chips + SSD controller + NVMe/PCIe data path. This layer handles high-frequency checkpoint writes, loading massive training datasets, and RAG hot data caching, making it the clearest incremental persistent storage demand for AI data centers.

  • Low-cost high-capacity storage layer (capacity foundation): made up of HDD, cold storage, and data lake archival systems. Faced with the exponential growth of multi-modal raw data, historical logs, and compliance backups, this layer still offers irreplaceable TCO (total cost of ownership) advantages.

  • AI storage system & data software (scheduling brain): includes high-performance parallel file systems, distributed object storage, vector databases, and a RAG data governance layer. What AI truly consumes is not raw hardware, but data availability that has been efficiently organized, indexed, and permissioned by the software stack.

As an ecosystem extension, decentralized storage does not directly rush into the millisecond-level hot data race. Instead, it anchors public dataset proofs, AI training data provenance, and long-term cold-memory archiving—thereby establishing its unique niche as a “trusted cold layer.”

HBM: the “bandwidth organ” closest to compute in the AI storage chain

High Bandwidth Memory (HBM) is not traditional storage, but a high-bandwidth memory layer near the GPU. Its core mission is not to store data, but to continuously “feed” data to compute with extremely high bandwidth. HBM is the most compute-adjacent and most deterministic link in the AI storage chain—it directly determines whether a GPU can be “fed,” and it is currently the most critical supply-chain bottleneck.

The core HBM architecture is “3D DRAM stacking + 2.5D advanced packaging.” Through TSV vertical stacking and CoWoS heterogeneous integration, it compresses the compute-storage distance to achieve a jump in bandwidth by a generation-level gap. Its industrial barriers are not only DRAM design, but a system engineering covering DRAM process technology, TSV, ultra-thin stacking, packaging, heat dissipation, testing, and customer certification. Any yield defect in any step can scrap the entire HBM stack.

Currently, only three global giants—SK hynix, Samsung, and Micron—can stably mass-produce, forming triple moats: top-tier DRAM process, packaging capability, and NVIDIA/AMD customer certifications.

| Generation | | --- | Capacity/chip | Bandwidth | Interface width | Mass production time | Main customers | | HBM2e | 8–16GB | 460 GB/s | 1024-bit | 2019–2020 | Mature generation used for previous AI accelerators like A100 | | HBM3 | 24GB | 819 GB/s | 1024-bit | 2022 | SK hynix gains significant first-mover advantage in the H100 cycle | | HBM3e | 24–36GB | 1.2 TB/s | 1024-bit | 2024–2025 | Current main volume for AI GPUs. Suppliers centered on SK hynix, Micron, and Samsung | | HBM4 | 32–48GB |

2 TB/s | 2048-bit | 2025–2026 | Planned production ramp-in and customer certification phase for the next generation AI in 2026 | | HBM4E | 64GB+ | 2 TB/s | 2048-bit+ | 2027 onwards | Planning/R&D after 2027 |

DRAM and CXL: system memory base and a memory pooling engine

HBM solves the extreme near-GPU bandwidth limit. DRAM solidifies the server system memory base, while CXL aims to break physical boundaries and reconstruct how data center memory resources are organized.

  • DRAM: primarily carries CPU-side cache, data pre-processing, intermediate state temporary storage, and system runtime—making it the most basic system memory layer for servers. The global DRAM market is highly concentrated among SK hynix, Samsung, and Micron; Changxin Memory (CXMT) is the key variable for China’s domestic DRAM substitution.

  • CXL (Compute Express Link): a new-generation cache-coherent interconnect protocol for data centers, designed to break limits of traditional DIMM slots, local memory capacity, and isolated server memory resource “islands,” pushing memory architecture toward expansion, pooling, and shared evolution. Currently, CXL is still in the early stage of moving from platform support to large-scale deployment; its architectural value over the mid-to-long term remains high. Key companies include Astera Labs and C^enT Technology.

Enterprise SSD: the data hub built by NAND, controllers, and NVMe

Enterprise SSDs are the most core high-throughput persistent storage increment in AI data centers. By providing extremely high throughput, ultra-low latency, and stable QoS, they continuously “feed” data to GPUs across the full lifecycle: training data loading, checkpoint writes, RAG retrieval, inference cache, and log回流.

In AI storage architecture, SSDs are not isolated hardware; they are a highly coupled system. It can be distilled into an industry formula:

enterprise SSD = NAND flash chips + SSD controllers + NVMe/PCIe data path.

The three layers represent independent industry-chain segments:

  • NAND flash chips (raw-material layer): determine storage density and unit cost; controllers manage performance release and lifespan management. Representative companies: Samsung, SK hynix (Solidigm), Micron, Kioxia, Western Digital, Yangtze Memory.

  • SSD controllers (performance enablement layer): determine performance release, data error correction, QoS stability, and wear-leveling balance. Representative companies: Phison (Phison Electronics), Silicon Motion (Silicon Motion Technology), Marvell, Maxio (Microchip Integration).

  • NVMe/PCIe (data path layer): determines the transmission efficiency from storage to compute. Combined with GPUDirect Storage, it reduces CPU memory bounce buffers and CPU involvement, significantly alleviating I/O bottlenecks. Representative companies: Broadcom, Marvell, Astera Labs

HDD / cold storage / archiving: the low-cost foundation of the AI data lake

AI will not eliminate HDDs. As multi-modal large models increase their demand for video and image data, and enterprises face exponential growth in compliance logs and historical datasets, cold storage demand at low cost is rising in parallel. In the AI storage architecture, SSDs and HDDs collaborate in layers based on business value: SSD handles hot data and high-throughput needs, while HDD handles low cost and long-cycle preservation. Representative companies include Seagate, Western Digital, Toshiba.

AI storage software stack: the scheduling core of data availability

What AI truly consumes is never bare disks, but “data services” precisely organized by the software stack. This architecture turns underlying hardware into knowledge assets that AI can directly call at the top layer. It is divided into four layers:

  • High-performance storage systems (provisioning system): focused on concurrent throughput and low latency; solves the “data hunger” problem for GPU clusters through parallel file systems, ensuring ultra-fast data flow for training and inference. Representative companies: VAST Data, WEKA, Pure Storage.

  • Object storage (raw data lake): centered on Object, Key, and Metadata management, carrying massive unstructured data. It does not chase extreme low latency; instead it builds a capacity foundation with low cost and cloud-native characteristics. Representative company: AWS S3

  • Vector database (semantic index layer): vector databases store, index, and retrieve vectors generated by embedding models, enabling AI to precisely locate relevant content from massive knowledge. Representative companies: Pinecone, Milvus

  • RAG data layer (knowledge calling layer): goes beyond single retrieval, covering data slicing, cleaning, permission control, and reference provenance, ensuring enterprise data can be called by large models safely, accurately, and with traceability. Representative company: Databricks

From AI hot storage to decentralized cold memory: maximize efficiency vs maximize trust

AI storage is an extreme-efficiency-driven system, with its value function focused on maximizing compute output. HBM bandwidth determines whether GPUs can be “fed,” SSD throughput determines dataset and checkpoint read/write efficiency, and low latency affects RAG and real-time inference experience. These metrics ultimately converge into GPU utilization and cost per Token, directly determining the business profit or loss of AI applications. The ultimate goal of AI storage is not to save, but to accelerate—it serves productivity.

Decentralized storage’s value function is entirely different. It asks whether data still exists after ten years, whether it has been tampered with, and whether it can withstand single-point censorship. Through cryptographic proofs and distributed networks, it constructs a public data base for open access and permanent preservation. Its ultimate goal is to defend the absolute truth and sovereign independence of data, serving fairness and anti-censorship needs, as well as civilization memory.

| Dimension | | --- | AI hot storage | Decentralized cold storage | | Core value | Maximize efficiency — accelerate data entering computation | Maximize trust — ensure data cannot be tampered with or deleted | | Data types | Hot/ warm/ real-time inference data | Cold / permanent archiving / public memory | | Key metrics | Bandwidth, throughput, latency, GPU utilization | Verifiable, censorship-resistant, immutable, permanent storage | | Revenue sources | Cloud vendors, AI labs, enterprise RAG systems | Public datasets, long-term archiving, on-chain applications | | Ultimate goals | Not to preserve, but to accelerate | Not speed, but trust | | Market status | Hot — a super cycle is underway | Silent — valuation collapses, narratives are being drained |

AI storage is “hot storage” that provides fuel for future productivity; decentralized storage is “cold memory” that preserves irremovable historical records for human civilization. The former serves efficiency and pursues extreme speed; the latter serves trust and defends silent memory. The former determines how fast the model runs; the latter determines whether memory will be deleted. Currently, market mechanisms reward productivity’s efficiency, so AI storage is at the center of the storm, while decentralized storage appears to be experiencing silence—valuation collapse and narrative being drained.

The vision and reality of decentralized storage

There are many decentralized storage projects, but in terms of industry mindshare and ecosystem depth, Filecoin and Arweave remain the core representatives. Although both are categorized as “decentralized storage,” their underlying architectural philosophies are almost two completely different paths—one uses market-based contracts to approximate AWS’s elasticity; the other uses a one-time social contract to approximate a library’s permanence.

  • Filecoin: builds the most complete verifiable economic system through PoRep and PoSt. It should not keep fighting AWS on consumer-grade cloud drives. Instead, it should pivot to AI data provenance (Provenance), public dataset hosting, and compliant archiving to provide verifiable chains for model audits and copyright proofs. The必由之路 is packaging it into an S3-compatible API and supporting fiat payments—upgrading from a “cheap storage market” to a “verifiable compute infrastructure.”

  • Arweave: with the narrative of “one-time payment, permanent storage,” it uses Blockweave and SPoRA to push incentives so miners preserve and quickly access as much data as possible, especially scarce historical data. Its best position is the foundation of human public memory—preserving records of human rights, war crime evidence, and cultural classics; archiving legal and financial history to provide permanent long-term memory accessible by AI agents. Arweave’s value is not speed, but its capacity to carry civilization memory across cycles.

| Dimension | | --- | Filecoin — Verifiable storage market (data as of June 2026 source: Filfox) | Arweave — Permanent public memory layer (data as of June 2026 source: ViewBlock) | | Core philosophy | “Storage is a market”—price discovery, elastic supply, contracts can expire without renewal | “Storage is a public good”—a one-time social contract; data exists without relying on any entity to keep paying continuously | | Underlying structure | Standard blockchain + IPFS content addressing; storage separated from the chain | Blockweave — each new block links to the previous block and a random historical “recall block” | | Consensus mechanism | Expected Consensus (EC): proof of replication (PoRep) + proof of spacetime (PoSt) | SPoRA (Succinct Proof of Random Access, iterated from PoA in 2021) | | Storage proof logic | PoRep proves miners generate the unique replica of the data; PoSt continuously proves that replica is still fully preserved | Mining requires proving access to random recall blocks—forcing miners to retain as much history as possible, including niche data | | Market structure | Two-layer market architecture: Storage Market (storage transactions) + Retrieval Market (retrieval transactions), on-chain matching and off-chain data transfer | Single-layer permanent writing; no independent retrieval market, relying on Gateway (e.g., AR.IO) to provide retrieval services | | Payment model | Storage market / contract-based — customers and storage providers (SPs) sign periodic leasing agreements, pay on demand | Endowment perpetual donation model — pay once; funds go into an interest-bearing pool, theoretically paying miners permanently | | Data availability guarantees | Guarantees integrity (data is not tampered with), but does not inherently guarantee retrieval speed; retrieval services need to be purchased separately | Guarantees persistence and accessibility; anyone holding a transaction ID can permanently view and download without relying on the original uploader’s wallet | | Typical scenarios | Enterprise cold archiving, compliance data proofs, verifiable snapshots of AI training data, on-demand elastic storage | Permaweb permanent websites, NFT metadata, historical archives, long-term memory for AI agents (AO compute layer) | | Token model | FIL’s maximum total supply is 2 billion; actual circulating amount is affected by block reward release, staking, penalties, and burning mechanisms | AR maximum supply is about 66 million; the circulating ratio is already close to the cap, so additional inflation has limited impact. | | Network scale | Quality Adjusted Power: about 14,848 PiB | Network Size: about 20.2 PiB | | Cumulative stored data | Active deals stored data: about 1,110 PiB | Weave Size about 0.345 PiB | | Miners / nodes | About 611 active miners | About 100 online nodes | | Storage cost | Filecoin Cloud $2.50/TiB/month/per replica | About 10.4–10.7 AR/GiB (≈$20–21) |

The困境 of decentralized storage projects like Filecoin and Arweave is not that their value propositions are wrong, but that there is a long-term mismatch among productization, retrieval experience, real demand, and Token incentives. This reveals a huge gap from geek ideals to mainstream commercial applications:

  • Supply-demand incentive mismatch: Early networks represented by Filecoin expanded rapidly via Tokens, but they did not build sufficiently strong paid demand on the demand side, resulting in massive capacity yet inadequate utilization and paid conversion. Rewards are “I can store,” not “I need you to store.”

  • Lack of enterprise-grade service capability: AWS’s moat is not the hard drives—it is the “data operating system” formed by APIs, SLAs, permission management, compliance audits, and technical support. Enterprises buy “peace of mind,” not experimental infrastructure that requires them to handle keys and node selection themselves.

  • Retrieval experience shortcomings: “put it in” does not mean “take it out” reliably with low latency. Dispersed nodes, complex topology, and the lack of unified SLAs make it hard to handle AI hot-data workflows; it is more suitable for trusted cold archiving and data provenance.

  • Insufficient privacy and compliance: Enterprise private data cannot simply be written into public permanent networks; the right to delete and permanent immutability are inherently conflicting. Decentralized storage is better suited for public data and long-term archives, not indiscriminately ingesting core private data.

  • Token economy amplification cycle: Bull markets’ financialization can mask insufficient demand; in bear markets, miners’ ROI declines and exposes commercialization shortcomings. Tokens can bootstrap cold-start supply, but they cannot automatically create demand or sustainable revenue.

Other decentralized storage projects often focus on a specific ecosystem or niche track: Storj/Sia’s cross-cycle industry mindshare and Web3 narrative influence are weaker than Filecoin/Arweave; BNB Greenfield/Walrus are tied to specific chain ecosystems of BNB or SUI; Celestia/EigenDA belong to the data availability (DA) layer, serving Rollup transaction confirmation rather than long-term archiving; AI/DA hybrid narrative projects like 0G attempt to integrate storage, data availability, compute, and AI agent settlement into a set of AI-native modular infrastructure, but their true demand, developer adoption, and a commercially closed loop still need verification.

The future opportunities for decentralized storage: a long pendulum between efficiency and trust

In the period of technical dividends exploding, capital chases efficiency relentlessly, and assets like GPUs and HBM are given massive premiums. Decentralized storage that advocates “trust and fairness” naturally gets marginalized. However, history’s pendulum will not stay forever on the efficiency side. Events such as unreasonable platform bans and content deletion, outbreaks of AI copyright lawsuits that force proof of data origins, data sovereignty conflicts triggered by geopolitical tensions, data monopolies causing public archives to disappear, and regulatory audit pressure regarding compliance of model training data may all lead to a re-pricing of “trusted storage.” And decentralized storage’s future opportunities can still show unique value in the following directions:

  • AI Data Provenance: combine cryptographic proofs to build “data lineage proofs,” addressing regulatory and audit pressure.

  • Public datasets and civilization archives: anchor censored archives and cultural heritage to build irreplaceable, non-deletable memory.

  • Trusted archiving and compliance evidence: enable trusted self-proof through hash proofs, providing high-grade digital notary services.

  • Fusion of ZK/TEE/DID technologies: resolve privacy tensions, upgrading from a single “storage protocol” to “trusted data infrastructure.”

  • Invisibility-style product route: provide S3-compatible APIs and fiat billing, letting users directly purchase “trusted archiving” services.

AI storage and decentralized storage—one pursues extreme efficiency to fuel our journey into the future; the other defends silent memory and protects our right to look back at the past. Today, the market unreservedly rewards efficiency, making decentralized storage look silent, even collapsed. But when the AI era further amplifies data monopolization, copyright disputes, and the fragility of historical memory, decentralized storage may come to a value reappraisal in the posture of a “trusted cold layer.” Memories that cannot be easily erased by platforms, companies, or any single power may transform—from the romance of idealism to essential infrastructure—as belief moves from the edge to necessity.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned