AI Compute Beyond GPUs: Why Packaging, Interconnects, and Storage Are the New U.S. Chip Themes

TradFi
Updated: 2026-09-22 07:26

On September 22, 2026, the Philadelphia Semiconductor Index widened its gain to 4%, closing at 12,398.17 points. According to Gate market data, Nvidia’s stock rose 2.3% to $227.38, Broadcom gained 1.6%, Advanced Micro Devices (AMD) climbed 9.95%, ASML rose 0.38%, and Intel increased 12.14%. AMD’s market cap first surpassed the $1 trillion mark, and Arm surged 17.16%.

Behind the surge in market sentiment, a deeper industry logic is coming into focus. Veteran chip designer and hardware analyst Mr. Bubble recently said in an interview on the Frictionless Podcast that the core bottleneck for AI compute infrastructure is shifting away from simply boosting chip compute power. It is moving toward advanced packaging, system interconnect latency, and memory hierarchy management. He noted that advanced process costs are rising rapidly, while gains in transistor density remain limited, leaving traditional Moore’s Law facing economic challenges. Instead of further shrinking process nodes, the industry is increasingly connecting multiple chips through advanced packaging to form a single computing system.

This article will break down how value distribution across the AI chip supply chain is being reshaped across four dimensions: process economics, packaging capacity, interconnect architecture, and memory hierarchy.

Economic challenges of process scaling and the rise of advanced packaging

For years, the market has measured chip progress by how transistor density roughly doubles every two years. But this framework faces a serious challenge. Mr. Bubble pointed out that the capital expenditures required to move to advanced process nodes have swollen to the level of several tens of billions of dollars, while transistor density improvements may be only 15% to 20%—"You can’t even double the number of transistors."

The industry’s response to this challenge is shifting from "shrinking transistors" to "expanding die area." Advanced packaging connects multiple chips into a single computing system, enabling them to function logically like one chip. This trend is reshaping the value distribution across the semiconductor supply chain.

The China Center for Information Industry Development (CCID) forecasts that, driven by continued growth in AI accelerator orders and the evolution of HBM stacks to 12 layers and beyond, the global advanced packaging market will reach $58.1 billion in 2026. By 2030, the market size is expected to approach $80 billion. Morgan Stanley’s calculations point to the same direction: the global advanced packaging market in 2026 is estimated at about $56.0 billion to $58.0 billion, representing the first time that the share of the overall OSAT/packaging and testing value chain exceeds 50%. Within that, AI-dedicated packaging margins are as high as 48% to 55%, making it the most profitable segment in the semiconductor value chain.

In its 2025 fourth-quarter earnings call, ASE said that due to semiconductor demand driven by AI adoption and an overall market recovery, 2026 advanced packaging and testing revenue is expected to double from $1.6 billion to $3.2 billion, with about 75% from packaging and 25% from testing. TSMC, meanwhile, said that 2025 advanced packaging revenue contribution is slightly above 10%. It expects growth to outpace the company’s average and that advanced packaging will account for "in the teens" of revenue in 2026. In addition, about 10% to 20% of 2026 capital expenditures will be allocated to advanced packaging, testing, photomask fabrication, and other projects.

CoWoS capacity expansion and the persistence of supply-demand gaps

The expansion pace of TSMC’s CoWoS capacity is a key window for observing advanced packaging bottlenecks. Based on TrendForce data, TSMC’s CoWoS monthly output is expected to continue ramping from about 70,000 wafers in 2025 to 120,000–140,000 wafers by end-2026. Monthly production in 2027 could rise further to around 200,000 wafers. JPMorgan’s outlook differs slightly: it expects CoWoS monthly capacity to reach 115,000 wafers, 190,000 wafers, and 225,000 wafers by the end of 2026, 2027, and 2028, respectively.

Even with rapid capacity expansion, supply-demand gaps remain. According to supply-chain information, Nvidia has already booked 800,000 to 850,000 wafers of CoWoS capacity at TSMC for 2026, securing more than half of the annual allocations for the year. After large consumption of capacity by top-tier customers, the packaging resources available for Broadcom, AMD, and cloud customers’ custom chips are relatively tight. DIGITIMES expects the share of CoWoS-L in TSMC’s CoWoS capacity to rise to around 75% in the fourth quarter of 2026, with new capacity mainly focused on larger-size AI chips.

Nvidia’s roadmap evolution directly reflects the real-world constraints of the packaging bottleneck. As Mr. Bubble analyzed, Nvidia initially planned to package four bare dies in the Rubin Ultra architecture. But due to physical limits of packaging, it was forced to step down to two-bare-die packaging per group, then connect two such packaged groups via a PCB. This means PCB manufacturers have become a new bottleneck in the compute interconnect chain—they need to ensure chips separated by several centimeters can operate together like a single chip.

CoWoS’s ongoing tightness is driving expansion of the advanced packaging supply ecosystem. Some back-end CoWoS stages are gradually being taken on by OSATs such as ASE/StatsChip, Amkor, and JCET. For example, ASE/StatsChip-related monthly capacity is expected to increase from about 30,000 wafers in 2026 to about 50,000 wafers in 2027, while Amkor is expected to rise from about 20,000 wafers to about 30,000 wafers. The industry structure is shifting from TSMC’s dominance at a single point toward a multi-player expansion model that includes wafer fabs, OSATs, substrate makers, and display panel companies.

Memory hierarchy reshaped: after HBM, competition shifts to flash and 3D DRAM

As AI inference expands its context window from 128K to the million-token scale, memory capacity faces unprecedented pressure. But Mr. Bubble proposed a practical breakdown: model weights should stay in HBM, while the massive historical context does not need to occupy expensive HBM space. Low-frequency context can be migrated to a larger-capacity, lower-cost flash storage tier.

Supply and demand dynamics for memory chips confirm how urgent this trend is. According to IDC’s forecast, in 2026 global DRAM supply will grow only 16% year over year, while NAND supply will grow 17%—both below the historical average. Meanwhile, demand growth rates are expected to be 45% and 38%, respectively. As a result, supply-demand gaps will keep widening through the first half of 2027. SK hynix disclosed that its entire 2026 HBM capacity was pre-sold by end-2025 to key customers including Nvidia and AMD. The industry broadly shows a "spot premium" phenomenon.

TrendForce data shows that DRAM industry revenue increased quarter over quarter by 59.5% in the second quarter of 2026. SK hynix’s revenue rose 37.9% quarter over quarter to $38.59 billion, with a 24.9% market share. Micron Technology’s revenue increased 65.5% quarter over quarter to $36.0 billion, and its market share rose to 23.3%. Choi Tae-won, chairman of Korea’s SK Group, recently said that because chip production faces systemic bottlenecks, a global shortage of memory chips is likely to continue until 2030.

Longer-term, the technology direction points to 3D DRAM—directly bonding DRAM above logic chips. By drastically shortening physical distances, it reduces latency, increases bandwidth, and lowers power consumption. At Hot Chips 2026, d-Matrix presented a Raptor 3D-DRAM generative inference accelerator. One card combines 32 GB of 3D DRAM with approximately 100 TB/s of memory bandwidth, requiring only about 0.37 picojoules per transmitted bit. In this architecture, the compute logic die is stacked directly on top of DRAM. Each chiplet provides 4 GB of memory and about 12.5 TB/s of bandwidth, and eight chiplets make up one card. Paper results show that Raptor achieves 4.71x higher inference throughput than an HBM configuration across multiple generative models and 2.44x higher than an SRAM reference configuration. d-Matrix expects to complete Raptor tape-out by end-2026 and has already integrated it into Nvidia’s NVLink Fusion ecosystem.

System-level design becomes a new barrier: paradigm shift from chips to racks

To Mr. Bubble, the fundamental unit of AI compute is no longer a single chip or even a single server. It is an entire rack—and sometimes multiple racks. The key rationale is that parallelization strategies for large-model training and inference make interconnect throughput and latency between chips decisive for cluster efficiency.

Nvidia addresses high-bandwidth interconnects through the NVL72 system and NVSwitch. But switches take up physical space that could otherwise be used for GPUs. The Google TPU team chose a completely different path: it builds part of the routing logic directly inside the chip, so the chip also acts as a pseudo-router. Its Scale-up expansion domain reaches 9,600 chips, which helps explain why a TPU’s per-card HBM configuration is lower than Nvidia GPUs’—yet it can still support training at massive scale.

Interconnect technology is going through a paradigm shift from copper cables to optical. In March 2026, three chip giants—AMD, Broadcom, and Nvidia—joined Microsoft, Meta, and OpenAI to form an optical compute interconnect multi-source protocol working group. The core mission is to develop a universal optical physical layer and unified components, replacing traditional copper cables with fiber optics. Under the plan, the standard will support vertical expansion connections both within AI systems and between racks. Ultimately, data transfer speeds will reach 3.2 Tb/s. Unlike previous proprietary, exclusive technology standards, the OCI specification follows a "protocol-agnostic" design philosophy. It can support both Nvidia’s NVLink and the UALink protocol led by AMD and Broadcom.

On the inference side, "prefill-decoding separation" is emerging as the mainstream architectural paradigm. Kimi’s reference design deploys the two stages on different nodes, and this architecture is being widely adopted. The prefill stage is compute-intensive, while the decoding stage is memory-bandwidth-intensive. Separating and deploying them can significantly improve overall system efficiency.

Long-term value at the hardware layer: why "owning the rails" is more reliable than "betting on the train"

Mr. Bubble’s investment logic has remained consistent: instead of betting on competing frontier AI labs, it is better to hold companies that own critical infrastructure capabilities, including hardware, packaging, storage, and other core components. This logic has been validated by recent market performance.

As of the close on September 22, 2026 (Beijing time), Gate stock market data shows that Nvidia traded at $227.38, up 2.30%. TSMC ADR was $445.14, up 2.41%. Micron Technology was $1,043.96, up 2.77%. SK hynix ADR was $188.86, up 0.73%. Intel rose 12.14%, and AMD rose 9.95%, with its market cap first surpassing the $1 trillion threshold.

Nvidia stock price, source: Gate stock market quotes

Intel stock price, source: Gate stock market quotes


AMD stock price, source: Gate stock market quotes

JPMorgan raised its CoWoS monthly capacity forecasts for TSMC at the end of 2026, 2027, and 2028 to 115,000 wafers, 190,000 wafers, and 225,000 wafers, respectively, and noted that the CoWoS shortage is widening. Evercore expects Broadcom’s electro-absorption modulated laser (EAM) capacity to increase from about 43 million to 44 million units in 2025 to 50 million units in 2026, reflecting structural growth in interconnect demand. These three companies—all of them, whether it’s Astera Labs’ view that the Scale-up single-chip interconnect content keeps rising, Marvell’s guidance on the growth rate of its connectivity business, or Intel’s emphasis on advanced packaging as the bottleneck for the entire industry—are all pointing to the same conclusion: interconnects, packaging, and memory are becoming new scarce components in AI infrastructure.

Conclusion

The competitive center of gravity for AI compute infrastructure is shifting from single-point breakthroughs driven by process miniaturization toward system-level coordination across packaging, interconnects, and storage. The challenge from Moore’s Law economics does not mean compute growth is stalling. It means value distribution is moving from "who has the most advanced process node" to "who controls the most critical bottleneck." Scarcity in advanced packaging capacity, tight HBM supply and demand, flash offload and the industrialization outlook for 3D DRAM, and the new bottleneck position of PCBs and system interconnects all form the new main line of US stock chip investing.

For investors following this trend, understanding the physical stacking logic from chips to racks may matter more than chasing short-term fluctuations in any single target. As Mr. Bubble put it—"the people who own the rails are more worth holding long-term than the companies building trains."

FAQ

Q1: Why is advanced packaging called the "new Moore’s Law"?

The traditional Moore’s Law relies on photolithography miniaturization to increase transistor density. But capital expenditures for advanced processes have ballooned to the level of several tens of billions of dollars, while the density improvement may only be 15% to 20%. Advanced packaging connects multiple chips into an integrated system, using "expanding die area" instead of "shrinking transistors," making it a more economically viable compute expansion path.

Q2: What are the storage technology directions after HBM?

In the near term, flash offload is a pragmatic solution for carrying massive inference contexts. In the mid term, 3D DRAM, by vertically bonding DRAM above the logic chip, is expected to simultaneously reduce latency, increase bandwidth, and lower power consumption. d-Matrix’s Raptor demonstrated its 3D-DRAM solution with over 100 TB/s bandwidth at Hot Chips 2026, and it is expected to complete tape-out by end-2026.

Q3: Why are PCB manufacturers becoming a new bottleneck for AI chips?

When a single package can’t accommodate more bare dies, chips need high-speed interconnects implemented through PCBs. Nvidia’s Rubin Ultra stepped down from a four-bare-die package to a two-bare-die double-packaged design. Signal integrity transmission between the two packaged groups depends on the PCB manufacturer’s process capabilities. This turns PCB from a passive connector into an active bottleneck that limits the pace of compute expansion.

Q4: After GPUs, which areas should AI chip investors focus on?

For packaging: focus on beneficiaries of CoWoS capacity expansion and companies taking on work in OSAT. For interconnects: focus on participants supporting the rollout of optical interconnect (OCI) standards. For storage: focus on HBM suppliers, early 3D DRAM technology innovators, and flash controller designers. At the system level: focus on companies leading in rack-level design capabilities.

Q5: Are valuations in the current US semiconductor sector reasonable?

As of September 22 (Beijing time), the Philadelphia Semiconductor Index is at 12,398.17. TSMC ADR’s P/E (TTM) is about 33x. Given that AI compute infrastructure is still in a structural expansion cycle, the supply-demand gaps in advanced packaging and HBM are expected to persist beyond 2027, and profit growth at leading companies may absorb current valuations. However, investors should watch potential risks, including export controls, geopolitical factors, and the possibility that capacity expansion ramps up slower than expected.

The content herein does not constitute any offer, solicitation, or recommendation. You should always seek independent professional advice before making any investment decisions. Please note that Gate may restrict or prohibit the use of all or a portion of the Services from Restricted Locations. For more information, please read the User Agreement

Share

sign up guide logosign up guide logo
sign up guide content imgsign up guide content img
Sign Up
Log In