About TrendForce News

TrendForce News operates independently from our research team, curating key semiconductor and tech updates to support timely, informed decisions.

[News] Google Says Memory Tops 75% of Server BOM; TPU-Software Dual Strategy Targets Memory Wall



Google has highlighted the growing “memory wall” facing AI computing at SEMICON Taiwan 2026. According to Commercial Times, Nikhil Cherian, Senior Director of Supply Chain Infrastructure at Google Cloud under Alphabet, said the rise of multimodal and mixture-of-experts (MoE) architectures is shifting AI computing from compute-bound to memory-bound, with high-performance memory now accounting for more than 75% of an AI server’s hardware bill of materials (BOM) cost.

Cherian said the rapid development of large language models and AI agent workflows is pushing the memory hierarchy against physical limits in capacity, bandwidth, latency, and power consumption. To address these bottlenecks, Google is taking a differentiated hardware approach, with the TPU 8i targeting low-latency inference and the TPU 8t designed for ultra-large-scale training, the report notes.

Google Takes a Dual-Track Approach to Break the Memory Wall

On the hardware side, the report highlights that the inference-focused TPU 8i features 288 GB of high-bandwidth memory (HBM) and triples on-chip SRAM capacity to 384 MiB, allowing dynamic conversational states and key-value (KV) cache to remain on-chip for zero off-chip latency. The training-focused TPU 8t, meanwhile, connects 9,600 chips into a massive compute cluster with a shared HBM pool of up to 2 PB, reducing off-chip data-transfer bottlenecks. It also incorporates TPU Direct Storage to help maintain high accelerator utilization.

On the software and sustainability fronts, the report notes that Google has developed TurboQuant, a training-free lossless quantization algorithm that compresses large-model KV cache from 32 bits to 3 bits. The approach reduces memory usage by 6x without sacrificing accuracy while accelerating attention computation by 8x.

As for its potential impact on the memory market, a March report from Economic Daily News notes that TurboQuant mainly compresses key-value (KV) cache during AI inference and does not affect the HBM used for model training and weight storage or long-term storage demand. Its impact on the broader memory market is therefore expected to be limited. However, by lowering computing costs, TurboQuant could help accelerate AI adoption and drive greater demand for end devices and edge computing.

Memory is taking up a growing share of AI infrastructure spending. According to TrendForce, DRAM and NAND Flash combined are projected to account for 47% of CSPs’ total CapEx in 2026, with the share rising to 68% in 2027, driven by rising memory contract prices and higher bit demand. TrendForce expects server DRAM contract prices to surge by around 270% in 2026, while enterprise SSD prices are projected to rise by approximately 235%.

Read more

(Photo credit: Google)

Please note that this article cites information from Commercial Times and Economic Daily News.


Get in touch with us