NVIDIA Kyber Rack Could Carry 340.4 TB of DRAM and Cost $41.6 Million

NVIDIA's future Vera Rubin Ultra based Kyber platform could push AI server memory capacity to an entirely new level, with one 144 GPU rack reportedly carrying approximately 340.4 TB of combined HBM4E and LPDDR5X memory. According to Bank of America estimates, the memory alone could represent more than $5 million of component cost, while the complete Kyber rack is estimated at approximately $41.6 million. These figures are analyst estimates rather than official NVIDIA pricing and could change as the Rubin Ultra configuration continues evolving.

The projected memory pool consists of approximately 124.4 TB of HBM4E attached to the 144 Rubin Ultra GPUs and another 216 TB of LPDDR5X used by the Vera CPU side of the platform. Bank of America estimates HBM4E at approximately $19.76 per GB, placing the HBM4E bill near $2.5 million. Surprisingly, the much larger LPDDR5X pool is estimated to cost around $2.8 million, meaning conventional CPU memory could cost more per rack than the HBM4E surrounding NVIDIA's accelerators. NVIDIA has already moved Vera away from conventional server RDIMM designs toward compact SOCAMM memory, with its current Vera CPU supporting up to 1.5 TB of LPDDR5X and dedicated Vera CPU racks scaling to hundreds of terabytes.

The 124.4 TB HBM4E estimate works out to roughly 864 GB per GPU across 144 accelerators, below the 1 TB per Rubin Ultra package originally associated with NVIDIA's future roadmap. That aligns with recent reports suggesting NVIDIA has been evaluating alternative Rubin Ultra memory configurations as HBM supply, manufacturing complexity, power, and cost become increasingly important constraints. NVIDIA has not officially confirmed a reduction, so the final Kyber specification remains subject to change. The company has publicly demonstrated Kyber hardware capable of placing 144 GPUs inside a single rack and has working prototypes that scale toward much larger Vera Rubin Ultra systems.

The scale becomes clearer when compared with the current Rubin generation. NVIDIA's standard Rubin GPU carries 288 GB of HBM4 with up to 22 TB/s of memory bandwidth, while Vera Rubin NVL72 combines 72 Rubin GPUs with 36 Vera CPUs. NVIDIA says Rubin is designed specifically around increasingly memory intensive agentic AI workloads, where long context inference, large KV caches, and high concurrency require far more capacity and bandwidth than previous generations.

Memory suppliers are already preparing for the next transition. SK hynix has begun sampling 48 GB HBM4E at up to 16 Gbps, while NVIDIA's expanding use of LPDDR5X has also increased the strategic importance of suppliers capable of supporting enormous CPU memory pools. NVIDIA is reportedly expanding its Vera Rubin LPDDR5X supply chain, showing that the memory challenge extends far beyond HBM alone.

An additional UBS analysis illustrates how heavily memory already influences Vera Rubin economics. UBS estimates a Rubin GPU package at approximately $9,247 including HBM4, packaging, interposer, and supporting components, while a Vera CPU configuration with SOCAMM2 memory is estimated around $20,059. These estimates reinforce an increasingly important shift in AI infrastructure: processors may attract most of the attention, but memory capacity, bandwidth, packaging, and supply are becoming some of the largest factors determining what a complete AI factory actually costs.

340.4 TB of memory inside a single Kyber rack illustrates how quickly AI infrastructure is moving beyond the traditional idea of a GPU server. NVIDIA is effectively building enormous shared compute environments where GPU HBM, CPU LPDDR5X, networking, power delivery, cooling, and packaging all have to scale together.

The surprising part is not simply the $2.5 million HBM4E bill. It is that LPDDR5X could cost even more. Agentic AI is forcing enormous memory capacity onto the CPU side because thousands of simultaneous environments, tools, context states, and orchestration workloads need somewhere to live. If Bank of America's estimates are close to the final configuration, memory manufacturers may become even more strategically important to NVIDIA's future AI roadmap than they already are today.

Is 340.4 TB of memory per rack the clearest sign yet that memory rather than raw GPU compute could become the biggest bottleneck for future AI infrastructure?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

CXMT DDR5 Yield Reportedly Breaks 90% as China Narrows the DRAM Manufacturing Gap

Next
Next

Supermassive Games Could Cut Up to 75 Jobs as Another Restructuring Hits the Directive 8020 Studio