HBF Could Store AI Model Weights While HBM Handles Dynamic KV Cache

High Bandwidth Flash could reshape AI accelerator memory by moving large, mostly static model weights away from expensive High Bandwidth Memory. Under the emerging hybrid architecture, HBF would provide hundreds of gigabytes of NAND based capacity for read intensive data, while HBM would remain responsible for dynamic information requiring lower latency and frequent updates.

The memory capacity problem becomes clear with large language models. A model containing 100 billion parameters at FP16 precision requires approximately 200 GB just to hold its weights. Current accelerators must also reserve HBM for the KV cache, which stores contextual information generated while a model processes prompts and produces responses. Longer context windows and larger user batches increase that cache requirement, sometimes forcing data center operators to add GPUs primarily for their memory capacity rather than their compute performance.

The first HBF standard from SK hynix and Sandisk supports up to 512 GB per package and bandwidth ranging from 0.4 TB/s to 3 TB/s. HBF stacks NAND dies vertically and connects them through Through Silicon Vias, while a controller coordinates many simultaneous reads to overcome the limited performance of individual NAND cells. The specification also uses UCIe to connect HBF with GPUs, CPUs, and other accelerator chiplets.

NAND still has considerably higher access latency and lower write endurance than the DRAM used by HBM. This makes HBF better suited to model weights, which normally remain unchanged throughout inference and are repeatedly read by the accelerator. Dynamic KV cache data is created and updated continuously for each request, making HBM the more appropriate location because it provides faster access and stronger write endurance.

The separation is not absolute. SK hynix researchers propose storing model weights and shared precomputed KV caches inside HBF because both can remain read only. Newly generated KV caches and other frequently modified data would remain inside HBM. A latency hiding buffer could prefetch predictable information from HBF before the accelerator requires it, reducing the effect of NAND access latency.

Simulations of the proposed H³ architecture paired an NVIDIA B200 accelerator with 8 HBM3E stacks and 8 HBF stacks. SK hynix reported up to 2.69 times higher performance per watt than an HBM only configuration. For a workload using a KV cache containing 10 million tokens, the hybrid system supported a batch size up to 18.8 times larger. These figures come from architectural simulations rather than commercial hardware and will require validation through physical products and production AI workloads.

HBF will also remain more expensive than conventional solid state storage because it requires advanced packaging, specialized controllers, and high density integration near the accelerator. Its purpose is therefore not to replace HBM or enterprise storage, but to establish another memory layer between them. HBM would provide immediate access to rapidly changing data, HBF would retain massive read intensive information close to the processor, and solid state drives would continue handling larger long term data sets.

Duck IT Take

Describing HBF as storage for model weights and HBM as memory for KV cache provides a useful starting point, but practical AI systems will use a more flexible division. Static weights and shared precomputed caches are strong candidates for HBF, while dynamic session data will continue benefiting from HBM.

The strategic advantage is capacity rather than raw latency. Moving hundreds of gigabytes of static data into HBF could allow each accelerator to support larger models and longer contexts without adding GPUs only to obtain more HBM. However, the technology depends on effective prefetching, predictable access patterns, software support, and sufficient NAND endurance.

HBF could become a valuable component of AI inference infrastructure, but it complements HBM rather than replacing it.

Question for readers

Could shifting AI model weights into HBF reduce the number of GPUs required for large scale inference?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

CXMT DRAM Revenue Surges 716% as China Challenges the Big Three

Next
Next

Crimson Moon Launches September 1 for $19.99 With Replayable Gothic Runs