Demystifying KV Cache: Expert Advice on Architectures, Best Practices, and Sizing

As generative AI shifts toward complex, long-context agentic workflows, the primary scaling bottleneck has moved from model training to inference deployment. Traditional architectures, constrained by local GPU HBM, force a 'recomputation tax', where GPUs waste valuable cycles recalculating conversation history every time context limits are hit. This creates latency, spikes Time-to-First-Token (TTFT), and limits concurrency. Join Anat Heilper, AI Architect at VAST Data, and Calvin Nieh, Technical Alliance Marketing Manager, to move beyond these limitations. Anat will dive into the engineering reality of deploying KV cache offloading, sharing technical insights from VAST’s collaboration with GPU and KV cache framework partners, as well as global cloud providers and AI-forward enterprises. 

Learn how to transform KV cache from limited and expensive local RAM into a high-value, persistent data asset using VAST’s Disaggregated Shared-Everything (DASE) architecture.

In this technical session, you will learn how to:

  • Eliminate the 'prefill tax' to achieve up to 20x faster TTFT for agentic and long-context RAG applications.

  • Maximize GPU utilization by offloading, parking, and fetching pre-computed KV tensors at network line rate.

  • Boost prefill efficiency with 6x–9.7x higher token throughput.

  • Apply expert-tested sizing heuristics and blueprints for optimizing high-performance, cost-efficient inference at scale across GPU, networking, and KV management stacks.

Join us on September 23rd at 12 pm ET | 9 am PT. We appreciate your participation!

* Required field.