That distinction matters. KV cache — the stored key-value history used during inference — can be one of the main GPU-memory bottlenecks for long-context large language models. But it is not the same thing as all the GPU memory needed to host and serve a model.
A defensible way to describe DeepSeek V4 is:
DeepSeek V4 uses Hybrid Attention, Compressed Sparse Attention and Heavily Compressed Attention to reduce KV cache pressure in long-context inference; current public evidence does not support saying that total VRAM requirements fall by 98% .
That is a much more useful statement for engineers, buyers and AI teams. It says where the improvement appears to be — long-context attention and KV cache — without turning one memory component into a blanket promise about the entire serving stack.
DeepSeek’s API news page lists DeepSeek-V4 Preview with a release date of April 24, 2026 . The DeepSeek V4 model card names DeepSeek-V4-Pro and DeepSeek-V4-Flash, and describes V4 as a Mixture-of-Experts language-model series that retains the DeepSeekMoE framework and Multi-Token Prediction strategy while adding architectural changes including Hybrid Attention .
The strongest memory-related evidence concerns attention. NVIDIA’s developer article says Compressed Sparse Attention uses dynamic sequence compression to compress KV entries and reduce the KV cache memory footprint, then applies DeepSeek Sparse Attention to make attention matrices sparser. It also describes Heavily Compressed Attention as a more aggressive method that consolidates KV entries across token sets into a single compressed entry, further reducing KV cache size .
In plain English: the documents support a claim about KV cache size and attention compute. They do not, on their own, prove that every part of a DeepSeek V4 deployment uses 98% less VRAM.
The most direct appearance of the 98% figure in the provided sources is a LinkedIn user-generated article titled “DeepSeek Sparse Attention Shrinks KV Memory by 98 Percent in Real World Serving” . That can be a useful breadcrumb for tracking a claim, but it should not be treated as an official DeepSeek specification.
A more checkable third-party number is 10% KV cache. Wccftech reported that, compared with DeepSeek V3.2, DeepSeek V4 requires 27% of the single-token inference FLOPs and 10% of the key-value cache . Read literally, that suggests about a 90% KV-cache reduction under that comparison. It does not mean every context length, batch size, hardware setup, serving engine or full deployment will see a 90% total VRAM reduction .
Another headline describes DeepSeek V4 as having 9.5x lower memory requirements . Even as simple arithmetic, 1/9.5 is about 10.5% of the original requirement, or roughly an 89.5% reduction. That still is not 98%, and the scope still has to be checked: KV cache, a long-context benchmark, or total deployment memory .
| Claim | Evidence status | Better reading |
|---|---|---|
| Total VRAM is down 98% | Not supported by the official materials cited here | Do not use it as a procurement, capacity-planning or marketing specification |
| KV cache is heavily compressed | Supported by technical descriptions | CSA and HCA target KV entries in long-context inference |
| 10% KV cache | Third-party report | Roughly a 90% KV-cache reduction versus DeepSeek V3.2, not a total VRAM figure |
| 9.5x lower memory | Third-party headline | About an 89.5% reduction if interpreted literally, but the measurement scope still matters |
KV cache savings are highly relevant for long documents, long conversations and agent workflows. Hugging Face describes long-running agentic workloads where every tool result is appended to the context, and later tokens pay the attention cost against a longer and longer history; single-token inference FLOPs and KV cache size both grow with sequence length . The GitHub version of the same discussion describes common failure modes: traces exceeding the context budget, KV cache filling the GPU, and tool-call round trips slowing down long tasks .
But full VRAM planning includes more than KV cache. The LinkedIn post associated with the 98% framing itself separates shared weights, expert weights, activations, KV cache and framework overhead . That separation is the key point: even if KV cache falls sharply in a particular long-context scenario, the rest of the memory budget does not automatically fall by the same percentage.
DeepSeek V4’s direction is still important. CSA and HCA are aimed at one of the most expensive parts of million-token inference: attention over long sequences. NVIDIA’s description points to compression of KV entries, sparse attention matrices and consolidation of multiple token sets into compressed entries as ways to reduce KV cache size and computation .
DeepSeek’s V4 technical report also mentions infrastructure optimizations, including a single fused kernel for MoE modules designed to overlap computation, communication and memory access . Those are meaningful efficiency improvements. They are not, however, direct evidence for the much broader claim that total VRAM use falls by 98% .
If you are evaluating DeepSeek V4 for long-file analysis, long conversations or agent workloads, the practical question is not whether a viral percentage is exciting. It is whether your workload is actually KV-cache-bound.
Use your own context length, batch size, concurrency target, serving engine and hardware configuration. If KV cache is the limiting factor, V4’s compression-oriented attention design may be valuable. If your bottleneck is model weights, activations, framework overhead or serving concurrency, KV-cache savings will not automatically translate into the same percentage of total VRAM savings .
The short version: DeepSeek V4 appears to make long-context KV cache much cheaper. The public evidence provided here does not justify saying DeepSeek V4 uses 98% less total GPU memory .