Long sessions, scarce memory
A long-running agent works with 100,000 tokens of context or more, and the GPU memory that holds that context is in short supply. Our new paper keeps each session's KV cache in host memory and lends it GPU memory while there is room.
The fast memory stacked next to every GPU, HBM (high-bandwidth memory), is in short supply, and the next generation of hardware shows that this is not about to change. NVIDIA’s Rubin Ultra will carry 192 GB of HBM per GPU, less than first planned and less than the 288 GB of today’s B300.
Recent model architectures already take this into account. DeepSeek’s Engram lets about a quarter of DeepSeek-V4.1-Flash’s weights live in CPU memory, or even on SSD, instead of HBM.
The other large consumer of HBM is the KV cache: the keys and values a model computes for each token of a session, kept in GPU memory and reused at every step to produce the next token. It grows with every token, and agents run long sessions. That is where our new paper comes in.
Our approach keeps each session’s complete KV cache in host memory, the server’s CPU memory, lends it GPU memory while there is room, and takes that memory back when other sessions need it. We developed it into a complete component of a modern serving stack: it works with the latest speculative decoding drafters, DSpark and DFlash 2, and with disaggregated serving, where prompts and answers run on separate GPUs.
As GPUs get less memory for their compute and agent sessions keep growing, keeping every KV cache in GPU memory leaves more and more of that compute idle. Putting that compute to use gets more tokens out of every GPU, and with demand for tokens growing faster than the supply of compute, that is how serving keeps up. We expect this kind of memory management to become a standard part of long-context serving.
We start with what long sessions ask of GPU memory, then walk through how sparse attention lets most of a session’s history live off the GPU, what HiSparse already does with that, what we changed, and what we measured.
1. Long sessions need a lot of GPU memory
A language model writes one token at a time, and each new token depends on everything that came before it, so the KV cache grows with the session. It sits in HBM, the GPU’s own memory, which is very fast, expensive and limited, and the model’s weights take their share of it first.
Agents make sessions long. A coding agent rereads its whole workspace before each step, the files it opened, the outputs of its tools and its own earlier reasoning, and then writes a few hundred tokens. In a typical agent workload, a request carries 100,000 tokens of context or more. On GLM-5.3, a single 100,000-token session needs about 4.4 GiB of KV cache.
A GPU serves many sessions at once. At each step, it reads the model’s weights and applies them to every session in its batch, so the more sessions it advances together, the more tokens it produces for the same work. That only holds while every session’s KV cache fits in memory. Once memory is full, new requests wait in line, and the GPU’s output stops growing even though its compute could do more.
Figure 1 of the paper shows this on GLM-5.3 with 32,000-token prompts. Resident serving, the conventional setup that keeps every session’s KV cache in GPU memory, grows with the load until about 64 simultaneous requests. Doubling the requests to 128 then adds less than 5% more output, while the average wait for a first token goes from under a second to over a minute, and DFlash 2 speculative decoding does not move that limit. The same GPUs keep growing with our memory system: at 128 requests, Adaptive HiSparse, which uses it alone, delivers 1.9 times the output of resident serving, and Adaptive DHiSparse, which adds speculative decoding, 1.6 times that of resident serving with DFlash 2.
So with long sessions, a GPU runs out of memory well before it runs out of compute. When each session needs only a small part of its KV cache on the GPU, more sessions run together, and the same hardware delivers more tokens.
2. Sparse attention and HiSparse
What makes that possible is a change in how recent models read their KV cache. They use sparse attention: at each step, attention reads only a small subset of the KV cache, the entries most relevant to the next token. A small, fast component called the indexer picks them by scoring every earlier position and keeping the highest-scoring ones. GLM-5.3 picks 2,048 positions per step. DeepSeek-V4.1-Flash picks 512, about 0.1% of a 500,000-token session.
The rest still has to stay within reach, since the indexer may pick any position at the next step and the picks change every time. But it does not have to sit in HBM. Host memory, the server’s CPU memory, is a second pool next to it. It is cheaper per gigabyte and slower for the GPU to reach, which suits the part of the history the model is not reading right now. HiSparse (Xie et al., 2026) keeps every session’s complete KV cache there. On the GPU, each session keeps the compact index the indexer searches, so the model still considers its whole history, and a small hot cache of recently used entries. At each step, picked entries already in the hot cache are used in place, and the missing ones are copied from host memory just before attention needs them. The model computes exactly what it would have computed with everything on the GPU; only the location of the bytes changes.
Offloading the history to host memory frees most of a session’s GPU memory. On GLM-5.3 at 100,000 tokens, a session needs 4.44 GiB on the GPU when its whole KV cache stays there, and 0.43 GiB with HiSparse, about ten times less.
HiSparse leaves one decision to whoever deploys it: the size of the hot cache, the same for every session and every load. A small cache fits more sessions, but sends more reads to host memory, even at quiet times when the GPU has memory to spare. A large cache saves those reads, but fits fewer sessions when the server is busy.
HiSparse also does not support draft models such as DSpark and DFlash 2, which production inference stacks now rely on for speculative decoding. Adaptive DHiSparse lifts both limits.
3. Adaptive DHiSparse: lending the free GPU memory
The paper rests on one split of responsibilities: the model decides which entries to read, and the serving system decides where those entries are kept. Adaptive DHiSparse never changes the first. It changes the second continuously, as the load changes.
Each session keeps a working set that is never lent: its hot cache and the newest tokens it is writing. Whatever GPU memory is left over holds extra copies of the sessions’ history, which we call borrowed history. When the server is quiet, a session can borrow room for most or all of its history, so most of the entries the indexer picks are already on the GPU: in a load test on GLM-5.3, lending cut the share of picks fetched from host memory at low load from 14.2% to 1.4%. When only part of the history fits, one pass on the GPU sorts the picks out in order: the newest tokens and borrowed history first, then the hot cache, and host memory only for what is still missing. That search stays as small as the hot cache, however much history a session has borrowed.
Borrowed history is given back when it is needed. When a new session needs room, the running sessions return pages, each in proportion to its history. Nothing is lost: a page is backed up to host memory before it is returned, and if it is picked later, it is copied back like any other miss. When a session ends, its space goes to the others.
The same load test shows the loan being recalled: when requests jumped from 16 to 256, the share of history held on the GPU fell from between 75 and 100% to under 1% within 16 seconds, and climbed back above 50% a little over two minutes after the load dropped.
As a result, the choice HiSparse leaves to deployment goes away. The hot cache can stay at the small size that suits a busy server, and at quiet times, borrowed history fills the memory a larger cache would have taken. At low load, sessions run almost as if their whole KV cache were in GPU memory. At high load, they shrink toward the HiSparse working set, and the server fits more of them. That is what we were after: a memory system that retunes itself as the load changes, so an inference stack can leave it on.
4. Speculative decoding
Many inference stacks also use speculative decoding: a small drafter proposes several tokens, and the model checks them all in one pass, which yields more tokens per step whenever the guesses are right. Our runs pair DeepSeek-V4.1-Flash with DSpark and GLM-5.3 with DFlash 2. HiSparse was not designed with a drafter in mind, so it cannot run one: it keeps no KV cache for a drafter and looks up a single query per session at a time. Adaptive DHiSparse brings draft-model speculative decoding to this kind of serving. Speculation makes memory harder in two ways. Each drafted token picks its own KV entries, so one check reads more of the history. And the drafter’s own KV cache, plus the scratch space for the checks, takes memory that could hold another session. A memory system that stays on has to handle both.
One copy per missing entry. When several drafts pick the same entry and it is not on the GPU, it is copied from host memory once, and each draft still attends to exactly its own picks. In our counters, this avoided 29% of the copies from host memory during checks.
A bounded drafter cache. The drafter’s KV cache lives in a ring sized to what the drafter can see. GLM-5.3’s drafter looks back 2,048 tokens, so its ring holds 2,176 rows per session, however long the session grows.
Reused scratch space. The space that holds picks during a check is handed from one group of layers to the next once its readers finish. On GLM-5.3, four such areas replace 78, which cuts that part of the memory by 95%.
Rejected drafts never enter the history, and neither drafter changes: DSpark and DFlash 2 propose and accept tokens exactly as they were designed to.
5. What we measured
We measured both models with long prompts whose prefix was already cached, from a few simultaneous requests to hundreds, and compared Adaptive HiSparse and Adaptive DHiSparse with resident serving, where every KV cache stays in GPU memory, and with HiSparse. The GLM-5.3 runs used disaggregated serving, with separate GPUs for prompts and answers.
On GLM-5.3 with 100,000-token prompts and DFlash 2, resident serving stops growing at about 16 simultaneous requests, near 220 tokens per second per GPU, and at 128 requests, the average wait for a first token reached about four minutes. Adaptive DHiSparse delivered 793 tokens per second per GPU at 128 requests, with a 5.1-second average wait, and 826 at 256.
On DeepSeek-V4.1-Flash with 500,000-token prompts and 512 simultaneous requests, Adaptive DHiSparse delivered 5,054 tokens per second per decode GPU, against 2,738 for resident serving with the same DSpark speculative decoding. Without speculative decoding, Adaptive HiSparse delivered 4,217 against 1,853.
The gain grows with the context. On DeepSeek at 512 requests without speculation, with 32,000-token prompts, where memory is not yet the limit, the two serve about the same: 6,320 and 6,489 tokens per second per decode GPU. At 200,000 tokens, Adaptive HiSparse serves 57% more.
Against HiSparse, which cannot use speculative decoding, the difference is larger. On GLM-5.3 at 128 requests, Adaptive DHiSparse delivered 45% more, 793 against 548 tokens per second per GPU. Under the same latency limits for every configuration (99% of first tokens within 10 seconds once the initial burst has cleared, and 40 milliseconds per token on average), it handled four times as many simultaneous requests as HiSparse, with 3.7 times the output.
Without speculative decoding, Adaptive HiSparse delivered 19 to 23% more than HiSparse at light load and 13 to 14% more at 128 and 256 requests on GLM-5.3, and 9% more on DeepSeek at 512 requests, where the first token came after 20 seconds instead of 28.
At 512 requests on DeepSeek with DSpark, requests wait 25 seconds for their first token, against 80 with resident serving.
6. What it costs
At light load, there is a small price. Without speculative decoding, adaptive serving was 3 to 7% slower than resident serving while every KV cache still fit: 3.6% on DeepSeek at 16 requests, about 5 to 7% on GLM-5.3. With DFlash 2 on GLM-5.3, the gap was about 9 to 10% at 4 and 8 requests, while with DSpark on DeepSeek, adaptive serving came out ahead even at 16.
Speculation has its own trade-off. Its extra memory leaves room for 140 GLM-5.3 sessions instead of 208. At 256 requests, Adaptive DHiSparse and Adaptive HiSparse delivered about the same, 826 and 823 tokens per second per GPU, within our run-to-run spread, but requests waited 65 seconds for their first token instead of 34, almost all of it queuing for one of the 140 slots. We also picked DFlash 2’s block size by hand for each load (8 up to 32 requests, 6 at 64, 4 at 128 and 256).
Our runs cover one deployment per model, with prompts whose prefix was already cached. What does not change is the computation: every step attends to the same entries, with the same bytes, as it would with everything on the GPU. The paper states the exact conditions.
7. What comes next
Working with scarce GPU memory is part of the job now, and we have come to enjoy it. Some of the most inventive work in AI today comes from that constraint. Sparse attention lets a model read only the part of a long history that the next token needs. DeepSeek-V4.1-Flash shares its KV cache across layers, and Engram moves part of its weights out of HBM altogether. Drafters such as DSpark and DFlash 2 can get several tokens out of each step. Each of these gets more out of the same GPUs, and they go further together when models and serving systems are built with each other in mind.
Adaptive DHiSparse is our contribution from the serving side. We hope to see more models built on the architectures we love most, such as sparse attention and cross-layer sharing of indexes and KV caches. They scale well to long sessions and let every GPU deliver more tokens. The same ideas may also help local machines, where GPU memory is tighter still and system RAM is often larger.
The full paper is coming to arXiv soon.
Sources
- Rubin Ultra and HBM per GPU: SemiAnalysis, Long live the short king: why 4-hi HBM wins, September 2026.
- Engram: Cheng et al., Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models, 2026.
- HiSparse: Xie et al., HiSparse: Scaling sparse-attention decoding with hierarchical KV cache management, 2026.
- The models: DeepSeek-AI, DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, 2026; Z.ai, GLM-5.3, 2026.
- Speculative decoding: Cheng et al., DSpark: Confidence-scheduled speculative decoding with semi-autoregressive generation, 2026; Inco AI, DFlash 2: Keep drafting parallel, 2026.
- Our earlier posts: The token gap, on why serving has to get more tokens from every GPU, and Eleven models in eight weeks, for the KV cache size of each model we serve.