Introduction
Every coding agent works the same way underneath. At each step it sends the model everything it has so far, the system prompt, the tool definitions, every file it opened and every command it ran, followed by one new tool result, and the model reads all of it to choose one small next action.
In the session we price later in this piece, the context starts at 12,000 tokens and ends at 85,500 after fifty steps, and the model is sent 2.44 million input tokens in order to write 20,000.
Almost no one of that is billed at full price. The provider has seen those tokens before, has kept what it computed the first time, and charges the repeat at a fraction of the normal rate. On Anthropic’s current price list the fraction is 10 percent for most models, 5 percent for Opus 5.5 and 2.5 percent for Fable 5.1 and Mythos 5.1. On DeepSeek it is 2 percent.
The discount comes with a clock: five minutes, thirty minutes, an hour or a day, depending on the provider and on what you pay.
I wanted to know where those tokens go between two calls, who pays to keep them, and why the numbers are what they are. Why a tenth and not a half. Why five minutes and not five hours. Why DeepSeek can charge a fiftieth while Google charges rent by the hour.
We ended up with one ratio that answers most of these questions, a replay of public production traces that answers most of the rest, and a view of where inference hardware is heading that neither of us expected when we started. It also produced a handful of rules that, in the agent session we model, cut the bill in half.
The short version
What is cached. Not your text: the key and value tensors every attention layer computes for every token. Depending on the architecture that is 890 bytes to 512 KiB per token, so a 100,000-token agent context occupies anywhere from 89 MB to 52 GB.
Why a read costs a tenth. A hit costs the provider a reload and some rent, not a recompute, and on an H100 even a single SSD returns cache faster than the GPU can regenerate it for every model here released since 2024. How cheap a read can be depends on how much compute each stored byte saves, which varies 700-fold across the models we measured; DeepSeek’s V4.1-Flash saves the most per byte and charges 2 percent.
Why five minutes. GPU memory is the most expensive place to keep anything: an H100 fills its own HBM with fresh cache in about a minute on a 70B-class model, host memory pays for roughly an hour and SSD for days. Five minutes is also where most reuse happens; in Kimi’s production traces it captures 79 percent of the possible reuse for chat and 93 percent for agents.
Why writes cost extra. Keeping bytes is rent, but only for entries nobody reads again, because every read refreshes the lifetime for free. Converted to dollars per gigabyte-hour, most providers price that rent like GPU memory, far above what DRAM costs.
What to do about it. Caching everything pays only if more than 22 percent of the tokens that pass through the cache are hits (53 percent with the one-hour lifetime). Gaps of 5 to 36 minutes are cheapest to cross with a zero-output ping every four and a half minutes, longer ones with the one-hour lifetime. In our fifty-step agent session with nine long tool calls, that is the difference between $1.88 and $0.97.
Where it is going. A Rubin GPU fills its own memory with fresh cache 2.4 times faster than a B200. That is why NVIDIA is adding petabytes of flash to every pod and why DeepSeek shrank its cache 437-fold in three years.
The price list in September 2026
Four providers sell the same physical thing in three different ways.
Anthropic charges a premium to write the cache and a small fee to read it. A write costs 1.25 times the base input price if the entry should live five minutes and 2 times if it should live an hour. A read costs 0.1 times the input price on most models, 0.05 on Opus 5.5 and 0.025 on Fable 5.1 and Mythos 5.1.
Every read refreshes the lifetime at no charge, and the clock starts when the request that wrote or read the entry begins, not when its response finishes, so a response that streams for four minutes leaves about one minute for the follow-up. Anthropic also states that the cached key and value tensors are held in memory only and never stored at rest.
OpenAI has moved to the same structure. From GPT-5.6 onward a write costs 1.25 times the uncached input rate, a read 0.1 times, and an entry stays eligible for at least thirty minutes after its latest write or reuse.
Earlier models charged nothing to write and kept entries in what OpenAI describes as volatile GPU memory for roughly five to ten minutes of inactivity; an extended policy kept them for up to 24 hours by offloading the key and value tensors to storage local to the GPU once memory filled up. The two companies differ in one detail that matters for agents at scale: Anthropic does not count cache hits against rate limits, and OpenAI does.
Google separates reading from keeping. Reading an explicit Gemini cache costs a tenth of the input price, and keeping the cache alive costs rent: $4.50 per million tokens per hour on Gemini 3.1 Pro and $0.50 on Gemini 3.8 Flash, a rate the price page schedules to double on January 1, 2027.
Gemini also caches without being asked: according to Google’s caching guide, implicit caching is on by default for Gemini 2.5 and newer models, discounting repeated prefixes of at least 4,096 tokens on the current 3.x models, with no storage fee and no guarantee of a hit.
DeepSeek charges for neither writing nor keeping. A hit on its current flash model costs $0.003 per million tokens off-peak against $0.15 for a miss, both doubling during two weekday windows. DeepSeek has offered context caching since August 2024 and serves its hits from a cache on disk.

Read one way, the market has converged: a read at a tenth of the input price is the default at Anthropic, OpenAI and Google. Read another way, it has already broken below that point, with the newest models from Anthropic and DeepSeek at 5, 2.5 and 2 percent. The charges for writing and keeping have not converged at all. Both patterns come from what the provider is physically holding.
What the provider keeps
A cached prompt is not text. When a transformer processes a token, every attention layer computes a key vector and a value vector for it, and every later token in the sequence reads them. Keeping those vectors lets the next request with the same beginning skip computing them again. OpenAI’s documentation says it directly: the cache holds key and value tensors, not the tokens themselves.
How many bytes that is depends on the attention design, and the range is wide. With grouped-query attention, the design used by Llama 3, Qwen3, GLM and MiniMax, every token costs two vectors per layer per key-value head:
bytes per token = 2 × layers × kv_heads × head_dim × bytes per element
Llama-3.1-70B: 2 × 80 × 8 × 128 × 2 (BF16) = 327,680 bytes = 320 KiBA 100,000-token agent context for that model is 32.8 GB, two fifths of an H100’s memory for a single conversation. Multi-head latent attention, which DeepSeek introduced in V2 and Moonshot reused in Kimi K2, stores one compressed latent per layer instead of per-head keys and values: 512 values plus 64 for position.
In BF16 that is 1,152 bytes per layer and 70,272 bytes per token across DeepSeek-V3’s 61 layers, so the same 100,000 tokens take 7.0 GB. DeepSeek-V3.2 adds a small FP8 key per token for its sparse-attention indexer but stores the latent itself in FP8, which by our layout accounting comes to 48,068 bytes.
The newest designs compress across tokens as well. DeepSeek’s V4 family keeps compressed and heavily compressed attention caches, about 3,600 bytes per token for V4-Flash by our reconstruction in August, and DeepSeek’s own report puts V4-Flash at 7 percent of V3.2’s cache at a million tokens.
V4.1-Flash, released on September 10, goes further: its decoder layers take their global cache from a projection of the final encoder state instead of computing their own, and the cache is stored in FP4 with one scale per 16 channels. The model card gives 890 bytes per token. A 100,000-token context is 89 MB.

Two things in that chart are easy to miss. The first is that cache size does not follow model size. Llama-2-7B, with plain multi-head attention, keeps 512 KiB per token, more than Llama-3.1-70B; MiniMax-M2, with 10 billion active parameters, keeps 248 KiB because all 62 of its layers use full attention with eight key-value heads; GLM-4.6 keeps 368 KiB across 92 layers.
The second is how far one lab has pushed. DeepSeek’s first model, the 67B release of late 2023, kept 389,120 bytes per token: 95 layers, eight key-value heads of dimension 128.
Divide by 890 and the result is 437.2, which is the 437-fold reduction DeepSeek reports in the V4.1 model card. We computed the V1 figure from its architecture and got the same ratio, a useful check on both numbers.

Sliding windows and linear attention trade differently. gpt-oss-120b keeps full history in only 18 of its 36 layers and a 128-token window in the rest; Qwen3-Next keeps full attention in 12 of 48 layers and a fixed-size recurrent state in the others.
Both shrink the per-token cache, and both add a per-sequence state that can only be reused at the exact position where it was saved, which makes prefix matching harder than it is for a paged cache.
Compute saved per byte
The size alone does not say whether keeping the cache is worth anything. What matters is the work the bytes save. Recomputing a token’s cache means running prefill for it again, which costs at least two floating-point operations per active parameter. Call that F and divide it by the bytes kept:
ρ = F / k prefill FLOPs avoided per byte of cache kept
F = 2 × active parameters in prefill (a floor: attention adds more at long context)For Llama-3.1-70B, F is 141 GFLOP per token and k is 327,680 bytes, so ρ is 431 thousand: every byte kept saves 431 thousand operations. DeepSeek-V3 is at 1.05 million, V4-Flash at 7.2 million, and V4.1-Flash, whose prefill runs through only 8 billion active parameters, at 18 million.
At the bottom are Llama-2-7B at 26 thousand and MiniMax-M2 at 79 thousand. Across the fifteen models in this piece ρ spans a factor of about 700.

Mixture-of-experts made caching relatively less valuable before DeepSeek made it more valuable. Sparsity cuts compute per token without touching attention, so a model like Qwen3-235B-A22B or MiniMax-M2 carries the cache of a large dense model and the prefill cost of a small one.
DeepSeek compressed attention in the same generations in which it sparsified the experts, and every release since V2 sits further toward the upper left of the chart.
Our F is a floor. With dense attention, recomputing a long prefix also pays for attention itself, which grows with position. Averaged over a 64,000-token prefix, attention adds about 84 GFLOP per token to Llama-3.1-70B’s 141, and about 160 GFLOP to DeepSeek-V3’s 74 when V3 runs prefill without absorbing its projections.
At agent context lengths the real ρ of dense-attention models is 1.6 to 3.2 times the floors we use, so every conclusion below gets stronger when those terms are added back.
How fast a GPU really writes cache
Put ρ next to a GPU and you get the quantity that ties the rest of the piece together. A GPU running prefill at peak throughput P and utilization u produces fresh cache at the rate
G = P × u / ρ bytes of new cache per secondWe use dense FP8 peak for P, 1,979 TFLOPS on an H100 and 4.5 PFLOPS on a B200, and u = 0.30. That is at the optimistic end of what has been published, so it is worth calibrating. SGLang’s open reproduction of DeepSeek’s serving system processed 52,300 input tokens per second per eight-GPU H100 node on 2,000-token prompts, which is about 0.26 of the FP8 peak once attention is counted.
On GB200 the same team reached 18,471 input tokens per second per GPU with FP8 experts, 0.29 of dense FP8 peak, and 26,156 with NVFP4 experts. DeepSeek’s own production prefill nodes, averaged over a day of real traffic, processed 73,700 input tokens per second per node including cache hits; removing the 56.3 percent that were hits leaves about 4,000 computed tokens per second per H800, roughly 0.16 of peak.
We keep 0.30 as the central value and carry 0.15 to 0.5 in every range below. On an H100 at our assumption, Llama-3.1-70B writes 1.38 GB of cache per second, DeepSeek-V3 0.56 GB, V4-Flash 83 MB and V4.1-Flash 33 MB.
G has one property worth stating separately. Move a model’s arithmetic from BF16 to FP8 and P doubles; store its cache in FP8 as well and k halves; G does not move.
Quantizing both sides together leaves the rate at which a GPU fills memory where it was, the same kind of precision invariance we actually found for the critical batch size in August, seen from the memory side. Quantizing only the arithmetic doubles it.
G answers three different questions. It is the break-even bandwidth for reloading, because any link faster than G delivers cached bytes faster than the GPU could regenerate them.
It sets the capacity bill for a lifetime, because holding everything a GPU produces for τ seconds takes G × τ bytes. And it sets the clock for the GPU’s own memory, because HBM capacity divided by G is how long the GPU takes to fill its HBM with fresh cache.
Reload or recompute
The first question has a lopsided answer. At 1.38 GB/s, a single Gen5 NVMe drive, which reads at about 14 GB/s, already returns Llama-3.1-70B’s cache faster than an H100 can recompute it, and every link above it in the hierarchy is faster still.
For the MLA and compressed models the margin over a single drive is 25 to 420 times. Only models at the bottom of the ρ scale need more than a drive: MiniMax-M2 produces 7.5 GB/s of cache on an H100 and about 67 GB/s on a Rubin GPU, a full PCIe Gen5 x16 link.

The latency tells the same story. Reloading a 100,000-token Llama-3.1-70B context, 32.8 GB, over a Gen5 x16 link at an effective 55 GB/s takes circa 0.6 seconds; recomputing it on eight H100s at our utilization takes about 3 seconds before attention is counted.
For V4.1-Flash the context is 89 MB and the reload takes a couple of milliseconds. Recomputation wins only where compute per byte is low and prefixes are short. A September characterization of SSD-backed caching for vLLM put the break-even at about 6,200 tokens for Qwen3-4B on its H100 setup.
Hiding the reload
A reload does not have to be waited on. The engine can load layer l + 1 of the cached prefix while it computes layer l of the new tokens, which is how the layer-wise pipelines in Mooncake and the systems that followed it work.
The reload disappears behind the computation when the link moves the cached bytes in no more time than the GPU spends on the new ones:
L_cached × k / W ≤ L_new × F / (P × u) which gives W ≥ (L_cached / L_new) × GThis is where agents are the hard case. A late agent step carries a long cached prefix and a short new suffix: 84,000 tokens cached and 1,500 new in the last step of our session, a ratio of 56.
Hiding that reload for Llama-3.1-70B on an H100 needs 77 GB/s, more than a Gen5 x16 link delivers, although attention over the long prefix adds compute to the new tokens and cuts the requirement by more than half.
DeepSeek-V3 needs 32 GB/s and V4.1-Flash under 2 GB/s. The ratio of cached to new tokens is exactly what an agent pushes up, which is why coherent CPU links and small caches matter more for agents than for chat.
The clock
Reloading is relatively cheap. Keeping is certainly not free, and it is paid for by the second. An entry is worth keeping as long as the expected saving from a future hit exceeds the rent. Recomputing one byte of cache costs the GPU’s price per second divided by G, and holding one byte costs the tier’s price per byte-second.
The holding time at which the two are equal is the break-even:
τ* = (GPU cost per second) / (G × s) s = holding cost of the tier per byte-second
τ*_HBM = H / G when HBM is priced at its share of the GPUThe HBM line needs one assumption. When serving is limited by memory, the normal state of a decode-heavy fleet, every gigabyte of HBM holding an idle cache entry is a gigabyte not holding an active request, so its price is its share of the GPU’s rent.
With that price the break-even reduces to something physical: the time the GPU takes to fill its own HBM with fresh cache. On an H100 serving Llama-3.1-70B that is 58 seconds. H/G is an upper bound, because part of HBM holds the weights: about 11 percent for Llama-3.1-70B spread over eight H100s, about 60 percent for a 671-billion-parameter FP8 model on eight H200s. The break-even shrinks in the same proportion.
For the other tiers we priced memory at 2026 levels, which are not those of a year ago. TrendForce expected conventional DRAM contract prices to rise 55 to 60 percent in the first quarter alone; by late August a 64 GB DDR5 server module cost 2.5 to 3 times its January price and a 3.84 TB enterprise NVMe drive 2.3 to 2.8 times, and TrendForce still expected both to rise in the third quarter.
We tend to assume $12 per GB of DRAM over four years, with 50 percent added for power, space and the CPU socket it hangs from, and $0.25 per GB of SSD over five years, doubled for the servers and network around it.
For the GPU we use an H100 at $2.50 an hour, inside the range specialist clouds charge in 2026.
Per byte-hour, HBM, DRAM and SSD come out at roughly 2,700 to 45 to 1. Each step down the ladder buys 45 to 60 times more holding time for the same money.

For Llama-3.1-70B the ladder reads about a minute in HBM, about an hour in DRAM and about two days on SSD. For DeepSeek-V3 it is two and a half minutes, two and a half hours and four and a half days. For V4.1-Flash it is 40 minutes, 41 hours and 77 days. MiniMax-M2 gets 11 seconds, 11 minutes and 8 hours.
Not a single one of these numbers is totally accurate. Across the ranges we consider plausible (utilization from 0.15, what DeepSeek’s production fleet achieves, to 0.5, an H100 from $1.50 to $4.00 an hour, the low and high ends of 2026 memory quotes, and 60 to 100 percent of HBM available to cache), the HBM break-evens move by a factor of 0.36 to 2.0, the DRAM ones by 0.2 to 5.8 and the SSD ones by 0.17 to 7.1. The whiskers in the chart show those ranges; the order of the tiers and of the models never changes.
Now lay the product menu over it. Five minutes is past the HBM break-even of every grouped-query model in our set even at the top of its range, and inside the DRAM break-even of all of them at our central assumptions, which says a five-minute cache for those models belongs in host memory rather than on the GPU.
One hour sits at the DRAM break-even of a 70B-class grouped-query model: 59 minutes at the center, between 12 minutes and almost 6 hours across the ranges. At the center, twenty-four hours is past the DRAM break-even of every model here except V4.1-Flash, and pays on SSD only for models above roughly 230 thousand FLOPs per byte, which leaves out most of the sparse grouped-query models.
The providers’ own descriptions match the ladder: OpenAI keeps short-lived entries in GPU memory and moves to local storage for the 24-hour policy, DeepSeek serves hits from disk, and Google charges rent for anything kept on request.
These are break-evens for an entry that is certain to be read again. With a probability p of reuse, each one shrinks by p. The probabilities are what the traces are for.
What the traces say
Moonshot published request traces from Kimi’s serving system with the Mooncake paper, which won the best paper award at FAST 2025. Each request carries an arrival time, its input and output lengths, and one hash per 512-token block, remapped so that equal hashes mean reusable cache.
The conversation trace holds 12,031 requests and the tool-and-agent trace 23,608, both sampled from one hour of production traffic. Average inputs are 12,035 and 8,596 tokens and average outputs 343 and 182, ratios of 35 and 47 to one.
We replayed both traces block by block against a cache with unlimited space whose entries expire a fixed time after their last use. With no expiry at all, 37.4 percent of the conversation trace’s input tokens and 57.1 percent of the agent trace’s could have come from cache. Those ceilings sit where others have found them.
Alibaba’s study of its production traces found ideal hit ratios of 62 and 54 percent over a day, and in its API trace single-turn requests produced 97 percent of the hits. DeepSeek reported that 342 billion of the 608 billion input tokens it served in one day of February 2025, 56.3 percent, hit its on-disk cache.

A five-minute lifetime captures 79 percent of the ideal reuse in the conversation trace and 93 percent in the agent trace. Ten minutes captures 94 and 99 percent. Thirty minutes captures essentially all of it. Most of the value is in the first five minutes, and by thirty there is little left to buy.
The gaps between reuses show where the two workloads differ. In the agent trace about half of all reused tokens were reused within one second of their previous use (52 percent over the whole trace, 50 percent in the edge-corrected sample below): requests arriving together and sharing a long prefix, the signature of an agent fanning out parallel calls.
In the conversation trace that share is about 10 percent, and the median gap is about two minutes, roughly the time a person takes to read an answer and type the next question.

Fan-out is also why Anthropic’s documentation notes that an entry becomes available only after the first response starts, and suggests waiting for it before sending parallel requests. Ten identical prefixes sent in the same second are ten cache writes, not one write and nine reads.
The trace covers one hour, which biases every gap statistic toward short gaps, so for the chart above we counted only blocks whose previous use came at least thirty minutes before the end. Among those, 75 percent of the conversation reuse and 91 percent of the agent reuse that happens within half an hour happens within the first five minutes.
Capacity is the other half of the question. An LRU cache sized to hold ten minutes of the trace’s unique traffic captures 89 percent of the ideal reuse for conversations and 98 percent for agents, and twenty minutes captures 98 percent of the conversation ideal.
The unique traffic arriving at a prefill GPU is exactly what that GPU computes, which is G. So the capacity a ten-minute window needs per GPU, with prefill running flat out, is G × 600 seconds: 827 GB for Llama-3.1-70B on an H100, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash.
An H100 has 80 GB of HBM, and an eight-GPU DGX H100 has 2 TB of system memory, about 256 GB per GPU. A ten-minute window for a 70B grouped-query model fits in neither, and has to live on SSD or in a pool shared across machines. For DeepSeek-V3 it spills just past DRAM. For V4.1-Flash it fits in HBM next to the weights. That comparison goes a long way toward explaining why DeepSeek can nearly give reads away and charge nothing for keeping.
It also explains why nobody lets you write everything for free. If every miss in the conversation trace were written at 1.25 times the input price and every hit read at 0.1, a five-minute cache would cut the input bill by 9 percent; on the agent trace the cut would be 36 percent.
The margin is in what gets written, which is why both Anthropic and OpenAI let callers place breakpoints, and why OpenAI’s new explicit mode charges no write for anything after the last breakpoint.
Where the bytes live
NVIDIA describes the hierarchy its Dynamo framework manages in four levels: G1 is GPU memory for the cache of requests in flight, G2 is host DRAM for staging, G3 is SSD inside a node and G4 is shared storage for what has to survive.
At CES in January it added a level between the last two. The CMX context memory storage platform, first shown as ICMS, puts Ethernet-attached flash managed by BlueField-4 processors into each pod, with petabytes of shared capacity, and NVIDIA claims five times the tokens per second and five times the power efficiency of general-purpose storage for this job. Its own description of the data is the best argument for a separate tier: transient, derived, and recomputable if lost.
The rest of the rack is moving the same way. A GB200 superchip pairs two Blackwell GPUs with a Grace CPU carrying up to 480 GB of LPDDR5X. A Vera Rubin NVL72 rack carries 54 TB of LPDDR5X, 1.5 TB per Vera CPU and 750 GB per GPU, reachable over a coherent NVLink-C2C link of 1.8 TB/s.
The two labs that price hits lowest built their own versions of this years ago. Mooncake pools the DRAM and SSD of Kimi’s GPU servers into one distributed cache; the paper reports more than 100 billion tokens a day on thousands of nodes, and 115 and 107 percent more requests on A800 and H800 clusters than the previous system.
DeepSeek’s 3FS file system, built on SSDs and RDMA, reached 6.6 TiB/s of aggregate reads on a 180-node cluster, and its README shows cache reads for inference peaking at 40 GiB/s.
Flash has one limit that bandwidth arithmetic hides, which is wear. A prefill GPU that persisted everything it computed would write G bytes per second to flash, all day. For Llama-3.1-70B on an H100 that is 119 TB a day; a 30 TB drive rated for one full write per day absorbs 30. Endurance alone would call for four drives per GPU before capacity or bandwidth enter the picture.
For DeepSeek-V3 the figure is 49 TB a day, for V4.1-Flash under 3 TB. At these write rates, admission control, writing only the blocks likely to be reused, decides whether the tier outlives its warranty.
The V4.1 model card makes the same point from the model side: its bounded replay of the sliding-window layers exists so that their cache never has to be persisted to SSD, and it cuts the persistent footprint to about an eighth of V4-Flash’s.
Rubin and the context tier
The part we did not expect is what the next GPU does to all of this. The HBM break-even is H/G, and G is proportional to peak compute, so the clock for GPU memory is set by a single hardware constant: HBM capacity per unit of compute.
In megabytes of HBM per dense FP8 TFLOPS it is 40.4 on an H100, 71.2 on an H200, 40.0 on a B200 and 37.2 on a GB200. On Rubin, with 288 GB against 17.5 dense FP8 PFLOPS on NVIDIA’s specification page, it is 16.5. NVIDIA revised that page on September 21, lowering HBM bandwidth to 19.2 TB/s and NVLink to 3 TB/s per GPU; neither enters this ratio, and capacity and compute are unchanged.
That one constant moves the whole ladder. For Llama-3.1-70B the time to fill HBM with fresh cache falls from 58 seconds on an H100 to 24 on Rubin, for DeepSeek-V3 from 142 to 58, for V4.1-Flash from 40 minutes to 16.
The capacity a ten-minute window needs per GPU grows with compute instead: 7.3 TB for Llama-3.1-70B, 3.0 TB for DeepSeek-V3 and 175 GB for V4.1-Flash, against 288 GB of HBM and 750 GB of LPDDR5X per Rubin GPU.


A Rubin GPU produces cache faster than anything in its own rack can hold for more than a minute or two, unless the model’s cache is small. Its share of the rack’s host memory holds about a minute of Llama-3.1-70B’s output and about two and a half minutes of DeepSeek-V3’s; for V4.1-Flash it holds 43 minutes.
The flash tier in the pod is the hardware answer to that gap, and a 437-fold smaller cache is the model answer. They are substitutes, and different companies are pursuing them.
The same arithmetic casts a different light on a product that disappeared. In September 2025 NVIDIA announced Rubin CPX, a GPU with 128 GB of GDDR7 built for the compute-heavy prefill of long contexts and due at the end of 2026. It was missing from the roadmap at GTC in March.
Commentators have pointed to tight 3nm capacity and the price of GDDR7 during the memory shortage. Our reading, which is an interpretation and nothing more, is that caching also shrank the market the chip was designed for.
When half or more of the input tokens in agent traffic are cache hits, the prefill that remains is the new part of each request, and a pool of flash that keeps the rest of the context close does more for time to first token than a pool of extra prefill compute.
How the lookup works
No one of this helps unless a request can find its cache, and finding it is stricter than it looks. Attention is causal: the cache for position i depends on every token before i. Change one token near the start of the prompt and every block after it has a different cache, even where the text afterward is identical.
So every system identifies cache by prefix. vLLM hashes each block of tokens together with the hash of the block before it, adding extra keys for images, adapters and a per-tenant salt, and a lookup walks that chain until the first miss.
The API products expose the same chain at a coarser grain. Anthropic writes an entry only at a breakpoint, as a cumulative hash of everything before it, and a read walks back at most 20 blocks looking for an entry an earlier request wrote; it will not discover stable content behind a changing block unless something wrote an entry there.
Minimum cacheable lengths run from 512 tokens on its newest models to 4,096 on Haiku 4.5. OpenAI’s GPT-5.6 has a 1,024-token minimum, places an implicit breakpoint at the end of the latest eligible message, allows up to four writes per request, and checks the first two and the latest fifty explicit breakpoints on a lookup. Gemini’s implicit cache starts at 4,096 tokens on its current 3.x models.
The minimums reflect bookkeeping. Every entry carries a hash, an index record and a routing decision, and below a few hundred tokens that overhead is comparable to recomputing. OpenAI’s guide works through the trade from the customer’s side and shows that padding a short shared prefix up to the 1,024-token minimum pays off once the prefix is longer than about 102 tokens and reused often enough.
Then the request has to arrive where the cache is. OpenAI’s cache lives on individual machines, and requests are routed by a hash of the initial tokens after its hidden system content plus an optional key; above about 15 requests per minute for one prefix and key, some requests overflow to machines that do not hold the entry.
Open-source stacks solve the same problem with cache-aware routers that trade some load balance for hits. For agents the difficult moment is the pause: while a tool runs, the engine is tempted to evict the conversation’s cache to make room for someone else.
Continuum, from Berkeley, pins the cache in GPU memory for a time-to-live predicted from the tool’s expected duration, and reports large reductions in job completion time on SWE-Bench and BFCL agent workloads.
Shared caches also leak. A hit is faster than a miss, and in audits run in September and October 2024, Stanford researchers detected prompt caching at 8 of 17 API providers and sharing across users at 7 of them, OpenAI among them.
The large providers now isolate: OpenAI by organization and processing region, Anthropic by workspace on its own API. OpenAI still recommends a separate cache key per end user inside one organization, which stops one customer from probing for another’s cached prefixes.
What the prices imply
With the physics in place, the price lists can be read as statements about cost. Start with the question a developer can answer from them directly:
how many times must a cached prefix be read before caching beats resending it?
With a write multiplier w and a read multiplier r, writing once and reading n times costs w + n·r input-equivalents, and not caching costs 1 + n.
break-even reads: n > (w - 1) / (1 - r)
5-minute write (w = 1.25, r = 0.1): n > 0.28 one read wins: 1.35 against 2.00
1-hour write (w = 2.00, r = 0.1): n > 1.11 one read loses: 2.10 against 2.00A five-minute write pays for itself with one read and an hour-long write needs two. The second read is the price of the longer clock, and the hourly cost of holding can be backed out of each price list. Anthropic’s five-minute premium is a quarter of the input price for five minutes, three times the input price per hour; its one-hour premium is one times the input price per hour.
OpenAI’s is a quarter for at least thirty minutes, at most half the input price per hour. Google’s rent is 2.25 times the input price per hour on Gemini 3.1 Pro and two thirds of it on Gemini 3.8 Flash. DeepSeek’s is zero.
Those hourly figures describe an entry that is written and never read again. An entry that keeps being read costs nothing more to keep, because at both Anthropic and OpenAI every read refreshes the lifetime without a new write charge.
For an active agent the write premium is a deposit, not rent. The provider charges for the chance that you walk away and gives the holding away to anyone who comes back, which is the pricing you would design if idle bytes in fast memory were the cost you worried about.
Converting those charges to dollars per gigabyte-hour requires the cache size per token of the model behind each price, and none of these providers publishes it for its closed models. So we plotted every charge against every plausible size.

Across the cache sizes of the modern open models in this piece, from V4.1-Flash’s 890 bytes to GLM-4.6’s 368 KiB, every one of these charges is at least two and a half times our DRAM cost, and most are more than ten times it. For caches under 64 KB per token, all of them except Gemini Flash’s are at or above what HBM costs to rent.
Sonnet 5’s hourly rate, $2 per million tokens, equals H100 HBM rent at 64 KB per token and exceeds DRAM cost for any cache under 3.9 MB per token, a size no production model approaches. None of the providers prices its clock at DRAM cost, whichever tier the bytes actually sit in.
DeepSeek is the one provider for which both sides of the calculation are public. At 890 bytes per token a million cached tokens occupy 0.89 GB, and holding them on SSD for a full day costs about $0.0002 at our rates. Recomputing them costs at least 16 PFLOP, 27 seconds of an H800 at our utilization, about $0.015 at $2 an hour.
The off-peak read price, $0.003, is a fifth of that recompute floor and about twelve times a day of SSD rent; the miss price, $0.15, is ten times the recompute floor. Nineteen months earlier DeepSeek charged $0.14 for an R1 hit against $0.55 for a miss: 25 percent, on a model that kept 79 times more cache per token.
We cannot see inside Anthropic’s models, so we cannot say why a read on Fable 5.1 costs a quarter of what it costs on Sonnet 5 as a share of the input price. The arithmetic says a read gets cheaper to serve as the cache per token shrinks relative to compute, and a factor of four is the size of step DeepSeek’s generations have taken.
It is also what a company would do to win agent workloads, where reads are most of the bill. The two explanations are not exclusive.
An agent session, priced
Here is what the clock does to a real bill. Take a coding agent that starts with a 12,000-token system prompt and tool catalog and takes fifty steps. Each step appends 1,100 tokens of tool output and 400 tokens of model output.
The context ends at 85,500 tokens; the session sends 2.44 million input tokens and receives 20,000 output tokens, a ratio of 122 to one. Manus, which runs this kind of loop at scale, reports a ratio around 100 to one and names cache hit rate as the metric it watches above all others.

Without caching the session costs $5.08, almost all of it input. With a five-minute cache and every tool call finishing inside the window it costs $0.88: the session reads 96.5 percent of its input tokens from cache, its effective input price falls to 14 percent of list, and 69 percent of what remains is reads.
Now let nine of the forty-nine gaps last eight minutes: the test suite, the build, the person who went for coffee. Each time the cache is gone, and the entire context is written again at 1.25 times the input price. The session costs $1.88, more than double. Writing everything with the one-hour lifetime brings it to $1.01; keeping the five-minute entries alive with one zero-output ping per gap brings it to $0.97.
The same session on Opus 5.5, with reads at 5 percent, costs $1.30 against $10.15 uncached. On DeepSeek’s V4.1-Flash off-peak it costs 3.2 cents against 38 cents. Because reads dominate an agent’s bill, the read multiplier matters more than any other number on the price list: halving it from 0.1 to 0.05 at Sonnet 5’s base price would cut the input side of this session by 34 percent.
Price lists also differ in structure, so the same session lands at very different fractions of each provider’s own input price. With a warm cache it runs at 0.14 of list on Sonnet 5 and GPT-5.6, 0.13 on Gemini’s implicit cache if every repeat hits, 0.09 on Opus 5.5, 0.07 on Fable 5.1 and 0.05 on DeepSeek’s V4.1-Flash.
The eight-minute gaps separate them further. GPT-5.6’s thirty-minute minimum absorbs them, and DeepSeek publishes no lifetime but serves hits from disk, so we actually assume its cache survives them too. Anthropic’s five-minute models need pings or the one-hour write to stay near their warm figures; left to expire, Sonnet 5 climbs to 0.34, Opus 5.5 to 0.30 and Fable 5.1 to 0.29. A lower read price does little for a session whose cache keeps expiring.

A field guide
The arithmetic above turns into rules that can be applied without redoing it. Each comes with the number behind it, so it can be rechecked when a price list changes.
When caching pays at all
If every input token either writes to the cache or reads from it, the input bill as a multiple of the uncached bill is (1 - h)·w + h·r, where h is the share of those tokens that are hits. Caching pays when that is below one:
break-even hit rate: h* = (w - 1) / (w - r)
5-minute write (w = 1.25, r = 0.10): h* = 21.7%
1-hour write (w = 2.00, r = 0.10): h* = 52.6%
Opus 5.5, 5-minute (r = 0.05): h* = 20.8%
DeepSeek (no write premium): any hit rate paysKimi’s chat trace at a five-minute lifetime has a hit rate of 29 percent, just above the line, which is why caching everything there saves only 9 percent. The agent trace sits at 53 percent and our modeled session at 96.5 percent, where the input bill falls to 14 percent of list. For chat-like traffic, mark what is shared across requests, such as the system prompt and the tool definitions, and let the rest go uncached; for agents, cache everything.

Crossing a long tool call
The expensive event in an agent session is an expiry in the middle of a long context, because it turns the cheapest tokens in the session into the most expensive ones. There are three ways across a gap longer than five minutes. Let the entry expire and rewrite the context at 1.25 times the input price.
Buy the one-hour lifetime and pay 2 times instead of 1.25 on the write. Or keep the five-minute entry alive with a request that repeats the prefix every four and a half minutes. Anthropic documents this as pre-warming: with max_tokens set to 0 the request reads the cache, refreshes its lifetime and generates nothing. Each ping costs one read, a tenth of the context’s input price.

Charging the whole one-hour premium to a single gap, pings are cheaper than the one-hour write for gaps up to about 36 minutes and cheaper than a rewrite up to about 54 minutes. When a session has several long gaps the one-hour premium is spread across them and wins sooner.
Pings have to reproduce the prefix and the settings exactly, thinking and effort configuration included, and Anthropic rejects the zero-output form when it is combined with streaming, extended thinking, structured outputs or a forced tool choice.
OpenAI’s GPT-5.6 needs none of this for gaps under half an hour, because its minimum lifetime is already thirty minutes. In our session, one ping per eight-minute gap brings the bill from $1.88 to $0.97.
Keeping the prefix identical
One changed token invalidates every block after it, so the rules are mechanical. Put what never changes first: tool definitions, then instructions, then reference material, then history. Append tool results and messages instead of editing or reordering earlier turns.
Serialize JSON with a fixed key order. Keep timestamps and per-request data at the end or in later messages. Do not toggle settings that are rendered into the prompt: Anthropic’s list includes tool definitions, web search and citation toggles, speed mode, and thinking or effort settings;
OpenAI’s includes tools, parallel tool calls, output format, reasoning effort and verbosity. Manus describes arriving at the first three rules after several rewrites of its agent framework.
Compaction is the deliberate exception. Replacing a context of C tokens with a summary of c tokens forfeits the cache and costs a summary generated at output prices. At Sonnet 5 prices, where output costs five times input, it pays back after about ((w - r + 5)·c + r·C) / (r·(C - c)) further steps. Compacting 80,000 tokens into a 5,000-token summary pays back in about five steps; into a 20,000-token summary, in about 22.
Measuring what you pay
Both APIs report the split. Anthropic returns cache_creation_input_tokens, cache_read_input_tokens and input_tokens, where input_tokens counts only tokens after the last breakpoint; OpenAI returns cached_tokens and cache_write_tokens inside input_tokens_details, and its input_tokens includes both.
The number to watch is the effective input multiplier:
effective input multiplier = f_uncached + w × f_written + r × f_readwhere the f are shares of all input tokens. Our session runs at 0.14 with a warm cache and 0.34 with nine expiries. A rising multiplier with a steady hit rate usually means more writes: a breakpoint on content that changes, or a conversation that grows more than 20 blocks past its last write, beyond the reach of Anthropic’s lookback.
Sizing a cache for your own fleet
For a self-hosted deployment the same formulas become a sizing recipe. Per prefill GPU, a cache that holds τ seconds of fresh output needs G × τ bytes, and in Kimi’s traces ten minutes of unique traffic captured 89 to 98 percent of the achievable reuse.
On an H100 that is 827 GB for Llama-3.1-70B, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash, which picks the tier for you: GPU or host memory for compressed caches, SSD or a shared pool for the rest. A flash tier has to be sized for writes as well as capacity, G × 86,400 bytes per GPU per day if everything is persisted, which is why admission control belongs in the design from the start.
And if reloads must hide behind prefill, the link needs (cached tokens / new tokens) × G: 77 GB/s for the last step of our session on Llama-3.1-70B, which a PCIe Gen5 x16 link cannot deliver and a coherent CPU link can.
Where this could be wrong
Utilization and prices move G and the clock linearly. We used 30 percent of dense FP8 peak and an H100 at $2.50 an hour. Double the utilization and every break-even time halves; halve the GPU price and the DRAM and SSD break-evens halve. The ordering of tiers and models does not change, and neither do the log-scale charts, but individual minutes and hours can be off by a factor of two. Published measurements put real prefill between 0.16 of peak in DeepSeek’s production fleet and 0.29 in SGLang’s GB200 benchmark, so our central value is on the optimistic side: at production utilization every HBM break-even roughly doubles and every capacity figure halves.
The HBM price is an opportunity cost. Pricing idle HBM at its full share of GPU rent is right when serving is memory-bound. A compute-bound prefill fleet values HBM less at the margin, and its HBM break-even is longer than H/G.
Closed models are closed. Everything we say about Anthropic, OpenAI and Google infrastructure is inference from prices and documentation. If their caches per token are far smaller than the open models’ caches, which DeepSeek has shown is possible, the implied storage charges sit even further above HBM rent; if far larger, closer to it.
The traces are one sampled hour of one service. Kimi’s traces come from chat and tool traffic, not from long-running coding agents with multi-minute tool calls. Agent sessions in 2026 probably have longer gaps and higher reuse, which would favor longer lifetimes than our replay suggests. Our correction for the one-hour window removes the truncation bias only for gaps up to thirty minutes. Alibaba has published two-hour traces from its Bailian service in a compatible block-hash format, including a coding trace and a reasoning trace; we have not replayed them here, and they are the obvious next check.
Prices are strategy as much as cost. A provider can price reads below cost to win agent workloads, and rent above cost because explicit caches are bought by customers with few alternatives. Reading prices as cost signals assumes competition does some of the work.
F is a floor, and we left latency out. Attention makes recomputation more expensive at long context, and a hit also buys time to first token, which has a value we did not price. Both make caching more attractive than our numbers show.
The rules are only as durable as the price lists. Pings work because reads refresh lifetimes for free and zero-output requests are documented, and the crossovers in the field guide move with every multiplier. Either could change with the next pricing update, which is why every rule above is written as a formula that can be rerun.
Predictions
Each of these can be checked against a public price list, model card or announcement.
By December 31, 2027, at least two of Anthropic, OpenAI and Google will list a cache-read price at or below 5 percent of the fresh input price for their flagship model. Today only Anthropic does.
By December 31, 2027, an open-weight model from a lab other than DeepSeek, competitive with frontier models on agentic coding benchmarks, will ship with a cache under 4 KB per token at its reference precision.
By December 31, 2027, DeepSeek will list a cache-hit price at or below 1 percent of its cache-miss price.
By December 31, 2027, at least two of the five largest GPU clouds will have announced a production flash tier dedicated to KV cache, built on CMX or an equivalent design.
Confidence dossier
Every number in this piece is listed below with a tier. A means a primary source we read directly: documentation, a model card, a paper or a company post. B means secondary reporting, a preliminary specification or market data. C means our own derivation or model from stated assumptions. D means our interpretation of closed systems, strategy or vendor claims. M means our own measurement on public data, either the Mooncake traces or published model configurations.
Method
Cache sizes come from each model’s published configuration: layer counts, key-value heads and head dimensions for grouped-query models, latent and positional dimensions for MLA models, and the model card for V4.1-Flash.
The V3.2 layout counts 512 FP8 latent values, 16 bytes of scales and 64 BF16 positional values per layer, plus a 128-value FP8 indexer key with a 4-byte scale. The V4-Flash figure comes from our August reconstruction of its compressed attention. F counts two operations per active parameter and nothing else.
G, the break-even times and the capacities follow from the formulas in the text, with P at dense FP8 peak, u at 0.30, an H100 at $2.50 an hour, a B200 at $4.50 and the tier costs in the table. Rubin figures use NVIDIA’s specification page as revised on September 21, 2026.
The trace replay sorts requests by arrival time and treats each hash as one 512-token block. A request hits its longest prefix of blocks present in the cache and stops at the first miss, and every block it touches has its last-use time set to the request’s arrival. Tokens are counted with the final partial block capped at the request’s input length. Lifetimes are sliding: an entry expires a fixed time after its last use. The capacity replay evicts the least recently used block once the cache is full. For the gap statistics we counted only reuses whose previous use came at least 1,800 seconds before the end of the trace.
The session model has a 12,000-token starting prefix, fifty steps, 1,100 tokens of tool output and 400 of model output per step, and Claude Sonnet 5 list prices. A step after a long gap pays a full write for its whole context.
The whiskers in the clock chart vary utilization from 0.15 to 0.5, the H100 price from $1.50 to $4.00 an hour, DRAM from $8 to $16 per GB with 25 to 100 percent overhead, SSD from $0.15 to $0.35 per GB with 50 to 200 percent overhead, and the share of HBM available to cache from 60 to 100 percent.
The gap model sends a ping every 4.5 minutes after the first five, charges each ping one read of the context, and charges the whole one-hour premium to the gap it protects; in the session with pings, each of the nine long gaps lasts eight minutes and takes one ping. No number in this piece required a GPU.
Sources
Anthropic. Prompt caching. Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/build-with-claude/prompt-caching
OpenAI. Prompt caching. OpenAI API documentation, accessed September 2026. https://developers.openai.com/api/docs/guides/prompt-caching
Google. Gemini Developer API pricing, accessed September 2026. https://ai.google.dev/gemini-api/docs/pricing
Google. Context caching. Gemini API documentation. https://ai.google.dev/gemini-api/docs/caching
DeepSeek. Models and Pricing. DeepSeek API Docs, accessed September 2026. https://api-docs.deepseek.com/quick_start/pricing
DeepSeek. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient. Release note, September 10, 2026; the news index dates Context Caching to August 2, 2024. https://api-docs.deepseek.com/news/news260910
DeepSeek. Day 6: One More Thing, DeepSeek-V3/R1 Inference System Overview. Open Infra Index, February 2025. https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md
Grattafiori et al. The Llama 3 Herd of Models. arXiv:2407.21783, 2024. https://arxiv.org/abs/2407.21783
DeepSeek-AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 2024. https://arxiv.org/abs/2405.04434
Moonshot AI. Kimi K2 model summary, README. https://github.com/MoonshotAI/Kimi-K2
DeepSeek-AI. DeepSeek-V3 inference configuration, config_671B.json. https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inference/configs/config_671B.json
DeepSeek-AI. DeepSeek-V3.2-Exp inference configuration, config_671B_v3.2.json. https://github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/main/inference/config_671B_v3.2.json
Bradanini and Tettamanti. DeepSeek V4-Flash: The Cost of Deciding What to Read. August 2026.
Simon Willison. DeepSeek V4, almost on the frontier, a fraction of the price. April 24, 2026 (quotes the V4 report on cache size at 1M tokens). https://simonwillison.net/2026/apr/24/deepseek-v4
DeepSeek-AI. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. Model card, September 2026. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Touvron et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288, 2023. https://arxiv.org/abs/2307.09288
MiniMax. MiniMax-M2 configuration, config.json. https://huggingface.co/MiniMaxAI/MiniMax-M2/resolve/main/config.json
GLM-4.6 configuration, as published in RedHatAI’s NVFP4 release of zai-org/GLM-4.6, config.json. https://huggingface.co/RedHatAI/GLM-4.6-NVFP4/blob/main/config.json
DeepSeek-AI. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism. arXiv:2401.02954, 2024. https://arxiv.org/abs/2401.02954
OpenAI. gpt-oss reference implementation, gpt_oss/torch/model.py. https://github.com/openai/gpt-oss/blob/main/gpt_oss/torch/model.py
Hugging Face Transformers v4.57.1. Qwen3-Next configuration defaults. https://github.com/huggingface/transformers/blob/v4.57.1/src/transformers/models/qwen3_next/configuration_qwen3_next.py
Qwen Team. Qwen3 Technical Report. arXiv:2505.09388, 2025. https://arxiv.org/abs/2505.09388
The SGLang Team. Deploying DeepSeek with PD Disaggregation and Large-Scale Expert Parallelism on 96 H100 GPUs. LMSYS blog, May 5, 2025. https://lmsys.org/blog/2025-05-05-large-scale-ep/
The SGLang Team. Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP (Part II): 3.8x Prefill, 4.8x Decode Throughput. LMSYS blog, September 25, 2025. https://www.lmsys.org/blog/2025-09-25-gb200-part-2
Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs. arXiv:2609.11744, September 2026. https://arxiv.org/abs/2609.11744
Qin et al. Mooncake: Trading More Storage for Less Computation, a KVCache-centric Architecture for Serving LLM Chatbot. USENIX FAST 2025, Best Paper Award. https://www.usenix.org/conference/fast25/presentation/qin
TrendForce. Memory Makers Prioritize Server Applications, Driving Across-the-Board Price Increases in 1Q26. January 5, 2026. https://www.trendforce.com/presscenter/news/20260105-12860.html
ITLDC. Server RAM and SSD market update, 2026 (supplier prices tracked from January 1 to August 21). https://itldc.com/es/blog/server-ram-ssd-market-update-2026
igor’sLAB. DRAM and NAND remain more expensive: TrendForce sees slowing but still rising memory prices in Q3 2026. https://www.igorslab.de/en/dram-and-nand-remain-more-expensive-trendforce-sees-slowing-but-still-rising-memory-prices-q3-2026/
IntuitionLabs. Data Center GPU Pricing 2026: The Full AI Pricing Index. July 20, 2026. https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026
Moonshot AI and Tsinghua University. Mooncake FAST’25 trace release: conversation, tool and agent, and synthetic traces. https://github.com/kvcache-ai/Mooncake/tree/main/FAST25-release
Wang et al. KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider. USENIX ATC 2025. https://arxiv.org/abs/2506.02634
NVIDIA Technical Blog. Introducing NVIDIA BlueField-4-Powered CMX Context Memory Storage Platform for the Next Frontier of AI. March 16, 2026. https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/
NVIDIA. Vera Rubin NVL72, product page, preliminary specifications. https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72
Tom’s Hardware. Nvidia launches Vera Rubin NVL72 AI supercomputer at CES. January 2026. https://www.tomshardware.com/pc-components/gpus/nvidia-launches-vera-rubin-nvl72-ai-supercomputer-at-ces-promises-up-to-5x-greater-inference-performance-and-10x-lower-cost-per-token-than-blackwell-coming-2h-2026
DeepSeek-AI. Fire-Flyer File System (3FS), README, performance section. https://github.com/deepseek-ai/3FS
NVIDIA. NVIDIA Unveils Rubin CPX: A New Class of GPU Designed for Massive-Context Inference. Press release, September 9, 2025, via AIwire. https://dev.aiwire.net/2025/09/09/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference/
Tom’s Hardware. Nvidia removes Rubin CPX accelerators from its roadmap. March 2026. https://tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed
ASCII weekly. Analysis of the 2026 NVIDIA roadmap and the missing Rubin CPX, in Japanese. https://weekly.ascii.jp/elem/000/004/387/4387523/3/
vLLM. Automatic Prefix Caching, design document. https://docs.vllm.ai/en/latest/design/prefix_caching/
Li et al. Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live. arXiv:2511.02230. https://arxiv.org/abs/2511.02230
Gu, Li, Kuditipudi, Liang and Hashimoto. Auditing Prompt Caching in Language Model APIs. ICML 2025. https://proceedings.mlr.press/v267/gu25b.html
Yichao Ji. Context Engineering for AI Agents: Lessons from Building Manus. July 18, 2025. https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus
Alibaba. Qwen-Bailian anonymized usage traces: to-C, to-B, thinking and coder workloads, README. https://github.com/alibaba-edu/qwen-bailian-usagetraces-anon










