When does prompt caching pay off? A statistical analysis of request dynamics
We design systems around the expected patterns by which people or machines interact with them. If the system is external and we control neither its internals nor its interface, it is sometimes sensible to reverse this and design the interaction pattern around the system. If it is a subsystem of something we are building, we can adjust the interface instead.
Large language models are an external component both of the applications people make and of the tools used to make them. Anyone calling LLM APIs produces a distribution of requests over time, and each distribution has a price tag attached to it. If you work in software, you have probably noticed how prominent the token cost discussions have become, both among developers and among the managers who need to account for this large new operating expense.
Maybe even more importantly, these systems are now widespread not only among software engineers but also among mathematicians, scientists, and other professionals. Since the agent-human interactions already make up a significant part of many people's daily routines, it is important to characterize them. While the future of the relationship is uncertain, it's hard to think of a scenario where this traffic doesn't continue to grow.
I attempt here to abstract away the content of the exchanged messages and deal only with the statistics of requests for a specific agent-human system. More specifically, I write about it in the context of prompt caching, where data centers store a partial internal representation of the specific interaction for a period of minutes to hours. The cache is an external system that we don't control in the sense described above: we can only adapt our interaction pattern and modify the caching configurations. The goal of the post is to use simple statistical models of the API request dynamics and to provide a framework for reasoning about the LLM economics.
I approach the problem through basic dynamical modelling and I also provide a snapshot of my monthly request distribution. The data depends on the tool and on personal usage patterns, and this is the main limitation of this approach. A more extensive study should collect an ensemble of such distributions from many users. Other important aspects of caching are covered elsewhere, such as controlling caches at the API level or the underlying hardware constraints that make caching cheaper but not free.
Prompt caching
Prompt caching is closely related to KV caching from the attention mechanism of LLMs. KV caching is an inference technique that caches K (key) and V (value) matrices computed for all tokens before the current one. The third (query) matrix is needed only for the current token (ignoring techniques such as speculative decoding). The K and V matrices computed up to some point are the same for the (i+1)-th token as they were for the i-th token. This means that we can reuse them for all subsequent token calculations.
Prompt caching extends this idea across multiple API requests, where the matrices for a prefix of the prompt are cached for minutes or even hours. This works because most of the LLMs that we use and that are discussed in this post are unidirectional. In other words, the next-token computation doesn't influence the attention relationships between previously generated tokens. A request hits the cache only if its prefix exactly matches that of the cached request. Throughout the post, I assume the prefix is stable, and misses come only from cache expiry, so that each hit resets the cache lifetime. Keeping the prefix stable is an important issue, but orthogonal to this post.
Coding tools that run without a lot of human intervention make frequent requests to LLM APIs. This means that in most situations they are hitting the cache. The same holds for batch LLM processing, where a single prompt is applied across a large workload. Since reading from cache is much cheaper than recomputing everything, this can lead to savings of up to 90 percent in some of these cases. To figure out how the token economics work, I will continue with the discussion of the pricing model used by large LLM providers.
The pricing model
Token expenditures are set by two quantities: token volume and token prices. When looking at the pricing tables (e.g., Azure), most people focus on two numbers: the input token price, which is a price per million tokens that are input into a model, and the output token price, a higher price charged on the tokens the model produces. Two numbers often skipped are cache reads and cache writes. This is a mistake, since, most likely, the majority of the tokens you send are cached.
Prompt caching strategy depends on the provider and sometimes even on the model. Since GPT-5.6, OpenAI seems to be converging toward the Anthropic model, so most of the following discussion uses the Anthropic caching model. Providers sometimes offer discounts for models; for example, the recently released Opus 5.5 charges for cache reads relative to standard . For simplicity, we only discuss the standard pricing.
To analyze the savings coming from prompt caching, let's compare the cost of a cached session to the cost of an uncached session. Prompt caching applies only to the cacheable prefix, which is the part of the context before the breakpoint. The breakpoint is a marker in the API request; it signals to the server which part of the context goes through the cache (either read or written). This can be understood through a schematic representation of the user/model conversation:
Initial context: System prompt + tool definitions
Turn 1: First user message. -> API request sent.
Turn 1: Assistant's response.
Turn 2: Second user message. -> API request sent.
---------------------------------
Turn 2: Assistant's response.
Turn 3: The current user message. -> API request sent.
Below the dashed line, the user is looking at the assistant's last message and typing in the new message. When the model is asked to generate the response to the user message from turn 3, it can only read from the cache everything up to the last API request in turn 2. The dashed line represents the breakpoint after turn 2. Ordinarily, tokens after the breakpoint would be billed as input, but here we assume that the breakpoint always moves to the end of the context. This means that every token is either a cache read or a cache write. In the case that the cache didn't expire, anything before the dashed line is billed as a cache read, and anything after it as a cache write.
There are several complications to this story:
- For the Anthropic API, the breakpoint can be set manually, but we assume that it is placed at the end of the context, which is what Claude Code does by default.
- Caching needs a minimum number of tokens to kick in, and this minimum depends on the model, which means that if you're near this boundary, it might be worth adding some tokens to the input to cross it.
- Claude Code can set multiple breakpoints automatically. This enables the preservation of the system prompt and tools cache even if CLAUDE.md or conversation history changes mid-session.
- The user doesn't necessarily initiate the API request; this can, for instance, come after the agent receives the output of a tool call.
Since these details are not crucial for the analysis, we assume that there is a single breakpoint at the end of the context. We also don't keep track of how the request was initiated.
Caching is turned on by default in Claude Code, but we use the non-caching case as the reference point. The latter case is not only a theoretical baseline; I have seen a setup where a round-robin load balancer was put in front of the API, caching had to be turned off, and the costs exploded.
In the non-caching case, each input token's cost is . Output tokens are charged at a significantly higher price. Since these are the tokens that the model has to compute anew, generation itself can never be served from cache, and therefore output tokens are excluded from the comparison. They don't leave the accounting entirely, since the model's output becomes part of the input context on the following request.
With caching turned on, the prices for cache reads and cache writes are set by multipliers of the base price :
| Cache lifetime | Read, | Write, |
|---|---|---|
| 5 minutes | 0.1× | 1.25× |
| 1 hour | 0.1× | 2× |
There are two cache lifetimes that have the same read price and different write prices: and . For comparison, GPT-5.6 guarantees a minimum lifetime of 30 minutes with the 5-minute multipliers. Paying a premium for cache writes compared to not using the cache can be understood as a simple storage cost. The cost model has only two pricing tiers, though, regardless of the model. This by itself indicates that a potential cost-saving strategy should be designed around these two timescales.
On the read side, we see that cache reads have a significant discount, but that they are not free. Even though the K/V tensors don't need to be recomputed between requests, while the model is generating new tokens it has to keep the tensors resident in high-bandwidth memory (HBM) and consume bandwidth towards the GPU compute units that other requests would otherwise use. This is the same whether the tensors were obtained from the cache or recomputed. K/V tensors depend on a specific conversation context, so they're unlikely to be shared between users, unless, for example, the users are initiating the request with same prefix from an LLM application. In contrast, a weight matrix in a dense layer doesn't depend on tokens in the context and can be used to run inference for different users.
Estimating the cost
We first choose an arbitrary session and fix it, so that every quantity discussed below describes a single session and not an average over many sessions. The token number counts the full input context of the -th request: the system prompt, the tools, every earlier turn, the assistant response to the previous request, and the new user message. The output of the -th request is not part of ; it is billed separately and enters as input.
If there are requests per session and the -th request has tokens, the input cost is
When caching is turned on, the input cost splits into cache reads and cache writes as
where and are read and write multipliers. For the -th API request, is the token count for the tokens read from cache, and is the number of tokens written to cache (at a more expensive rate). Since we assume that the breakpoint is always at the end of the context, the total number of tokens is split into these two quantities,
We now define the cost ratio by dividing the cached cost by the non-cached cost,
This is a dimensionless quantity that determines the cost efficiency of a session. The base input token price drops out, since the multipliers fully determine the price ratios. Here, the definition of cache miss rate naturally appears as the fraction of total tokens that are cache writes,
When the pricing model is fixed (the multipliers and the TTL), the cost efficiency of a session is fully determined by . Note that is a ratio of token sums, not an average fraction of API requests that miss the cache. The two differ whenever requests vary in size.
The cache misses might come from reviewing the model output, leaving for lunch during the coding task, or the model reasoning for too long. Regardless of the origin, the cost ratio of the session is
The multipliers are constants set by the caching configuration, and we get the linear dependence of the cost ratio on the cache miss rate:
Caching becomes cost-efficient when the cost ratio is smaller than , which puts an upper bound on the cache miss rate:
To put it concretely, for 5-minute cache lifetimes the cache miss rate needs to be below and for 1-hour cache lifetimes below . As shown later, for most reasonable use cases, one would have to make a significant effort to reach a cache miss rate above 78 percent, so on the 5-minute tier, prompt caching almost always leads to significant cost savings. The threshold for the 1-hour tier sits at roughly half that.
Through the rest of the post, I discuss how different usage patterns influence the cache miss rate .
The rapid-exchange limit
Assume first that the cache never expires, meaning that the gap between consecutive requests never exceeds the cache lifetime (TTL). We make an additional assumption to keep the estimates simple: each turn appends the same number of tokens .
At the -th request, the context is tokens. Of these, were present on the previous request and are served as reads, while the newest are written. Over requests, we write
since each request writes exactly new tokens. The number of read tokens is
which makes a total number of tokens equal to .
The cache writes don't come from cache expiry, but from a structural constraint: the newest tokens have never been sent before, so no cache could hold them, and they are written whatever you do, even with a cache that never misses. The miss rate equals
This is a floor on the fraction of tokens billed at the higher price (cache writes), and it decays as . It also means the floor stops mattering at the session lengths agentic tools actually reach. By the time a session approaches auto-compaction, the tokens written on any single request are negligible against the accumulated context being read alongside them, and a session that hits perfectly costs almost nothing beyond the reads.
The uniform miss rate
The rapid-exchange limit assumed the cache never expires. Now let each request miss with the same probability , keeping the assumption that every turn appends the same number of tokens . We first estimate the number of tokens read and calculate the tokens written to the cache by subtracting from the total.
On the -th request, the previous context is served as reads whenever the cache is hit, so the expected read count is the probability of a cache hit multiplied by the number of tokens sitting in cache from the previous request, i.e., . Everything else in the context is written: the new tokens, which no cache could hold because they have never been sent before, plus the previous context whenever the cache misses. The tokens written over the session are
where . Since the cache miss rate is the fraction of all tokens billed as writes, and the totals from the previous section still apply, this gives
The structural floor has not gone away; it has been damped by the fraction of requests that hit. At we recover the floor, and at every token is written, as it must be when nothing is ever cached.
The result clarifies the difference between the token-weighted cache miss rate and the probability that the request misses the cache. In the limit of a large context (), the contribution from writing the new tokens into cache at each round becomes negligible. This means that and have the same value. However, this happens only due to the approximation that uses the same probability for each step and the same number of tokens appended with each request. In general, these two quantities are different even in this limit, and this difference is precisely stated in the appendix. is a token-based average over the whole session, while is in general not a single number but a set of probabilities, .
It is worth writing the break-even condition in terms of a single quantity. Define
as the ratio of write surplus to read discount, taking as the multiplier of the non-cached case. The quantity is set by providers based on TTL; they charge more for writes if cache lifetimes are longer, consistent with the longer storage of K/V matrices.
The condition for which prompt caching becomes cost-efficient, , becomes
If exceeds , no session is long enough to make caching pay. The failure happens in the infinite-length context limit since the structural term vanishes as grows and approaches . With and , this reproduces the ceilings of roughly and percent from the previous section. As we already discussed, under our current approximation, the values of and coincide in this limit.
If the cache miss rate is reasonably low, we can look at the condition on the request number . A session that never misses () needs at least two requests to break even on the 5-minute tier and four on the 1-hour tier, and both are reached almost immediately. Even if half of requests miss, only 3 requests are needed to break even in the 5-minute tier. Both assumptions here are strong: every request appends the same number of tokens, and every request misses with the same probability. Neither holds in practice, and the appendix at the end of this post works out what happens when both are relaxed. The answer turns out to depend on a quantity this section cannot express, namely whether the misses fall early or late.
Measuring the request distribution
Coding agents store the history of API requests in local directories, for example, .claude for Claude Code, or .copilot for GitHub Copilot. The gaps between the requests can be extracted from the timestamps stored in these files. Here, I present a monthly snapshot of my Claude Code usage statistics and characterize the data using quantities defined above. If you want to analyze the data for a period longer than a month, you first need to change the default retention period.
Here I show the time gaps between requests for 8,121 request pairs across 89 sessions in terms of percentiles:
| percentile | gap |
|---|---|
| p10 | 3.4s |
| p50 | 12.6s |
| p80 | 1m 02s |
| p90 | 2m 11s |
| p95 | 4m 16s |
| p99 | 26m 31s |
| max | 2d 17h |
The same data is represented through gap fractions lying between two times as:
| gap | share of pairs | 5-minute cache | 1-hour cache |
|---|---|---|---|
| under 30s | 67.8% | hit | hit |
| 30s to 5m | 27.7% | hit | hit |
| 5m to 1h | 3.98% | miss | hit |
| 1h to 12h | 0.31% | miss | miss |
| over 12h | 0.17% | miss | miss |
In the table, 5 minutes are treated as a hard cutoff, but they are actually a guaranteed minimum. In my data, about half of request gaps between 6 and 8 minutes still hit the cache. Here are a few general properties that can be observed from the data and from the figure below.
- Most of the requests lie below the shorter cache expiry time of 5 minutes. The 95th percentile is at 4m 16s, and only 4.46 percent of request gaps are above 5 minutes (3.24 percent of requests actually miss the cache).
- For request gaps above ~12 hours, the request gaps match the daily work pattern. The maximal request gap of 2d 17h comes from the resumption of a session after the weekend.
- The long-time distribution has heavy tails, with a long time limit that behaves broadly like a power law.
The distribution of request gaps on a log-log scale. The inset figure shows the distribution on the linear scale.
We can now estimate the cache miss rate and the cost ratio for this monthly snapshot of usage. I used the default 5-minute caching tier, but I also show the data for the counterfactual possibility of using a 1-hour tier:
| cache expiry time (TTL) | ||
|---|---|---|
| 5 minutes (observed) | 4.83% | 0.156 |
| 1 hour | 2.43% | 0.146 |
The cache miss rate is very low with both TTLs, and the cost ratio is around in both cases, leading to significant savings of around 85 percent on input cost (output cost is ignored here). As discussed in the simplified model from before, part of comes from the new tokens that are written on each request even when the cache is hit. In my data, this number is around 1.77% out of 4.83%. Interestingly, if I used the longer cache expiry time of 1 hour, my savings would be slightly larger (this might have been the period where I didn't run a lot of autonomous sessions). Depending on your usage, it also might make sense to play around with different tools and strategies for keeping the cache warm (such as claude-thermos).
In the concluding parts of the post, I want to introduce a simple statistical model of the dynamics and show how it compares with the data. But before that, I want to stress the importance of conscious strategies for handling long context, which is outside of the scope of the post. Compacting the long context both makes the misses cheaper and improves the output quality by avoiding context rot. In light of the previous discussion on the utilization of the cache, one should identify where long gaps fall in long-running sessions, and not let them land on a large context. This is achieved most efficiently through compaction or creating memory artifacts such as textual files that preserve the essence of conversation. I prefer the latter because it creates a file whose contents can be read and modified, giving much more control.
A simple model of request dynamics
Here I want to describe a toy model of the request dynamics. While I will later show how the model fits the actual data, the main goal of the section is to give a formal dynamical description so that we can reason about the problem quantitatively.
The previous sections took as given. It is the probability that the request was sent to the API after the cache has expired. However, the underlying quantity that determines the dynamics is the probability distribution of request times. If the gaps between requests are independent and identically distributed, which is a dubious but simplifying assumption, the requests form a renewal process. This enables us to look only at a single time distribution of requests, which answers a simple question:
What is the probability density that a request will be sent at time after the previous request?
The renewal assumption is that we don't need to know anything about the previous requests to answer this question. This is unlikely to be true for this process: the tool calls are usually clustered; the agent runs a command whose output it uses for the next command, and so on. The human returns from a 1-hour-long break and asks Claude ten quick questions, not reading most of its output.
Now, we make a model of the distribution of requests as one that consists of two time scales, the slow time scale and the fast time scale . The simplest two-scale distribution that one might devise is the mixture of two exponential distributions,
We can understand this distribution as arising from the following process: the next request is sent either due to a slow process with probability or through a fast process with probability . Each of these processes follows an exponential distribution with a single parameter, which is the mean time. These distributions can then be interpreted as conditional probabilities, i.e., the probability that the request happens at time provided that the underlying process is slow/fast. The exponential distribution is the simplest single-parameter continuous distribution that we might think of. It arises by dividing time into parts. At each interval , the probability of an event happening (such as sending an API request) is the same, , and is independent of what happened before. This means that after such time intervals, the probability that the request wasn't sent is , which in the limit of an infinitely small interval () tends to the exponential , from which we derive the exponential probability density.
The reason for introducing such a deliberately simple model is that we want to eliminate all the complex details but leave the two time scales as a mathematical description of two processes that we intuitively understand: the agentic tool calls on the fast side, and the time spent thinking about the output or coffee breaks on the slow side. The complexities are not ignored only because they are hard to model (although this is true), but also because they will depend on the LLM tool that is used, the way that someone uses it, and the task for which it is used. The model that applies to a single person doing a specific task with a specific agent is not very useful since it lacks generality.
The slow scale and cache expiry
Once we have the time distribution of requests, we can extract the cache miss probability by integrating the probabilities that a request arrived after the cache expiry time ,
Here, we make a simplifying assumption that can also be used as the definition of the fast time scale. This is the time scale of processes that are much faster than ,
This means that the cache miss probability comes dominantly through slow processes, since the integral of the fast component is negligible. If we accept this separation of the dynamics into two scales, it's reasonable to make this assumption. A significant part of the requests, especially tool calls, usually happen on the scale of seconds and not minutes, which is also reflected in my personal usage data that I showed earlier.
The cache miss probability is approximately
since the fast term is exponentially small. In the most conservative estimate, when the slow scale is much larger than the TTL, we get that is essentially the fraction of requests coming from the slow distribution. This means that in the large context limit (), the cache miss rate upper bound is set by the fraction of requests coming from the slow component. If the slow-component mean time is comparable to TTL, the cache miss rate will be in single digits even if 20 percent of the requests come from the slow component, which is a conservative number, since most of the calls run in a tight loop.
We can now use this model to compare the two TTLs. As we showed, the efficiency of the strategy is set by , where the write cost multiplier and cache miss rate are the two quantities that depend on the choice of TTL. The 1-hour strategy wins when its is smaller, i.e.,
In the limit of large , the cache miss rate is equal to with our assumptions (constant and same ), and we evaluate using the expression from above for two TTLs. Slow-mode weight and slow mean time are the properties of a specific interaction pattern and unrelated to TTL. This means that the condition is
The longer TTL pays off when
When is longer than 110 minutes, neither cache expiry configuration can avoid misses from the slow process. Since the longer TTL comes with a higher write price, it pays off to use the shorter TTL.
The reason that the longer TTL doesn't always win under this condition is that we ignored the fact that new tokens need to be written to the cache at each request (finite case). This surcharge is higher for longer TTL. If is short, the gains from fewer cache misses become less likely to compensate for this surcharge.
I will skip the calculation, but for finite , is the time for which the 1-hour tier has the largest advantage compared to the 5-minute tier. At this , the shortest session where the 1-hour tier wins is , which is around 14 requests for my , which I estimate in the next section. Ignoring the fact that the model is not completely right, these scales suggest that in my case the 1-hour tier wins: my slow scale is somewhat shorter, but my average number of requests is significantly higher.
Throughout the post, I made some assumptions to arrive at simple results. The interesting conclusions might come from the manner in which some of these assumptions break.
Fitting the data
Here is a short overview of how the model actually fits the data.
Below I show the plot that fits the exponentials to the real usage data. The x-scale is logarithmic as in the previously shown plot, but the y-scale here counts the number of requests in bins that grow with time, while the previous figure showed the per-second rate. This is the main reason that the histograms look so different. With these plots, it's easier to separate the two time scales. We don't fit the part of the graph above , which we choose to be 12 hours. The slow scale depends on the choice of , but there is a plateau in request gap numbers around this time, which justifies the truncation.
It's clear that the exponential distribution is not the right family of distributions to fit the data. However, the separation into two timescales improves the fit significantly, which becomes clear in the lognormal plot that I show later. Here are the parameters of the plot,
| parameter | value |
|---|---|
| (slow weight) | 0.145 |
| 20.5s | |
| 477s |
Most of the weight is in the fast distribution (~85.5 percent), and the slow scale is ~8 minutes, which somewhat justifies the relative effectiveness of the 1-hour tier on my usage data.
The fit of two exponentials to the request gap data
I found that a much better fit uses a mixture of two lognormal distributions. The lognormal distribution comes from the multiplicative process involving independent and identically distributed variables. It's essentially a multiplicative version of a Gaussian, which comes out as a limiting distribution for a sum of independent and identically distributed variables.
The fit shown below confirms the separation between two timescales. This distribution is easier to fit to the data because each component of the mixture has a width that is an additional parameter we can control (the fit parameters shown in the figure). While the fit looks good, I stop the analysis here, since a fit to a single user's usage snapshot is hard to generalize. Before further analysis of the structure of request dynamics, or an attempt to identify a process that generates this distribution, I would like to see an ensemble of distributions across many users and tools.
The fit of a double lognormal distribution to the request gap data.
Conclusion
The cost ratio measures the cost efficiency of the session. Besides the prices set by providers, it is determined by the cache miss rate , which is the central quantity that I evaluate. The most straightforward contribution comes from the probability that the gap between two requests is longer than the cache expiry time (TTL). The additional contribution is from the fact that we need to pay to store newly added tokens in the cache. What I find is that prompt caching pays off when the miss rate is below 78 percent for a 5-minute TTL, and below 47 percent for a 1-hour TTL. Both are well above my miss rate of approximately 4.8 percent, leading to around 85 percent in input savings.
In the appendix, I lay out the general (but denser) calculation of the cache miss rate. Relaxing the assumption that each request appends the same number of tokens with the same miss probability adds one more term. The covariance between the request miss probability and the context size of the previous request
measures in which part of the session cache misses occur. If the misses happen more often when the session is already long (large ), the covariance is positive, which increases the cost. For my distribution, I found that the covariance is negligible; relative to the average request size , it is , which is much smaller than my .
I hope that these quantities are useful to anyone who wants to quantify how they use LLMs. In the analysis of the real usage data, one needs to be careful to define a cache miss precisely. For example, a naive analysis might count running the /clear or /compact commands as a cache miss, even though they shouldn't be understood that way (if we follow the framing of the post). Furthermore, the pricing models are subject to change, so it is important to check the current official pricing info before running the analysis.
These quantities and models should be seen as a complement to good context management practices. In the future, I would like to run the analysis on various tools and also see how my usage evolves with time.
Appendix: the general case
Everything before assumed that each request appends the same number of tokens and that each request misses with the same probability. Here we drop both of these assumptions.
Let there be requests and let the number of new tokens appended in the -th request be . Each request is stateless and has to send the full context with tokens, so is the token count difference between the two consecutive requests,
where we set , so that for the first call. The total number of tokens that are sent after requests is
For each request we sum up all the tokens that have been added so far and, after this, sum up all the requests. The new tokens from the -th request appear in the sum times, since they are present in every request after they are added into the context.
The number of tokens read from cache in the -th request is the product of the probability that the cache was hit and the number of tokens present in the previous request. If is the probability that the cache was missed in the -th request, then is the cache hit probability, and the total number of read tokens after requests is
Nothing is cached before the first request, so the first request writes the full context into the cache. This means that doesn't influence the cost calculation (it's always multiplied by ); we can set it to an arbitrary number. To keep the definitions simple, we set , where the average is over . This allows us to keep the sum from to as the definition of .
The cache write tokens are then the tokens that are not read, , or
The two terms are worth examining more closely. The first term comes from the fact that each request has to write its new tokens at least once, with probability . The second is the cache miss probability multiplied by the token count of the -th request. If the cache misses on the -th request, we set and the number of tokens written is , the whole context. If it hits, and only the new tokens are written.
The fraction of tokens written to the cache is the cache miss rate, defined per token and not per request:
Using the cost multipliers and , the cost efficiency condition for caching () can be written as
If we send the same number of tokens for all with the same probability , we recover the earlier result by expressing the left-hand side as
from which
with as before. If there is no at which caching saves money, as the limit of the expressions above shows. These are the ceilings of roughly and percent, following from and .
How the general case differs
It is convenient to define the shifted token count , the previous request's context size, with , which lets every sum start at . We define the averages
where averaging is done over requests, and the deviations are defined as
The cache miss rate is then
and the sum in the numerator evaluates to
where the terms linear in deviations vanish by identity. The last term contains
the covariance between the cache miss probability and the shifted token count. The average shifted token count is
With this, the token-weighted cache miss rate is
The first two terms come from average quantities of probabilities and request sizes. Without correlations, these terms partition the writes: a token is written either because the request it belongs to missed the cache, or because the request hit and the token is new and has never been written before. The second contribution is therefore present even when the cache is always hit.
If each request comes with the same size , and the miss probabilities show no linear trend with the position in the session (so that ), we recover the earlier formula with an average probability, . The second term becomes small because the context added per request stays constant while the total number of tokens sent grows as , since a context of order is sent order times. In practice, the equal-size assumption fails more clearly on the first request, since contains the system prompt and tool definitions, in addition to the first message.
The covariance is the last piece of the puzzle, and it says that the total cost depends not only on how often the cache misses but on where in the session the misses fall. A miss is charged against the context that has accumulated up to that point, so the same number of misses, moved later in the session, costs more. The practical consequence is that a long gap before the first tool call is almost free, while the same gap at a hundred thousand tokens of context is not.