Full caching vs. dynamic retrieval: the answer changed
Once you have a context layer grounding your CRM AI, there are two ways to get those rules to the model.
Load them upfront. Put the whole rule set in the system prompt. With prompt caching, that prefix gets cached, and cache reads run at a fraction of the price of fresh input tokens. Call this Full.
Fetch them on demand. Let the agent pull only the rules a question actually needs as tool calls during the conversation. Fewer rules per answer, fewer tokens. Call this Dynamic.
My first benchmark said Full was the obvious winner. The newer data says that was only true for the way I shaped the test.
The practical answer is now clearer: Full is better for cold, one-off questions. Dynamic is better when people ask follow-up questions and hold an ongoing conversation about their data.
The first result was real, but incomplete
I started with an 18-question golden set. Each question ran as an isolated test.
| Per 18-question run | Full | Dynamic |
|---|---|---|
| Total input tokens read | 3.54M | 2.67M (-25%) |
| Billed cost (weighted) | 842K | 1.24M (+48%) |
| Wall-clock time | 1,033s | 1,158s (+12%) |
| Correctness | same | same |
Dynamic read fewer tokens and still cost more. It looked decisive: the full rule set sat in a cheap cached prefix, while dynamically fetched rules landed in expensive message history and were sent again on later turns.
That caching math is still true. What changed was the rest of the system and the shape of the workload.
The production optimization worked
We added conversation caching and a deterministic query validator, then reran the same 18-question set against what is actually running in production.
| Same test set | Before | After | Change |
|---|---|---|---|
| Cost per run | $3.32 | $2.02 | -39% |
| Round trips | baseline | 12% fewer | -12% |
| Response time | baseline | 9% faster | -9% |
Across the production workloads we have measured, the caching changes have reduced spend by as much as 40% and response time by as much as 20%, with no loss in answer correctness. That is no longer a projection from token math. It is showing up in the running system.
The validator matters to that result. In one measured run, 13 of 33 generated queries failed, and nine of those failures came from two mistakes our written guidance already explained. The model had loaded the rules and ignored them, so more prompting was not the answer.
The validator now corrects syntax only when there is exactly one valid reading and rejects ambiguous intent with an actionable message. That removes avoidable agent turns before they become avoidable cost.
Full vs. Dynamic depends on the conversation
With those changes in place, I compared Full and Dynamic again under four workload shapes.
| Test shape | Full | Dynamic | Winner |
|---|---|---|---|
| 18 isolated questions, cold caches | $2.02 | $2.23 | Full by 10.5% |
| Same questions, steady state | $1.533 | $1.272 | Dynamic by 17% |
| 3 conversations x 3 follow-ups | $1.208 | $1.064 | Dynamic by 12% |
| Same conversations, steady state | $0.775 | $0.583 | Dynamic by 25% |
Correctness was identical in every comparison.
The isolated-question benchmark flatters Full. Dynamic needs a separate cached prefix for each record type. In this run, that meant 233K cache-write tokens for Dynamic versus 107K for Full. An 18-question cold run gives those writes almost no time to amortize. Production pays them once per cache window and benefits from the smaller prefix on every question after that.
The multi-turn result surprised me more. I expected Dynamic to fetch its rules again on every follow-up and lose badly. It did not. The prior answer carried the reasoning forward, so only one of six follow-ups needed another retrieval call. Dynamic then used an approximately 18K-token prefix on each turn instead of Full's approximately 38K-token prefix.
That is why Dynamic wins as conversations get longer: it pays a higher setup cost, then carries less context through every subsequent turn.
Cache the conversation, not only the system prompt
The largest broadly applicable gain still came from caching the conversation itself.
Most implementations cache the system prefix and stop there. But the prior messages can become a cacheable prefix for the next turn, so each call reads the existing conversation at the cache rate and writes only its new delta.
Use a short time-to-live for that conversation breakpoint. A conversation only needs to remain warm while someone is actively using it. In my tests, a five-minute TTL beat an hour because it avoided paying a larger write premium for cache lifetime the workflow did not need.
This is especially important for Dynamic. Its advantage appears after the initial retrieval, when follow-up questions can reuse both the conversation and the reasoning already established inside it.
Where I landed
The choice is not really Full versus Dynamic in the abstract. It is one-off usage versus conversational usage.
- Use Full when most requests are isolated, caches are often cold, or you cannot yet measure how people use the system.
- Use Dynamic when users ask several related questions, continue from prior answers, and operate against warm caches.
- In either mode, cache the conversation and validate deterministic query rules before spending another model turn on them.
Our org default remains Full until usage logging tells us how often real sessions become multi-turn conversations. That is a deployment decision, not uncertainty about whether the optimization works. The real-world result is already strong: up to 40% lower spend and up to 20% faster responses. The remaining question is which traffic shape should receive which caching strategy by default.