Introduction
Prefix caching (often called KV cache reuse) lets the inference engine skip recomputing attention state for the portion of a prompt it has already processed. In a multi-turn agent loop, the system prompt, tool definitions, and prior turns are byte-identical from one request to the next, so only the newest tokens need fresh compute.
On Crusoe, Serverless Inference runs on MemoryAlloy, a cluster-wide memory fabric with cache-aware routing. Requests that share a prefix are steered toward the KV state already built for it, which is why a correctly structured conversation should keep hitting cache even when your traffic is spread across many requests.
Cached input tokens are billed at the cached tokens rate and reduce time-to-first-token, so the hit rate directly affects both cost and latency for conversational and agentic workloads.
Serverless Inference reports cache reuse per request in the response usage object. usage.prompt_tokens_details.cached_tokens is the number of prompt tokens served from cache; dividing it by usage.prompt_tokens gives that request's hit rate. This value is computed by the engine, not the client, so it is the authoritative number to test against.
This guide shows you how to measure the hit rate for your own prompts, how to make the measurement reproducible, and how to read the result. By the end you will have a per-turn table of prompt tokens, cached tokens, and hit rate for a conversation shaped like your production traffic.
ℹ️ Note: Cache granularity varies by model. Some models credit cached tokens in fixed-size segments rather than per token, so short prompts can report
cached_tokens: 0even when the prefix is identical on every call. Always measure with your real system prompt and tool definitions rather than a short synthetic one.
Prerequisites
- Intelligence Foundry API Key (See Retrieving Your Intelligence API Token)
- Python 3.8 or Later
- Network Access to
https://api.inference.crusoecloud.com/v1 - Your Production System Prompt Saved Locally as
system_prompt.txt - Your Tool Definitions Saved Locally as
tools.json(Tool-Calling Workloads Only)
Instructions
Step 1: Confirm the Model Identifier
List the models your key can access and copy the exact identifier for the model you want to test. Identifiers are case- and prefix-sensitive; a wrong spelling returns 404 model_not_found.
export CRUSOE_API_KEY=<YOUR_API_KEY> curl -s -H "Authorization: Bearer $CRUSOE_API_KEY" \ https://api.inference.crusoecloud.com/v1/models \ | python3 -c 'import json,sys; [print(m["id"]) for m in json.load(sys.stdin)["data"]]'
Use the identifier exactly as returned by this command in the script below.
💡 Tip: The same
/v1/modelsresponse lists each model's per-token pricing, includingpricing.input_cache_reads, the price applied to cached tokens. See How-To List Model Features via curl and the Inference API for filtering the catalog to one model.
Step 2: Build a Conversation That Matches Your Workload
- Start with your real system prompt and tool definitions. Add a unique user message per turn (a counter or nonce is enough) so that only the true shared prefix can be reused.
- Append the assistant's reply to the message list after each turn, exactly as your agent would. The prefix must grow turn over turn or you are not testing multi-turn reuse.
- Set
temperature: 0and a smallmax_tokens(32 is plenty) so assistant replies are short and the prompt length is predictable.
Step 3: Send the Turns Sequentially and Record Usage
Save the following as cache_test.py. It sends N turns one after another within a single session and prints prompt tokens, cached tokens, and hit rate for each.
import json, os, sys, urllib.request
KEY = os.environ["CRUSOE_API_KEY"]
URL = "https://api.inference.crusoecloud.com/v1/chat/completions"
MODEL = "<MODEL_ID>" # exact id from /v1/models, e.g. provider/model-name
TURNS = int(sys.argv[1]) if len(sys.argv) > 1 else 5
system = open("system_prompt.txt").read()
tools = json.load(open("tools.json")) if os.path.exists("tools.json") else None
def ask(messages):
body = {"model": MODEL, "messages": messages, "temperature": 0, "max_tokens": 32}
if tools:
body["tools"] = tools
req = urllib.request.Request(URL, data=json.dumps(body).encode(), headers={
"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=120) as r:
return json.load(r)
msgs = [{"role": "system", "content": system}]
print(f"{'turn':>4} {'prompt_tokens':>14} {'cached_tokens':>14} {'hit_rate':>9}")
for t in range(TURNS):
msgs.append({"role": "user", "content": f"[turn {t}] What is the status of order {100000+t}?"})
resp = ask(msgs)
usage = resp["usage"]
p = usage["prompt_tokens"]
c = (usage.get("prompt_tokens_details") or {}).get("cached_tokens", 0)
print(f"{t:>4} {p:>14} {c:>14} {100*c/p:>8.1f}%")
reply = resp["choices"][0]["message"]
msgs.append({"role": "assistant", "content": reply.get("content") or ""})Run it with the number of turns you want to test:
python3 cache_test.py 10
Turn 0 is a cold request and will normally show cached_tokens: 0. Turns 1 onward show how much of the growing prefix was reused.
ℹ️ Note: The script appends only the assistant's text reply and drops any
tool_calls. That is fine for measuring cache reuse (the prefix still grows every turn) but it is not a faithful replay of your agent's tool loop.
Step 4: Vary One Thing at a Time
- To check whether concurrency affects your rate, run several sessions in parallel (each with its own message list) and compare against the sequential run. If the per-turn values are the same, request routing is not a factor.
- To check whether prefix length matters, pad the system prompt to different sizes and rerun. If
cached_tokenssteps up in fixed increments rather than tracking prompt length closely, the model credits cache in segments; see the Note in the Introduction. - Keep the model, temperature, and tool definitions fixed across runs so the comparison is apples-to-apples.
Each invocation of cache_test.py keeps its own message list, so running three copies at once gives you three independent parallel sessions:
for i in 1 2 3; do python3 cache_test.py 5 > run_$i.txt & done; wait
ℹ️ Note: Parallel sessions count against your project's requests-per-minute limit for the model. Organizations without a payment method on file have a much lower default RPM, so a large parallel run can return
429 Too Many Requests. See Serverless rate limits.
Step 5: Read the Result
- A healthy multi-turn session shows
cached_tokensrising every turn after the first and never dropping to zero mid-session. - Aggregate hit rate for a session is
sum(cached_tokens) / sum(prompt_tokens)across all turns, which is the figure to compare against your usage reports. - If
cached_tokensis zero on every turn, check that the prefix really is identical (whitespace, tool ordering, and JSON key order all matter) and that your prompt is long enough for the model's cache granularity.
💡 Tip: When comparing against another provider, run the identical conversation on both and compare
cached_tokensper turn, not just the aggregate. Providers count cache credit differently, and a per-turn view shows whether a gap comes from granularity or from missed reuse.
Example
A team runs a customer-support agent with a long system prompt and two tool definitions. They save the prompt to system_prompt.txt, the tools to tools.json, and run python3 cache_test.py 3:
turn prompt_tokens cached_tokens hit_rate 0 <PROMPT_TOKENS> 0 0.0% 1 <PROMPT_TOKENS> <CACHED_TOKENS> <RATE>% 2 <PROMPT_TOKENS> <CACHED_TOKENS> <RATE>%
Turn 0 is cold, so cached_tokens is 0. On turns 1 and 2 the engine reuses the shared prefix and reports a non-zero cached_tokens value. If that value is the same on both turns and the same across repeated runs, the engine is caching deterministically and the rate you see is the expected rate for a prompt of this length on this model. Running the same three turns with several sessions in parallel and getting identical numbers rules out request routing as a factor.
Related Articles
- How-To List Model Features via curl and the Inference API
- How-To Get Started With Text Generation on Crusoe Managed Inference (Python)
- How-To Fix 429 RateLimitError When Interacting With Crusoe Managed Inference Service