AI delivery / Operating economics

Prompt caching in practice: Diagnosing writes without reuse

How I investigated Launcherry’s generation-repair caching and verified 99% prompt-token reuse in two measured calls, with the limits made explicit.

Explore the Launcherry series
Written by

Ivan Pedai · Product strategy, UX and AI workflow architecture

Article context

Firsthand methods from Launcherry · Local development and deployment evidence distinguished

The cache was receiving writes without delivering reuse

Launcherry’s production usage records exposed a problem in repeated generation work: calls were writing prompt tokens to cache while reading nothing back within the run. The Google Search copy bank and analysis records showed the pattern. Caching existed in the system, but the expected reuse was not happening.

I wanted to know whether later calls reused the shared context. The answer had to come from cache-read usage records. Turning caching on was only the start of the investigation.

The September 17 verification records the investigation and correction. In each of two later generation-repair calls, 3,879 of 3,932 input tokens were read from cache. That is 98.65%, rounded to 99%. The result describes prompt-token reuse in those measured calls, not a product-wide reuse rate or a 99% cost saving.

Ask what is being reused and what is changing

Generation and repair concern the same campaign, with a specific correction added to each repair request. Some context can stay stable while the correction changes. I examined how the request preserved that shared material.

I tested through the application’s adapter and production provider path. The recorded probe compared requests with stable conditions against requests that changed them, helping locate what interfered with reuse.

The transferable lesson is to test the request the product actually sends. Two prompts that seem similar to a human can differ in ways that matter to the provider’s cache. The model, request configuration and surrounding adapter behaviour belong in the investigation alongside the written instructions.

Provider conditions also differ. OpenRouter’s prompt-caching documentation distinguishes cached-token reads from cache writes and describes provider-specific behaviour. I use that documentation to frame a check, then rely on the application’s measured response to establish its result.

Keep the correction bounded by product requirements

I directed a correction to the generation request design so repair calls could reuse shared context. Generation requirements and validation stayed intact: the optimisation had to produce the same usable deliverable.

The implementation was accompanied by regression cases covering the relevant request behaviour and the relationship between draft and repair. Those checks provide repeatable evidence about how the application constructs its work. A live rerun then provided the evidence about cache reads.

Regression tests checked how the application constructed requests. The live rerun checked whether the provider returned cache reads. I needed both to connect the code change to the observed reuse.

Read the measured result without inflating it

The recorded rerun includes preparation calls before the two successful repair reads. Those earlier calls wrote cache entries. Repairs 2 and 3 then each read 3,879 of 3,932 input tokens from cache and wrote no new cache tokens. The numerator and denominator belong to each of those calls.

The calculation is straightforward: divide 3,879 by 3,932 and multiply by 100. The resulting 98.65% can be rounded to 99% for a compact outcome label, provided the sample stays visible. Reporting the denominator makes the label inspectable.

The whole run also included the earlier cache-writing calls. Its total bill includes output tokens and other usage, so this input-token ratio applies only to the two repair calls. Product-wide savings would require a full-run cost comparison.

I also avoid equating reuse with latency improvement. The cache evidence establishes reads. A latency claim needs timing measurements under relevant conditions. The same discipline applies to claims about throughput, margins or customer savings.

Connect caching to economics without confusing the measures

AI usage is an operating cost in Launcherry. Caching can affect that cost, but total economics also depend on the provider’s rates, output generation, the number of calls and the amount of repair required. A highly reused input can still accompany an unhelpful result or an expensive sequence.

My acceptance decision therefore keeps quality and economics connected but separately measured. Output evaluation assesses the deliverable. Production usage establishes the cost and cache behaviour observed in that environment. The workflow architecture determines which work happens and how often.

For a product team considering optimisation, I would begin with a complete user outcome and its actual usage records. Locate repeated work, identify reusable context and measure the effect of a bounded change. Preserve the original quality requirements so a cheaper run cannot pass simply by doing less useful work.

Require observed reuse before claiming an optimisation

I can now trace this result from the original usage pattern through a request-design correction, regression coverage and live cache reads. Two repair calls are a small sample, but they answer the operating question I set out to test.

The useful capability is being able to trace an operating assumption into application behaviour, test it and correct it. It connects AI workflow construction to the cost of delivering the product, while keeping the evidence understandable.

The Launcherry case study carries the compact outcome. This article preserves the scope behind it. For another system, I would ask the same concrete question before calling caching successful: which later calls read cached context, how much did they read and what did the complete run cost to produce an acceptable result?

Next in the series

Building a SaaS with AI

Read the next article

Put AI to work
with product judgment

Discuss AI workflow construction, AI-enabled product delivery or business process optimisation.