Skip to content

feat(llm): report prompt-cache hits as LLMUsage.cached_tokens - #41

Merged
treo merged 1 commit into
mainfrom
paul/usage-cached-tokens
Oct 2, 2026
Merged

treo merged 1 commit into
mainfrom
paul/usage-cached-tokens

Conversation

@treo

@treo treo commented Oct 1, 2026

Copy link
Copy Markdown
Member

Every major provider now tells you how much of the prompt it served from its cache, and LLMUsage had nowhere to put that: it is three ints, so providers dropped the number at the point where they read the response. Consumers that bill or report usage (the xac coding harness is the one that prompted this) could not show a cache hit rate at all.

cached_tokens defaults to 0, which is what makes the field non-breaking: responses built without it (older code, plain dicts, MockLLM scripts) simply validate and read 0. Two rules the field carries by convention, both stated in its docstring:

  • it is a subset of prompt_tokens, never an extra spend, so total_tokens stays prompt + completion;
  • 0 means "nothing cached", not "unknown" — there is deliberately no third state, because a rate is the only thing the number is used for.

Each provider reads its own gateway's name for the same figure: OpenAI-compat usage.prompt_tokens_details.cached_tokens (the details object is absent when nothing was cached, hence the getattr chain), Anthropic cache_read_input_tokens (billed cache reads are already counted inside input_tokens), Gemini cached_content_token_count, and Bedrock, which carries the Anthropic name in its raw usage dict. LLMResponse.merge sums the count alongside the others, or a merged response would lose it.

UsageEvent inherits the field for free — it extends LLMUsage, so SimpleToolOrchestrator's UsageEvent(**usage.model_dump()) now emits it.

Tests cover the two properties consumers actually depend on: the count survives an LLMResponse dump/reload round-trip (a subclass attribute would not — pydantic serializes a nested model through its declared type), and a merge sums it. The provider test builds a real ChatCompletion and skips when the openai extra is not installed.

Every major provider now tells you how much of the prompt it served from
its cache, and `LLMUsage` had nowhere to put that: it is three ints, so
providers dropped the number at the point where they read the response.
Consumers that bill or report usage (the xac coding harness is the one that
prompted this) could not show a cache hit rate at all.

`cached_tokens` defaults to 0, which is what makes the field non-breaking:
responses built without it (older code, plain dicts, MockLLM scripts) simply
validate and read 0. Two rules the field carries by convention, both stated in
its docstring:

- it is a *subset* of `prompt_tokens`, never an extra spend, so `total_tokens`
  stays prompt + completion;
- 0 means "nothing cached", not "unknown" — there is deliberately no third
  state, because a rate is the only thing the number is used for.

Each provider reads its own gateway's name for the same figure: OpenAI-compat
`usage.prompt_tokens_details.cached_tokens` (the details object is absent when
nothing was cached, hence the `getattr` chain), Anthropic
`cache_read_input_tokens` (billed cache reads are already counted inside
`input_tokens`), Gemini `cached_content_token_count`, and Bedrock, which
carries the Anthropic name in its raw usage dict. `LLMResponse.merge` sums the
count alongside the others, or a merged response would lose it.

`UsageEvent` inherits the field for free — it extends `LLMUsage`, so
`SimpleToolOrchestrator`'s `UsageEvent(**usage.model_dump())` now emits it.

Tests cover the two properties consumers actually depend on: the count
survives an `LLMResponse` dump/reload round-trip (a subclass attribute would
not — pydantic serializes a nested model through its declared type), and a
merge sums it. The provider test builds a real `ChatCompletion` and skips when
the openai extra is not installed.
@treo
treo requested a review from wmeddie October 1, 2026 19:48
@treo
treo merged commit 6f4167d into main Oct 2, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant