Repository navigation
feat(llm): report prompt-cache hits as LLMUsage.cached_tokens - #41
Merged
Merged
Conversation
Every major provider now tells you how much of the prompt it served from its cache, and `LLMUsage` had nowhere to put that: it is three ints, so providers dropped the number at the point where they read the response. Consumers that bill or report usage (the xac coding harness is the one that prompted this) could not show a cache hit rate at all. `cached_tokens` defaults to 0, which is what makes the field non-breaking: responses built without it (older code, plain dicts, MockLLM scripts) simply validate and read 0. Two rules the field carries by convention, both stated in its docstring: - it is a *subset* of `prompt_tokens`, never an extra spend, so `total_tokens` stays prompt + completion; - 0 means "nothing cached", not "unknown" — there is deliberately no third state, because a rate is the only thing the number is used for. Each provider reads its own gateway's name for the same figure: OpenAI-compat `usage.prompt_tokens_details.cached_tokens` (the details object is absent when nothing was cached, hence the `getattr` chain), Anthropic `cache_read_input_tokens` (billed cache reads are already counted inside `input_tokens`), Gemini `cached_content_token_count`, and Bedrock, which carries the Anthropic name in its raw usage dict. `LLMResponse.merge` sums the count alongside the others, or a merged response would lose it. `UsageEvent` inherits the field for free — it extends `LLMUsage`, so `SimpleToolOrchestrator`'s `UsageEvent(**usage.model_dump())` now emits it. Tests cover the two properties consumers actually depend on: the count survives an `LLMResponse` dump/reload round-trip (a subclass attribute would not — pydantic serializes a nested model through its declared type), and a merge sums it. The provider test builds a real `ChatCompletion` and skips when the openai extra is not installed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every major provider now tells you how much of the prompt it served from its cache, and
LLMUsagehad nowhere to put that: it is three ints, so providers dropped the number at the point where they read the response. Consumers that bill or report usage (the xac coding harness is the one that prompted this) could not show a cache hit rate at all.cached_tokensdefaults to 0, which is what makes the field non-breaking: responses built without it (older code, plain dicts, MockLLM scripts) simply validate and read 0. Two rules the field carries by convention, both stated in its docstring:prompt_tokens, never an extra spend, sototal_tokensstays prompt + completion;Each provider reads its own gateway's name for the same figure: OpenAI-compat
usage.prompt_tokens_details.cached_tokens(the details object is absent when nothing was cached, hence thegetattrchain), Anthropiccache_read_input_tokens(billed cache reads are already counted insideinput_tokens), Geminicached_content_token_count, and Bedrock, which carries the Anthropic name in its raw usage dict.LLMResponse.mergesums the count alongside the others, or a merged response would lose it.UsageEventinherits the field for free — it extendsLLMUsage, soSimpleToolOrchestrator'sUsageEvent(**usage.model_dump())now emits it.Tests cover the two properties consumers actually depend on: the count survives an
LLMResponsedump/reload round-trip (a subclass attribute would not — pydantic serializes a nested model through its declared type), and a merge sums it. The provider test builds a realChatCompletionand skips when the openai extra is not installed.