Repository navigation
feat(llm): reasoning levels (LLMOptions.reasoning_effort) - #42
Merged
Merged
Conversation
Thinking models all have a "how much may you think" dial and none of them
spell it the same way. `LLMOptions` had temperature, top_p, max_tokens,
stop_sequences, functions and vendor_specific, so a caller who wanted to ask
for less thinking either hand-wrote a provider-specific payload or could not
ask at all — and a harness that streams the model's reasoning (the xac coding
agent, which prompted this) has no way to label or control what it is showing.
`ReasoningEffort` + `LLMOptions.reasoning_effort` is the framework vocabulary.
It defaults to None, which is what makes the addition non-breaking: an unset
level sends no field at all, so every existing request is byte-identical. The
ladder is `none/minimal/low/medium/high/xhigh/max` — the union of what
OpenAI-compatible APIs accept today; `none` is the explicit off switch for
models that reason by default.
The mapping lives in one new module, `core/models/reasoning.py`, rather than
in each provider, because the hard part is not the name but the *shape*:
Anthropic's `thinking.budget_tokens` is deprecated on Claude 4.6 and rejected
by 4.7+, which take `thinking:{type:adaptive}` + `output_config.effort`;
Gemini takes a `thinkingLevel` enum on 3+ and a `thinking_budget` on 2.5;
Bedrock Converse carries Anthropic's fields verbatim but spells the output
ceiling `inferenceConfig.maxTokens`. So a level alone cannot determine what to
send — a module also needs to know which shape its model speaks, which is why
each provider takes a `reasoning_mode` config key defaulting to the shape its
current models accept.
Three rules every mapping obeys:
- unset sends nothing: a model not asked how much to think uses its own
default, and inventing a value would make a UI's "default" a lie;
- a level the model cannot express is lowered, never dropped — `xhigh` on a
Gemini whose enum stops at HIGH sends HIGH, `minimal` (not on Anthropic's
ladder) sends `low`. Silently sending *no* thinking is the wrong direction
of error when the user asked for more;
- where a model has no off switch, the mapping says so: Gemini 3 cannot
disable thinking, so `none` takes its lowest rung instead of a budget of 0
that a 3.x model rejects.
Budget-shaped requests also get their output ceiling raised to fit: thinking
tokens count toward `max_tokens` and Anthropic requires the budget strictly
below it, so a level that leaves no answer room would truncate or 400. A
`thinking`/`output_config`/`thinking_config` set in config or `vendor_specific`
still wins — someone who wrote it by hand meant it — and `reasoning_mode` is
excluded from each module's `default_kwargs` so it never reaches an API body.
`TextBasedToolCallAdapter` rebuilt `LLMOptions` field-by-field and would have
dropped the level on its way to the wrapped provider; it now forwards it.
Tests are request-builder tests (no key, no network) because the failure being
guarded is silent: a level configured, displayed and then vanished on the way
out. They cover the enum contract, the field landing in both streaming and
non-streaming kwargs, and each provider's shape including the two modes, the
clamps, the ceiling arithmetic and `vendor_specific` precedence.
…dules The first cut put every provider's spelling of `reasoning_effort` in `core/models/reasoning.py`, which made the framework's data layer the one place that knows how Anthropic, Gemini and Bedrock each phrase a request. That is exactly what the protocol/implementation split exists to keep out of the core: `OpenAILLM` already owns OpenAI's message format and usage fields, and nothing about it leaks into `core`. A thinking dial is the same kind of detail. So the core now carries only the vocabulary — `ReasoningEffort` and the option field — and each implementation owns its mapping: - `llm/google.py`: `_thinking_config()` plus the two tables it reads (the `thinkingLevel` clamp and the budget ladder), next to the `_prepare_config` that already maps every other option to Gemini's names; - `llm/claude_thinking.py`: the Anthropic-shaped mapping, shared with Bedrock's Converse `additionalModelRequestFields`. It is a module rather than a function on `AnthropicLLM` because a provider must not import a sibling provider to reuse anything: `llm/__init__.py` swallows the `ImportError` of any module whose SDK is missing, so `from .anthropic import …` fails without the `anthropic` extra even when the code needed nothing from it; - `llm/openai.py`: unchanged, because there is nothing to move — the framework field already is that API's field name. The three rules a provider owes the option (unset sends nothing, an inexpressible level is lowered not dropped, a model with no off switch takes its lowest rung) were the shared module's docstring; they now live on `ReasoningEffort` itself, where an implementer of the protocol reads them, and each mapping points back instead of restating them. `level_value()` — the enum-or-string normalization — is gone. `ReasoningEffort` is a `str` enum, so a member and its name are the same key for a dict lookup, the same value for a membership test, and serialize identically; the helper was guarding a distinction that does not exist. Tests split along the same seam: `test_claude_thinking.py` covers the mapping and both providers' request builders, `test_gemini_thinking.py` covers Gemini's two shapes. The provider half of `test_reasoning_effort.py` (the OpenAI-compat verbatim pass-through) is untouched, and nothing outside the provider modules imports a mapping now.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Thinking models all have a reasoning level dial, and none of them spell it the same way.
LLMOptionshadtemperature,top_p,max_tokens,stop_sequences,functionsandvendor_specific, so asking for less thinking meant either hand-writing a provider-specific payload or not asking at all.This adds one framework-level option and lets each provider implementation translate it into the shape its API expects.
What's in the core
ReasoningEffort—none / minimal / low / medium / high / xhigh / max, the union of what these APIs accept today — and one optional field. Unset is the default and sends nothing at all, which is what makes this non-breaking: arequest that doesn't ask for a level carries no reasoning field (verified on all four providers' request builders, see Testing).
The core carries the vocabulary only. How a level reaches the wire is each provider's business, the same way its message format and usage fields already are — the protocol/implementation split is the reason. Three rules the option
asks of every implementation are documented on the enum itself, where someone writing a provider will read them:
default; inventing a value would make a UI's "default" a lie;
more thinking and silently getting none is the wrong direction of error;
NONE— take its lowest rung instead ofsending a disable flag it rejects.
What each provider sends
reasoning_modeOpenAILLMreasoning_effort— already that API's own name, so there is nothing to mapAnthropicLLMthinking: {type: adaptive}+output_config.effort, orthinking: {type: enabled, budget_tokens}+ a raisedmax_tokens"effort"(default) /"budget"BedrockLLMadditionalModelRequestFields; the ceiling isinferenceConfig.maxTokens"effort"/"budget"/"off"GoogleLLMthinkingConfig.thinking_level(enum) orthinkingConfig.thinking_budget(tokens)"level"(default) /"budget"The shape is a constructor choice because the same vendor changed spelling between generations: Anthropic's
thinking.budget_tokensis deprecated on Claude 4.6 and 400-rejected by 4.7+, which takeoutput_config.effort; Gemini'sthinkingLevelerrors on 2.5, which take a budget. A level alone cannot pick the shape, so the module asks which one its model speaks. Defaults are the shape current models accept — the deprecated form is the opt-in.Clamping, per the rules above:
minimal→lowon Claude (not on its ladder),xhigh/max→HIGHon Gemini (its enum stops there),none→LOWon Gemini 3 (which cannot disable thinking) butthinking: {type: "disabled"}onClaude (which can). Budget-shaped requests get their output ceiling raised to fit the budget: thinking tokens are billed against
max_tokens, which must be strictly above it, or the response comes back truncated.vendor_specificalways wins for the same field — someone who wrotethinkingorthinking_configby hand meant it, so the level is applied first and only fills a gap.reasoning_modeis excluded from each module'sdefault_kwargs, so it never leaks into a request body.