Skip to content

feat(llm): reasoning levels (LLMOptions.reasoning_effort) - #42

Merged
treo merged 2 commits into
mainfrom
paul/reasoning-effort
Oct 3, 2026
Merged

treo merged 2 commits into
mainfrom
paul/reasoning-effort

Conversation

@treo

@treo treo commented Oct 2, 2026

Copy link
Copy Markdown
Member

Thinking models all have a reasoning level dial, and none of them spell it the same way. LLMOptions had temperature, top_p, max_tokens, stop_sequences, functions and vendor_specific, so asking for less thinking meant either hand-writing a provider-specific payload or not asking at all.

This adds one framework-level option and lets each provider implementation translate it into the shape its API expects.

from xaibo.core.models.llm import LLMOptions

response = await llm.generate(messages, LLMOptions(reasoning_effort="high"))

What's in the core

ReasoningEffort — none / minimal / low / medium / high / xhigh / max, the union of what these APIs accept today — and one optional field. Unset is the default and sends nothing at all, which is what makes this non-breaking: a
request that doesn't ask for a level carries no reasoning field (verified on all four providers' request builders, see Testing).

The core carries the vocabulary only. How a level reaches the wire is each provider's business, the same way its message format and usage fields already are — the protocol/implementation split is the reason. Three rules the option
asks of every implementation are documented on the enum itself, where someone writing a provider will read them:

  • unset sends nothing — a model not asked how much to think uses its own
    default; inventing a value would make a UI's "default" a lie;
  • a level the model cannot express is lowered, never dropped — asking for
    more thinking and silently getting none is the wrong direction of error;
  • a model with no off switch has no NONE — take its lowest rung instead of
    sending a disable flag it rejects.

What each provider sends

provider fields reasoning_mode
OpenAILLM reasoning_effort — already that API's own name, so there is nothing to map —
AnthropicLLM thinking: {type: adaptive} + output_config.effort, or thinking: {type: enabled, budget_tokens} + a raised max_tokens "effort" (default) / "budget"
BedrockLLM the same Anthropic fields, in additionalModelRequestFields; the ceiling is inferenceConfig.maxTokens "effort" / "budget" / "off"
GoogleLLM thinkingConfig.thinking_level (enum) or thinkingConfig.thinking_budget (tokens) "level" (default) / "budget"

The shape is a constructor choice because the same vendor changed spelling between generations: Anthropic's thinking.budget_tokens is deprecated on Claude 4.6 and 400-rejected by 4.7+, which take output_config.effort; Gemini's thinkingLevel errors on 2.5, which take a budget. A level alone cannot pick the shape, so the module asks which one its model speaks. Defaults are the shape current models accept — the deprecated form is the opt-in.

Clamping, per the rules above: minimal → low on Claude (not on its ladder), xhigh/max → HIGH on Gemini (its enum stops there), none → LOW on Gemini 3 (which cannot disable thinking) but thinking: {type: "disabled"} on
Claude (which can). Budget-shaped requests get their output ceiling raised to fit the budget: thinking tokens are billed against max_tokens, which must be strictly above it, or the response comes back truncated.

vendor_specific always wins for the same field — someone who wrote thinking or thinking_config by hand meant it, so the level is applied first and only fills a gap. reasoning_mode is excluded from each module's default_kwargs, so it never leaks into a request body.

treo added 2 commits October 2, 2026 16:20
Thinking models all have a "how much may you think" dial and none of them
spell it the same way. `LLMOptions` had temperature, top_p, max_tokens,
stop_sequences, functions and vendor_specific, so a caller who wanted to ask
for less thinking either hand-wrote a provider-specific payload or could not
ask at all — and a harness that streams the model's reasoning (the xac coding
agent, which prompted this) has no way to label or control what it is showing.

`ReasoningEffort` + `LLMOptions.reasoning_effort` is the framework vocabulary.
It defaults to None, which is what makes the addition non-breaking: an unset
level sends no field at all, so every existing request is byte-identical. The
ladder is `none/minimal/low/medium/high/xhigh/max` — the union of what
OpenAI-compatible APIs accept today; `none` is the explicit off switch for
models that reason by default.

The mapping lives in one new module, `core/models/reasoning.py`, rather than
in each provider, because the hard part is not the name but the *shape*:
Anthropic's `thinking.budget_tokens` is deprecated on Claude 4.6 and rejected
by 4.7+, which take `thinking:{type:adaptive}` + `output_config.effort`;
Gemini takes a `thinkingLevel` enum on 3+ and a `thinking_budget` on 2.5;
Bedrock Converse carries Anthropic's fields verbatim but spells the output
ceiling `inferenceConfig.maxTokens`. So a level alone cannot determine what to
send — a module also needs to know which shape its model speaks, which is why
each provider takes a `reasoning_mode` config key defaulting to the shape its
current models accept.

Three rules every mapping obeys:

- unset sends nothing: a model not asked how much to think uses its own
  default, and inventing a value would make a UI's "default" a lie;
- a level the model cannot express is lowered, never dropped — `xhigh` on a
  Gemini whose enum stops at HIGH sends HIGH, `minimal` (not on Anthropic's
  ladder) sends `low`. Silently sending *no* thinking is the wrong direction
  of error when the user asked for more;
- where a model has no off switch, the mapping says so: Gemini 3 cannot
  disable thinking, so `none` takes its lowest rung instead of a budget of 0
  that a 3.x model rejects.

Budget-shaped requests also get their output ceiling raised to fit: thinking
tokens count toward `max_tokens` and Anthropic requires the budget strictly
below it, so a level that leaves no answer room would truncate or 400. A
`thinking`/`output_config`/`thinking_config` set in config or `vendor_specific`
still wins — someone who wrote it by hand meant it — and `reasoning_mode` is
excluded from each module's `default_kwargs` so it never reaches an API body.

`TextBasedToolCallAdapter` rebuilt `LLMOptions` field-by-field and would have
dropped the level on its way to the wrapped provider; it now forwards it.

Tests are request-builder tests (no key, no network) because the failure being
guarded is silent: a level configured, displayed and then vanished on the way
out. They cover the enum contract, the field landing in both streaming and
non-streaming kwargs, and each provider's shape including the two modes, the
clamps, the ceiling arithmetic and `vendor_specific` precedence.
…dules

The first cut put every provider's spelling of `reasoning_effort` in
`core/models/reasoning.py`, which made the framework's data layer the one place
that knows how Anthropic, Gemini and Bedrock each phrase a request. That is
exactly what the protocol/implementation split exists to keep out of the core:
`OpenAILLM` already owns OpenAI's message format and usage fields, and nothing
about it leaks into `core`. A thinking dial is the same kind of detail.

So the core now carries only the vocabulary — `ReasoningEffort` and the option
field — and each implementation owns its mapping:

- `llm/google.py`: `_thinking_config()` plus the two tables it reads (the
  `thinkingLevel` clamp and the budget ladder), next to the `_prepare_config`
  that already maps every other option to Gemini's names;
- `llm/claude_thinking.py`: the Anthropic-shaped mapping, shared with Bedrock's
  Converse `additionalModelRequestFields`. It is a module rather than a function
  on `AnthropicLLM` because a provider must not import a sibling provider to
  reuse anything: `llm/__init__.py` swallows the `ImportError` of any module
  whose SDK is missing, so `from .anthropic import …` fails without the
  `anthropic` extra even when the code needed nothing from it;
- `llm/openai.py`: unchanged, because there is nothing to move — the framework
  field already is that API's field name.

The three rules a provider owes the option (unset sends nothing, an
inexpressible level is lowered not dropped, a model with no off switch takes
its lowest rung) were the shared module's docstring; they now live on
`ReasoningEffort` itself, where an implementer of the protocol reads them, and
each mapping points back instead of restating them.

`level_value()` — the enum-or-string normalization — is gone. `ReasoningEffort`
is a `str` enum, so a member and its name are the same key for a dict lookup,
the same value for a membership test, and serialize identically; the helper was
guarding a distinction that does not exist.

Tests split along the same seam: `test_claude_thinking.py` covers the mapping
and both providers' request builders, `test_gemini_thinking.py` covers Gemini's
two shapes. The provider half of `test_reasoning_effort.py` (the OpenAI-compat
verbatim pass-through) is untouched, and nothing outside the provider modules
imports a mapping now.
@treo
treo requested a review from wmeddie October 2, 2026 19:38
@treo
treo merged commit 56a04d3 into main Oct 3, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant