Bug description
max_tool_steps is enforced on the STT pipeline tool loop but not on the realtime one. With a realtime model, a chain of tool calls keeps issuing tool calls for as many rounds as the tools produce output — the documented ceiling never applies.
livekit-agents/livekit/agents/voice/agent_activity.py, current main (15b4bc84c):
- pipeline,
_pipeline_reply_task_impl — gate present:
max_steps_reached = speech_handle.num_steps >= self._session.options.max_tool_steps + 1
if max_steps_reached:
logger.warning("maximum number of function calls steps reached, ...")
speech_handle._num_steps += 1
- realtime,
_realtime_generation_task_impl — increments the same counter with no gate anywhere in the task:
speech_handle._num_steps += 1
The reply that continues the chain sets tool_choice="auto" ("none" only when draining), so the model is free to call another tool, and max_tool_steps is read in exactly one place in the package — agent_activity.py:3930, the pipeline branch.
Why it matters
The option is documented without a modality exception:
max_tool_steps (int): Maximum consecutive tool calls per LLM turn. Default 3. — voice/agent_session.py:474
So the same AgentSession(max_tool_steps=3) bounds tool chaining for an STT pipeline model and does nothing for a realtime one. A tool that keeps returning output runs the turn indefinitely on realtime, with no warning, since the logger.warning("maximum number of function calls steps reached") line is on the pipeline side only. Each round re-enters generate_reply on the realtime session's own billing path.
I read the counter's intent from #7431, where longcw explained that the + 1 is deliberate: the first generation answers the user and is not a tool step, so max_tool_steps=N permits the first reply plus N tool replies. I am not raising the off-by-one again — the point here is only that the realtime path has no ceiling of either size.
Direction question
I did not want to guess at the fix, because realtime models differ here: _realtime_reply_task_impl reads capabilities.per_response_tool_choice, and for a model without it tool_choice has to be set on the session via update_options with tools swapped by update_tools. Which is intended for the ceiling?
- Match the pipeline — compute
max_steps_reached from num_steps in the realtime task and pass tool_choice="none" on the reply after the last permitted tool round. Straightforward for models with per_response_tool_choice; needs the session-level path for the others.
- Realtime tools are expected to terminate themselves, and the option applies to the pipeline only. Then the fix is a docs change stating the exception.
- Something I have not considered.
I lean towards 1, since the option is documented as a general per-turn ceiling and realtime is the stricter modality — but this is your call, and I would rather agree on the semantics than push a patch that guesses. Happy to write the regression test either way once it is settled.
Verification detail and scope
I did not stand up a live realtime session, so I have not observed an unbounded run end to end. What is verifiable statically is that _num_steps is incremented at agent_activity.py:4639 and that no code between there and the end of _realtime_generation_task_impl reads max_tool_steps. Grepping max_tool_steps across the package returns only agent_activity.py:3930, agent_session.py (declaration, docstring, plumbing), remote_session.py and report.py (both reporting the configured value), plus two test files that only set the option.
I searched for an existing report and did not find one covering realtime and max_tool_steps; #7431 is the pipeline off-by-one and is closed.
Bug description
max_tool_stepsis enforced on the STT pipeline tool loop but not on the realtime one. With a realtime model, a chain of tool calls keeps issuing tool calls for as many rounds as the tools produce output — the documented ceiling never applies.livekit-agents/livekit/agents/voice/agent_activity.py, currentmain(15b4bc84c):_pipeline_reply_task_impl— gate present:_realtime_generation_task_impl— increments the same counter with no gate anywhere in the task:The reply that continues the chain sets
tool_choice="auto"("none"only when draining), so the model is free to call another tool, andmax_tool_stepsis read in exactly one place in the package —agent_activity.py:3930, the pipeline branch.Why it matters
The option is documented without a modality exception:
So the same
AgentSession(max_tool_steps=3)bounds tool chaining for an STT pipeline model and does nothing for a realtime one. A tool that keeps returning output runs the turn indefinitely on realtime, with no warning, since thelogger.warning("maximum number of function calls steps reached")line is on the pipeline side only. Each round re-entersgenerate_replyon the realtime session's own billing path.I read the counter's intent from #7431, where
longcwexplained that the+ 1is deliberate: the first generation answers the user and is not a tool step, somax_tool_steps=Npermits the first reply plus N tool replies. I am not raising the off-by-one again — the point here is only that the realtime path has no ceiling of either size.Direction question
I did not want to guess at the fix, because realtime models differ here:
_realtime_reply_task_implreadscapabilities.per_response_tool_choice, and for a model without ittool_choicehas to be set on the session viaupdate_optionswith tools swapped byupdate_tools. Which is intended for the ceiling?max_steps_reachedfromnum_stepsin the realtime task and passtool_choice="none"on the reply after the last permitted tool round. Straightforward for models withper_response_tool_choice; needs the session-level path for the others.I lean towards 1, since the option is documented as a general per-turn ceiling and realtime is the stricter modality — but this is your call, and I would rather agree on the semantics than push a patch that guesses. Happy to write the regression test either way once it is settled.
Verification detail and scope
I did not stand up a live realtime session, so I have not observed an unbounded run end to end. What is verifiable statically is that
_num_stepsis incremented atagent_activity.py:4639and that no code between there and the end of_realtime_generation_task_implreadsmax_tool_steps. Greppingmax_tool_stepsacross the package returns onlyagent_activity.py:3930,agent_session.py(declaration, docstring, plumbing),remote_session.pyandreport.py(both reporting the configured value), plus two test files that only set the option.I searched for an existing report and did not find one covering realtime and
max_tool_steps; #7431 is the pipeline off-by-one and is closed.