Skip to content

Bound markdown chunks - #370

Open
abelfx wants to merge 7 commits into
singnet:mainfrom
abelfx:OMEGA-369-bound-markdown-chunks
Open

abelfx wants to merge 7 commits into
singnet:mainfrom
abelfx:OMEGA-369-bound-markdown-chunks

Conversation

@abelfx

@abelfx abelfx commented Sep 29, 2026 •

Copy link
Copy Markdown

Description

Fixes #369.
Fixes #378.

_chunk_markdown returned a heading-less file as one chunk and skipped MAX_CHUNK_CHARS. A file with headings could still exceed the cap when a single paragraph was longer than the limit, because the splitter only cuts on blank lines.

  • A file with no # heading is now one section and goes through the same size pass as headed files.
  • Anything still over MAX_CHUNK_CHARS after the paragraph split is cut on a newline or space, or mid-token when there is no whitespace.
  • A warning is logged when a file has no headings, and when a character-level cut is applied.
  • The paragraph loop counts the "\n\n" join against the cap. Without this it built chunks of 6001–6002 characters, and the character-level cut then left a 1–13 character fragment that was embedded as its own chunk.
  • Text before the first heading is added as a section with the filename as its breadcrumb, and HEADING_RE matches [ \t]+ instead of \s+ so a lone # on its own line is no longer read as a heading (_chunk_markdown drops all text before the first # heading #378).

How Has This Been Tested?

python3 -m pytest tests/test_rag_chunking.py

  • A heading-less file larger than MAX_CHUNK_CHARS is split, and every chunk stays within the cap.
  • The same text with one heading stays within the cap.
  • One paragraph with no blank lines is cut.
  • A string with no whitespace is cut mid-token and the pieces join back to the original text.
  • Cuts land on word boundaries when spaces are present.
  • Heading-less chunks keep the filename as the breadcrumb.
  • A short heading-less file stays one chunk.
  • Heading breadcrumbs (Top, Top > Nested, Second) are unchanged.
  • Warnings are logged for a heading-less file and for a character-level cut.
  • Two paragraphs landing within the separator of the cap produce [2999, 3000], not [5995, 5].
  • Text before the first heading appears in the chunks, and a long preamble becomes its own chunk.
  • A bare # line is not treated as a heading.

Checklist

  • PR contains autogenerated code
  • Self-review completed
  • Test scenarios above are passed with the version of the code from PR

@alyona-snet alyona-snet changed the title [OMEGA-369] Bound markdown chunks Bound markdown chunks Sep 30, 2026
@abelfx abelfx changed the title Bound markdown chunks [OMEGA 369] Bound markdown chunks Oct 1, 2026
@abelfx abelfx changed the title [OMEGA 369] Bound markdown chunks [OMEGA-369] Bound markdown chunks Oct 1, 2026
@timur-ashkenov

Copy link
Copy Markdown
Collaborator

Issue was reproduced. Have 0 questions for PR. Ready to QA.

@alyona-snet

Copy link
Copy Markdown
Collaborator

@vsbogd Can it be covered in #358 or we can apply this changes independently?

@MartinEbner

Copy link
Copy Markdown

Thanks for picking this up so quickly. I ran this branch's _chunk_markdown over the corpus that
produced #369 — 32 book-length markdown files, 22 of them without any ATX heading — and compared it
with main's.

It does what it says:

main this PR
chunks over MAX_CHUNK_CHARS 36 0
largest heading-less file (4.4 MB) 1 chunk of 4,417,137 chars 981 chunks, largest 6,000

And it loses nothing: the whitespace-normalised text of every file's chunks is identical to main's
on all 32 files.

One thing it introduces. The run produces 40 chunks under 20 characters, where main produces none;
the smallest are a single character. They come from the paragraph loop, which compares
len(chunk_text) + len(p) against the cap but then joins with "\n\n", so it builds chunks of
6,001–6,002 characters. On main that was a harmless overshoot. Here _cut_to_limit splits each of
them into a ~6,000-character chunk plus a 1–13-character fragment, and each fragment is embedded as
a chunk of its own.

Counting the separator removes them:

if chunk_text and len(chunk_text) + 2 + len(p) > MAX_CHUNK_CHARS:
as is with + 2
chunks under 20 characters 40 0
chunks over MAX_CHUNK_CHARS 0 0
indexed text identical to main 32 / 32 32 / 32

A unit test is unlikely to hit this by accident, because it needs two paragraphs landing within two
characters of the cap. This one does:

a = " ".join(["word"] * 600)            # 2,999 chars
b = " ".join(["word"] * 600) + "s"      # 3,000 chars
chunks = _chunk_markdown("# H\n\n" + a + "\n\n" + b, "t.md")
print([len(c["text"]) for c in chunks])
# as is:      [5995, 5]
# with + 2:   [2999, 3000]

Separately, while checking for lost text I found that main drops everything before the first
heading — older than this PR and untouched by it, so I've filed it as #378 rather than add it here.
The fix belongs in _sections_from_headings, so it may be convenient to fold into this PR, but that
is your call.

@abelfx

abelfx commented Oct 5, 2026

Copy link
Copy Markdown
Author

Thanks — confirmed both. The + 2 turns your case from [5995, 5] into [2999, 3000], and no chunk needs a character-level cut there anymore.

I folded #378 in. with this PR a heading-less file is indexed in full, but a file with one stray # still lost everything before it, so the two seemed better landed together. The regex change is included: HEADING_RE.search("#\n35\n") previously returned '35' as a heading title and now returns no match.

Both are covered by tests in tests/test_rag_chunking.py.

@MartinEbner

Copy link
Copy Markdown

A small heads-up: the push on 7 October contains only the merge from main, so the three changes
from your comment of 5 October aren't in the branch yet. At e226b41:

  • the two-paragraph case still comes out as [5995, 5] (the + 2 isn't there);
  • "Text before the first heading.\n\n# Heading\n\nText after the heading.\n" still comes back as
    ['Text after the heading.'], so _chunk_markdown drops all text before the first # heading #378 isn't covered;
  • HEADING_RE still uses \s+, and HEADING_RE.search("#\n35\n").group(2) still returns '35'.

tests/test_rag_chunking.py at that commit has no tests for them either. Probably a commit that
stayed local.

@abelfx

abelfx commented Oct 8, 2026

Copy link
Copy Markdown
Author

Pushed now. The description was ahead of the branch; the fixes are in 52ae97b. Sorry about that.

@alyona-snet alyona-snet changed the title [OMEGA-369] Bound markdown chunks Bound markdown chunks Oct 9, 2026
@alyona-snet alyona-snet added backlog The issue has been included in the backlog discussion labels Oct 9, 2026
@vsbogd

vsbogd commented Oct 9, 2026

Copy link
Copy Markdown
Member

@vsbogd Can it be covered in #358 or we can apply this changes independently?

No #358 doesn't cover this

@MartinEbner

Copy link
Copy Markdown

Thanks, confirmed at 52ae97b. All three cases come out right now: the text before the first
heading is kept, the two-paragraph case gives [2999, 3000], and a bare # line is no longer read
as a heading. Over the same 32-file corpus, no chunk exceeds MAX_CHUNK_CHARS, none is under 20
characters, and every file's text is indexed in full, apart from the heading lines, which go to the
breadcrumbs. On main none of the 32 is both complete and within the cap: 22 come out as one
oversized chunk each, and the other 10 lose the text before their first heading.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backlog The issue has been included in the backlog discussion

Projects

None yet

5 participants