Describe the bug
The heading walk in _chunk_markdown (src/rag.py; moved into _sections_from_headings by #370)
builds each section from the end of one heading to the start of the next:
for i, m in enumerate(matches):
...
start = m.end()
end = matches[i + 1].start() if i + 1 < len(matches) else len(text)
body = text[start:end].strip()
Nothing ever reads text[:matches[0].start()]. Whatever precedes the first heading is not put in any
section, so it never reaches a chunk and is never indexed.
To Reproduce
from rag import _chunk_markdown
doc = "Text before the first heading.\n\n# Heading\n\nText after the heading.\n"
print([c["text"] for c in _chunk_markdown(doc, "t.md")])
# ['Text after the heading.']
On real input, over a corpus of 32 book-length markdown files, 10 lose text this way. Comparing the
chunks with the source, what each of those 10 keeps is exactly the concatenated text between
headings. The heading lines themselves go to the breadcrumb by design, so the text actually lost is
the region before the first heading:
| ATX headings in the file |
text lost |
| 1 |
78.3% (a 1.4 MB file) |
| 1 |
48.7% |
| 1 |
48.6% |
| 6 |
47.4% |
| 1 |
46.0% |
| 9 |
28.8% |
| 13 |
27.2% |
| 2, 6, 6 |
19.5%, 18.9%, 8.7% |
Nothing in the log shows it. The file is reported as indexed N chunks with a plausible N, and
the missing part is simply absent.
Why ordinary text triggers it
In all four single-heading files above, the one match is not a heading. Two are a sentence that
happened to begin with # ; the other two are a # standing alone on a line. The second case gets
through because HEADING_RE uses \s+ after the hashes, and \s also matches a newline:
HEADING_RE.search("#\n35\n").group(2) # '35' — the next line becomes the heading's title
In the worst case above, that is the whole story: a lone # line 78% of the way into a book, and
everything before it is dropped.
This interacts with #369 / #370. After #370, a file with no heading is chunked in full. A file
with one stray # line is not: it loses everything before that line. So, with #370 merged, a
single stray # becomes the difference between a book that is fully indexed and one that is three
quarters missing.
Expected behavior
Text before the first heading is indexed like any other section.
Possible fix
Add the preamble as a first section before the walk, with the filename as its breadcrumb — the same
breadcrumb a heading-less file gets:
sections = []
stack = {} # level -> heading text
preamble = text[:matches[0].start()].strip()
if preamble:
sections.append({"text": preamble, "breadcrumb": filename, "heading": ""})
("heading": "" keeps the table-of-contents filter, which reads s["heading"], working.) The
existing merge and size passes then treat it like any other section — a short preamble is merged
into the first section, a long one is split.
Tested over the same 32 files, counting a file as fully indexed when every character except the
heading lines themselves (which go to the breadcrumb) appears in its chunks:
| variant |
files fully indexed |
current main |
22 / 32 |
main + preamble section |
32 / 32 |
| #370 + preamble section |
32 / 32 |
Separately, and smaller: changing \s+ to [ \t]+ in HEADING_RE stops a bare # line from being
read as a heading with the following line as its title. A line genuinely written as # Title still
matches; # followed by a line break no longer does. That reduces how often this is triggered, but
the preamble section is what fixes it.
The # Table of Contents filter also removes a section, but that one is deliberate, so it is not
part of this report.
Environment
Current main, unchanged by #370 — verified on the same corpus: with and without #370 the indexed
text is identical on all 32 files. src/rag.py only; independent of OS, channel, LLM provider and
embedding provider.
Describe the bug
The heading walk in
_chunk_markdown(src/rag.py; moved into_sections_from_headingsby #370)builds each section from the end of one heading to the start of the next:
Nothing ever reads
text[:matches[0].start()]. Whatever precedes the first heading is not put in anysection, so it never reaches a chunk and is never indexed.
To Reproduce
On real input, over a corpus of 32 book-length markdown files, 10 lose text this way. Comparing the
chunks with the source, what each of those 10 keeps is exactly the concatenated text between
headings. The heading lines themselves go to the breadcrumb by design, so the text actually lost is
the region before the first heading:
Nothing in the log shows it. The file is reported as
indexed N chunkswith a plausibleN, andthe missing part is simply absent.
Why ordinary text triggers it
In all four single-heading files above, the one match is not a heading. Two are a sentence that
happened to begin with
#; the other two are a#standing alone on a line. The second case getsthrough because
HEADING_REuses\s+after the hashes, and\salso matches a newline:In the worst case above, that is the whole story: a lone
#line 78% of the way into a book, andeverything before it is dropped.
This interacts with #369 / #370. After #370, a file with no heading is chunked in full. A file
with one stray
#line is not: it loses everything before that line. So, with #370 merged, asingle stray
#becomes the difference between a book that is fully indexed and one that is threequarters missing.
Expected behavior
Text before the first heading is indexed like any other section.
Possible fix
Add the preamble as a first section before the walk, with the filename as its breadcrumb — the same
breadcrumb a heading-less file gets:
(
"heading": ""keeps the table-of-contents filter, which readss["heading"], working.) Theexisting merge and size passes then treat it like any other section — a short preamble is merged
into the first section, a long one is split.
Tested over the same 32 files, counting a file as fully indexed when every character except the
heading lines themselves (which go to the breadcrumb) appears in its chunks:
mainmain+ preamble sectionSeparately, and smaller: changing
\s+to[ \t]+inHEADING_REstops a bare#line from beingread as a heading with the following line as its title. A line genuinely written as
# Titlestillmatches;
#followed by a line break no longer does. That reduces how often this is triggered, butthe preamble section is what fixes it.
The
# Table of Contentsfilter also removes a section, but that one is deliberate, so it is notpart of this report.
Environment
Current
main, unchanged by #370 — verified on the same corpus: with and without #370 the indexedtext is identical on all 32 files.
src/rag.pyonly; independent of OS, channel, LLM provider andembedding provider.