Skip to content

_chunk_markdown drops all text before the first # heading #378

Description

@MartinEbner

Describe the bug

The heading walk in _chunk_markdown (src/rag.py; moved into _sections_from_headings by #370)
builds each section from the end of one heading to the start of the next:

for i, m in enumerate(matches):
    ...
    start = m.end()
    end = matches[i + 1].start() if i + 1 < len(matches) else len(text)
    body = text[start:end].strip()

Nothing ever reads text[:matches[0].start()]. Whatever precedes the first heading is not put in any
section, so it never reaches a chunk and is never indexed.

To Reproduce

from rag import _chunk_markdown

doc = "Text before the first heading.\n\n# Heading\n\nText after the heading.\n"
print([c["text"] for c in _chunk_markdown(doc, "t.md")])
# ['Text after the heading.']

On real input, over a corpus of 32 book-length markdown files, 10 lose text this way. Comparing the
chunks with the source, what each of those 10 keeps is exactly the concatenated text between
headings. The heading lines themselves go to the breadcrumb by design, so the text actually lost is
the region before the first heading:

ATX headings in the file text lost
1 78.3% (a 1.4 MB file)
1 48.7%
1 48.6%
6 47.4%
1 46.0%
9 28.8%
13 27.2%
2, 6, 6 19.5%, 18.9%, 8.7%

Nothing in the log shows it. The file is reported as indexed N chunks with a plausible N, and
the missing part is simply absent.

Why ordinary text triggers it

In all four single-heading files above, the one match is not a heading. Two are a sentence that
happened to begin with # ; the other two are a # standing alone on a line. The second case gets
through because HEADING_RE uses \s+ after the hashes, and \s also matches a newline:

HEADING_RE.search("#\n35\n").group(2)    # '35'  — the next line becomes the heading's title

In the worst case above, that is the whole story: a lone # line 78% of the way into a book, and
everything before it is dropped.

This interacts with #369 / #370. After #370, a file with no heading is chunked in full. A file
with one stray # line is not: it loses everything before that line. So, with #370 merged, a
single stray # becomes the difference between a book that is fully indexed and one that is three
quarters missing.

Expected behavior

Text before the first heading is indexed like any other section.

Possible fix

Add the preamble as a first section before the walk, with the filename as its breadcrumb — the same
breadcrumb a heading-less file gets:

sections = []
stack = {}  # level -> heading text
preamble = text[:matches[0].start()].strip()
if preamble:
    sections.append({"text": preamble, "breadcrumb": filename, "heading": ""})

("heading": "" keeps the table-of-contents filter, which reads s["heading"], working.) The
existing merge and size passes then treat it like any other section — a short preamble is merged
into the first section, a long one is split.

Tested over the same 32 files, counting a file as fully indexed when every character except the
heading lines themselves (which go to the breadcrumb) appears in its chunks:

variant files fully indexed
current main 22 / 32
main + preamble section 32 / 32
#370 + preamble section 32 / 32

Separately, and smaller: changing \s+ to [ \t]+ in HEADING_RE stops a bare # line from being
read as a heading with the following line as its title. A line genuinely written as # Title still
matches; # followed by a line break no longer does. That reduces how often this is triggered, but
the preamble section is what fixes it.

The # Table of Contents filter also removes a section, but that one is deliberate, so it is not
part of this report.

Environment

Current main, unchanged by #370 — verified on the same corpus: with and without #370 the indexed
text is identical on all 32 files. src/rag.py only; independent of OS, channel, LLM provider and
embedding provider.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions