Skip to content

fix(readers): reject chunk overlap that prevents forward progress - #1354

Open
linhongyu510 wants to merge 1 commit into
Future-House:mainfrom
linhongyu510:fix/chunker-overlap-infinite-loop
Open

linhongyu510 wants to merge 1 commit into
Future-House:mainfrom
linhongyu510:fix/chunker-overlap-infinite-loop

Conversation

@linhongyu510

Copy link
Copy Markdown

What hangs

chunk_pdf and chunk_code_text advance their buffers with the same expression:

buffer = buffer[chunk_chars - overlap :]

When overlap == chunk_chars, the slice starts at zero. The buffer never shrinks, so both while len(buffer) > chunk_chars loops run forever.

ParsingSettings.reader_config is an unconstrained dict[str, Any] and is passed directly into these functions by read_doc, so no earlier validation prevents this configuration.

Reproduction with the real functions

I ran each function in its own spawned process and used a parent-side timeout so a hung child could not stall the probe itself. The same 5,000-character input and chunk_chars=100 were used for the control and boundary cases:

pdf  overlap=50:  returned 99 chunks
pdf  overlap=100: FUNCTION DID NOT RETURN in 20s
code overlap=50:  returned 99 chunks
code overlap=100: FUNCTION DID NOT RETURN in 20s

This is not merely a slow edge case: the step is exactly zero, so no iteration can make progress.

Fix

Add one shared chunk-parameter validator:

  • chunk_chars > 0
  • 0 <= overlap < chunk_chars

and call it from chunk_pdf, chunk_text, and chunk_code_text.

The infinite loop exists in the PDF/Office/image and code paths. chunk_text does not use the same shrinking loop, but validating it too is intentional: the same reader_config is dispatched by file extension, and an invalid configuration should not hang PDFs/code while being silently accepted for .txt/.html files.

chunk_chars=0 remains supported through its documented meaning: read_doc handles it before dispatch and returns the full document without calling a chunker.

The code fails fast with an actionable ValueError; it does not silently clamp or reinterpret the user's configuration.

Tests

One parameterized regression test covers all three chunkers and four invalid boundaries:

  • overlap equals chunk size (the non-terminating case)
  • overlap larger than chunk size
  • negative overlap
  • zero chunk size when directly invoking a chunker

That is 12 explicit cases. The test has a 5-second timeout so a reintroduction of the zero-step loop cannot hang CI indefinitely.

Reverse verification:

  • With readers.py restored to baseline, the six non-hanging invalid cases (larger/negative overlap across all three chunkers) all fail with DID NOT RAISE ValueError.
  • The two real zero-step loops were separately verified in isolated processes as shown above.
  • With the fix applied, all 12 cases pass in under one millisecond each.

Verification

Target invalid-boundary tests                   12 passed
Target tests + existing missing-page regression 13 passed
Valid read/chunk paths                           3 passed
Ruff 0.14.4 (repo-pinned)                        passed
Black 26.1.0 (repo-pinned)                       passed
mypy src/paperqa/readers.py                      passed

Full tests/test_paperqa.py, compared against an unmodified checkout in the same environment:

baseline: 121 passed, 51 failed, 9 errors, 5 subtests passed
fixed:    133 passed, 51 failed, 9 errors, 5 subtests passed

The failure/error sets are identical. They are existing environment-dependent tests requiring LLM credentials or unavailable Office parsing dependencies; this change adds exactly 12 passing tests and no new failures.

I also attempted the complete pre-commit collection. Its first run spent over ten minutes installing the many isolated hook environments, so I stopped that setup rather than reporting it as passed. The code-relevant gates from that configuration (Ruff, Black, and the local mypy hook) were run directly with the pinned versions and passed. git diff --check is clean and uv.lock is unchanged.

AI disclosure

This change was prepared with AI assistance. The defect was identified from the zero-progress slice, reproduced using the real chunkers under isolated process timeouts, and verified against the repository's tests and quality gates by the contributor.

`chunk_pdf` and `chunk_code_text` advance their buffers with
`buffer[chunk_chars - overlap:]`. When overlap equals chunk_chars, the slice
starts at zero, the buffer never shrinks, and both functions loop forever.

The configuration reaches these functions directly through
`ParsingSettings.reader_config`, which is an unconstrained `dict[str, Any]`, so
there is no earlier validation. With chunk_chars=100 and overlap=100, both real
functions failed to return within 20 seconds; with overlap=50 they returned 99
chunks.

Validate the shared chunking contract in all three chunkers:

- chunk_chars must be positive (zero is the documented no-chunking sentinel and
  is handled by read_doc before dispatch)
- overlap must be non-negative and strictly less than chunk_chars

Validating `chunk_text` too keeps the same reader_config semantics across PDF,
text, and code inputs instead of letting an invalid configuration hang only
some file types.

The parameterized regression tests cover all three chunkers and four invalid
boundaries: equal overlap, larger overlap, negative overlap, and zero size.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant