Skip to content

Detect HTML in maybe_is_html regardless of case or leading whitespace - #1358

Open
MohammadHijjawi97 wants to merge 1 commit into
Future-House:mainfrom
MohammadHijjawi97:fix-maybe-is-html-case
Open

MohammadHijjawi97 wants to merge 1 commit into
Future-House:mainfrom
MohammadHijjawi97:fix-maybe-is-html-case

Conversation

@MohammadHijjawi97

Copy link
Copy Markdown

maybe_is_html compared the first 4 bytes against {b"<htm", b"<!DO", b"<xsl", b"<!X"}, so only an uppercase <!DOCTYPE or a lowercase <htm was recognized. These inputs were all missed:

  • a lowercase <!doctype html>, the common HTML5 form;
  • an uppercase <HTML> tag;
  • a UTF-8 BOM;
  • leading whitespace or newlines.

Docs.aadd_file wrote such files to a .txt temp file and indexed the raw markup instead of parsing it with html2text.

>>> maybe_is_html(BytesIO(b"<!doctype html><html></html>"))
False

This change reads a small prefix, drops a UTF-8 BOM and leading whitespace, and compares case-insensitively.

Added a parametrized offline test. It covers lowercase and uppercase doctypes and tags, leading whitespace and a BOM, plus PDF and plain-text negatives, and checks that the file is rewound. The four new positive cases fail on main and pass here. ruff and black are clean.

`maybe_is_html` compared the first four bytes against uppercase-only
signatures, so a lowercase `<!doctype html>` (the common HTML5 form),
an uppercase `<HTML>` tag, a UTF-8 BOM, or leading whitespace all went
undetected. `Docs.aadd_file` then saved such files as `.txt` and indexed
the raw markup instead of parsing it with `html2text`.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant