Detect HTML in maybe_is_html regardless of case or leading whitespace - #1358
Open
MohammadHijjawi97 wants to merge 1 commit into
Open
MohammadHijjawi97 wants to merge 1 commit into
MohammadHijjawi97 wants to merge 1 commit into
Conversation
`maybe_is_html` compared the first four bytes against uppercase-only signatures, so a lowercase `<!doctype html>` (the common HTML5 form), an uppercase `<HTML>` tag, a UTF-8 BOM, or leading whitespace all went undetected. `Docs.aadd_file` then saved such files as `.txt` and indexed the raw markup instead of parsing it with `html2text`.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
maybe_is_htmlcompared the first 4 bytes against{b"<htm", b"<!DO", b"<xsl", b"<!X"}, so only an uppercase<!DOCTYPEor a lowercase<htmwas recognized. These inputs were all missed:<!doctype html>, the common HTML5 form;<HTML>tag;Docs.aadd_filewrote such files to a.txttemp file and indexed the raw markup instead of parsing it withhtml2text.This change reads a small prefix, drops a UTF-8 BOM and leading whitespace, and compares case-insensitively.
Added a parametrized offline test. It covers lowercase and uppercase doctypes and tags, leading whitespace and a BOM, plus PDF and plain-text negatives, and checks that the file is rewound. The four new positive cases fail on
mainand pass here.ruffandblackare clean.