Skip to content

Commit dd6b11a

Browse files
tompngclaude
andauthored
Remove the never-matching BOM rule from the Markdown grammar (#1820)
kpeg does not interpret `\u` escapes in string literals, so `BOM = "\uFEFF"` generated `match_string("uFEFF")` and the rule has never consumed a BOM since it was introduced. Files are read through RDoc::Encoding, which already strips the BOM, and no other markup parser handles it, so drop the rule instead of fixing it. ```ruby # :markup: markdown # uFEFF is not a bom because backslash is missing class A end ``` ↓ ```html <!-- before --> <section class="description"> <p>is not a bom because backslash is missing</p> <!-- after --> <section class="description"> <p>uFEFF is not a bom because backslash is missing</p> ``` No test added: the only behavioral change is that a leading string "uFEFF" is no longer eaten. I don't think we want a test to ensure that "uFEFF" remains. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
1 parent f7e570c commit dd6b11a

3 files changed

Lines changed: 4 additions & 23 deletions

File tree

‎lib/rdoc/markdown.kpeg‎

Lines changed: 1 addition & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -569,7 +569,7 @@
569569

570570
root = Doc
571571

572-
Doc = BOM? Block*:a { RDoc::Markup::Document.new(*a.compact) }
572+
Doc = Block*:a { RDoc::Markup::Document.new(*a.compact) }
573573

574574
Block = @BlankLine*
575575
( BlockQuote
@@ -1180,7 +1180,6 @@ Digit = [0-9]
11801180

11811181
Alphanumeric = /\p{Word}/
11821182
AlphanumericAscii = /[A-Za-z0-9]/
1183-
BOM = "\uFEFF"
11841183
Newline = /\n|\r\n?|\p{Zl}|\p{Zp}/
11851184
Spacechar = /\t|\p{Zs}/
11861185

‎lib/rdoc/markdown.rb‎

Lines changed: 2 additions & 20 deletions
Original file line numberDiff line numberDiff line change
@@ -962,21 +962,11 @@ def _root
962962
return _tmp
963963
end
964964

965-
# Doc = BOM? Block*:a { RDoc::Markup::Document.new(*a.compact) }
965+
# Doc = Block*:a { RDoc::Markup::Document.new(*a.compact) }
966966
def _Doc
967967

968968
_save = self.pos
969969
while true # sequence
970-
_save1 = self.pos
971-
_tmp = apply(:_BOM)
972-
unless _tmp
973-
_tmp = true
974-
self.pos = _save1
975-
end
976-
unless _tmp
977-
self.pos = _save
978-
break
979-
end
980970
_ary = []
981971
while true
982972
_tmp = apply(:_Block)
@@ -14777,13 +14767,6 @@ def _AlphanumericAscii
1477714767
return _tmp
1477814768
end
1477914769

14780-
# BOM = "uFEFF"
14781-
def _BOM
14782-
_tmp = match_string("uFEFF")
14783-
set_failed_rule :_BOM unless _tmp
14784-
return _tmp
14785-
end
14786-
1478714770
# Newline = /\n|\r\n?|\p{Zl}|\p{Zp}/
1478814771
def _Newline
1478914772
_tmp = scan(/\G(?-mix:\n|\r\n?|\p{Zl}|\p{Zp})/)
@@ -16624,7 +16607,7 @@ def _DefinitionListDefinition
1662416607

1662516608
Rules = {}
1662616609
Rules[:_root] = rule_info("root", "Doc")
16627-
Rules[:_Doc] = rule_info("Doc", "BOM? Block*:a { RDoc::Markup::Document.new(*a.compact) }")
16610+
Rules[:_Doc] = rule_info("Doc", "Block*:a { RDoc::Markup::Document.new(*a.compact) }")
1662816611
Rules[:_Block] = rule_info("Block", "@BlankLine* (BlockQuote | Verbatim | CodeFence | Table | Note | Reference | HorizontalRule | Heading | OrderedList | BulletList | DefinitionList | HtmlBlock | StyleBlock | Para | Plain)")
1662916612
Rules[:_Para] = rule_info("Para", "@NonindentSpace Inlines:a @BlankLine+ { paragraph a }")
1663016613
Rules[:_Plain] = rule_info("Plain", "Inlines:a { paragraph a }")
@@ -16838,7 +16821,6 @@ def _DefinitionListDefinition
1683816821
Rules[:_Digit] = rule_info("Digit", "[0-9]")
1683916822
Rules[:_Alphanumeric] = rule_info("Alphanumeric", "/\\p{Word}/")
1684016823
Rules[:_AlphanumericAscii] = rule_info("AlphanumericAscii", "/[A-Za-z0-9]/")
16841-
Rules[:_BOM] = rule_info("BOM", "\"uFEFF\"")
1684216824
Rules[:_Newline] = rule_info("Newline", "/\\n|\\r\\n?|\\p{Zl}|\\p{Zp}/")
1684316825
Rules[:_Spacechar] = rule_info("Spacechar", "/\\t|\\p{Zs}/")
1684416826
Rules[:_HexEntity] = rule_info("HexEntity", "/&\#x/i < /[0-9a-fA-F]+/ > \";\" { rdoc_escape([text.to_i(16)].pack('U')) }")

‎lib/rdoc/markdown/byte_runtime.rb‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -26,7 +26,7 @@ class Markdown
2626
# +current_column+, ...) are left as-is and would misreport locations when
2727
# given byte offsets. They are unreachable: markdown is deliberately
2828
# designed to parse any input somehow rather than fail (the root rule
29-
# `Doc = BOM? Block*` cannot fail), so a parse failure means a bug in the
29+
# `Doc = Block*` cannot fail), so a parse failure means a bug in the
3030
# grammar itself, and nothing in RDoc invokes +raise_error+ or
3131
# +show_error+. Make these helpers byte-aware before using them for
3232
# anything.

0 commit comments

Comments
 (0)