Skip to content

bugc: end each source range at its last token - #362

Closed
gnidan wants to merge 2 commits into
mainfrom
bugc-trim-ranges
Closed

gnidan wants to merge 2 commits into
mainfrom
bugc-trim-ranges

Conversation

@gnidan

@gnidan gnidan commented Oct 7, 2026

Copy link
Copy Markdown
Member

Most code ranges in bugc's output ended with whitespace, and sometimes a comment. In a small game contract, at both -O 0 and -O 2, that held for the left operand of a binary operator (block.number in block.number + …), an assignment target (motd ), postfix chains (msg.data[0:4] as bytes4 ), if and return statements (up to the next statement, across any comment between them), and the create and code blocks (up to the next block). A debugger that highlights a range showed the extra text.

Every token parser in the BUG grammar is a lexeme that also consumes the whitespace and comments after it, and located took a node's end from Parsimmon's mark, which is the index after them. A node ends with a token, so its end is wrong whenever that token is followed by whitespace or a comment.

Now lexeme records, for each token, the index where it ends, keyed by the index after the whitespace and comments it skips. located and the postfix chain look up the end there, so a node's range ends at its last token. The map is per parse and cleared at the start of parse. Node ids come from the ranges, so they change too, and they stay unique (the bugc suite checks that for postfix chains).

The new parser test parses every example program and checks that no node's range ends with whitespace or inside a comment.

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://ethdebug.github.io/format/pr-preview/pr-362/

Built to branch gh-pages at 2026-10-07 02:01 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

gnidan added a commit that referenced this pull request Oct 7, 2026
* bugc: compile `!` as a logical not

* bugc: link the changelog entry to #353

* bugc: end each source range at its last token

* bugc: link the changelog entry to #362

* bugc: add block.prevrandao

* bugc: link the changelog entries to #363

* bugc: hash several value-type words with keccak256

* bugc: link the changelog entries to #365

* bugc: slice a fixed-size bytes value by its bytes

* bugc: link the changelog entry to #357

* bugc: hash the data of dynamic bytes and strings

* bugc: link the changelog entry to #360

* bugc: give a caller's call JUMP the invoke's arguments

* bugc: list variables before a function's first statement

* bugc: store a memory string's bytes when assigning it to storage

A storage write took its value as a word, so a memory string or bytes
reference wrote its memory address into the slot. Code generation now
encodes the referenced bytes as Solidity does: inline with length * 2
up to 31 bytes, else length * 2 + 1 in the slot and the data from
keccak256(slot), with the last word's trailing bytes cleared. A slice
now carries the bytes type, so its value is recognized too.

* bugc: link the changelog entry to #355

* bugc: say why the bytes test slices without naming the copy width

* bugc: copy all of a slice's bytes

irgen built a slice's data as a read and a write, and a read or write
moves at most one word, so only the first 32 bytes were copied. Add a
copy IR instruction (CALLDATACOPY from calldata, MCOPY from memory) and
emit it for the slice's data.

* bugc: link the changelog entries to #359

* bugc: copy a struct read from storage into memory

* bugc: link the changelog entries to #361

* bugc: lay storage arrays out as Solidity does

* bugc: add push to dynamic arrays in storage
@gnidan

gnidan commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Included in #368.

@gnidan gnidan closed this Oct 7, 2026
@gnidan
gnidan deleted the bugc-trim-ranges branch October 7, 2026 03:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant