Skip to content

Overhaul span compression - #163857

Open
nnethercote wants to merge 1 commit into
rust-lang:mainfrom
nnethercote:overhaul-span-compression
Open

nnethercote wants to merge 1 commit into
rust-lang:mainfrom
nnethercote:overhaul-span-compression

Conversation

@nnethercote

@nnethercote nnethercote commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Currently there are span representations if you have a context and no parent, or a parent and no context. But if you have both a context and a parent, the span is always interned. In incremental builds this is something like 20-40% of all spans.

This commit uses a new compression scheme that gives small perf wins for non-incremental and bigger perf wins for incremental. The following table briefly summarizes the formats.

-----------------------------------------------------------------
              old                        new
-----------------------------------------------------------------
InlineCtxt    lo32,len15,ctxt16,parent0  lo32,len15,ctxt16,parent0
InlineParent  lo32,len15,ctxt0,parent16  lo32,len14,ctxt0,parent15
InlinePair    n/a                        lo24,len7,ctxt15,parent16
PartialInt.   ctxt16                     everything else
FullInt.      everything else            n/a

Things to note.

  • The old scheme had a 32-bit field and two 16-bit fields. The new scheme has one 64-bit field and uses repr(packed(4)). (Keeping the alignment at 4 is important to keep AST node sizes low.)

  • The new scheme distinguishes the four formats with a single variable-length prefix in the high bits. This is simpler than the old combination of tag bits and special values.

  • The new SpanField type is used to describe the various field widths for each format, and provide operations on them.

  • The new InlineParent is very slightly worse than the old one, with one less bit for each of len and parent. This doesn't affect many spans.

  • The new InlinePair format handles most cases where both a context and a parent are present. It has to compromise some on lo and len, but still gets a lot of them. This is the big performance win.

  • The partially/fully interned cases are merged, because the Interned format can now always hold a full context inline, and the new Interned is equivalent to the old PartiallyInterned format. (This requires limiting ctxts to 29 bits, but we'll hit other problems, such as OOM, long before that limit is reached.) This also means we don't need to store the ctxt field in the interner, and we use the new SpanDataNoCtxt type for that.

r? @petrochenkov

@rustbot rustbot added S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue. labels Oct 6, 2026
@nnethercote

Copy link
Copy Markdown
Contributor Author

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Oct 6, 2026
@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Oct 6, 2026
@rust-bors

rust-bors Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

☀️ Try build successful (CI)
Build commit: d2adbd4 (d2adbd4acd6d3b565c7a003a9fc9439c67d42cf6)
Base parent: b57eb9a (b57eb9a5fd94f30697f51fc88fbe898e67c7aae6)

@rust-timer

This comment has been minimized.

@rust-log-analyzer

This comment has been minimized.

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (d2adbd4): comparison URL.

Overall result: ❌✅ regressions and improvements - please read:

Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf.

Next, please: If you can, justify the regressions found in this try perf run in writing along with @rustbot label: +perf-regression-triaged. If not, fix the regressions and do another perf run. Neutral or positive results will clear the label automatically.

@bors rollup=never rustc-perf
@rustbot label: -S-waiting-on-perf +perf-regression

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
0.9% [0.5%, 1.5%] 16
Improvements ✅
(primary)
-1.0% [-4.0%, -0.1%] 261
Improvements ✅
(secondary)
-0.8% [-5.6%, -0.1%] 244
All ❌✅ (primary) -1.0% [-4.0%, -0.1%] 261

Max RSS (memory usage)

Results (primary -1.5%, secondary -1.1%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
2.5% [2.3%, 2.9%] 3
Improvements ✅
(primary)
-1.5% [-3.3%, -0.5%] 32
Improvements ✅
(secondary)
-2.2% [-4.1%, -1.0%] 10
All ❌✅ (primary) -1.5% [-3.3%, -0.5%] 32

Cycles

Results (primary -3.5%, secondary -3.5%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
4.1% [3.9%, 4.4%] 2
Improvements ✅
(primary)
-3.5% [-9.8%, -1.6%] 71
Improvements ✅
(secondary)
-4.1% [-9.5%, -1.7%] 23
All ❌✅ (primary) -3.5% [-9.8%, -1.6%] 71

Binary size

This perf run didn't have relevant results for this metric.

Bootstrap: 489.693s -> 492.198s (0.51%)
Artifact size: 409.33 MiB -> 408.67 MiB (-0.16%)

@rustbot rustbot added perf-regression Performance regression. and removed S-waiting-on-perf Status: Waiting on a perf run to be completed. labels Oct 6, 2026
Comment thread compiler/rustc_data_structures/src/stable_hash.rs Outdated
@nnethercote

Copy link
Copy Markdown
Contributor Author

icounts results look good. Cycles results are even better, probably because max-rss/faults/cache-misses/branch-misses are all improved.

@nnethercote

Copy link
Copy Markdown
Contributor Author

LLM disclosure: I used an LLM for some of the ideas and analysis in this PR. I wrote the code and text myself.

@nnethercote
nnethercote force-pushed the overhaul-span-compression branch from 302cea2 to 1596682 Compare October 6, 2026 08:43
@nnethercote
nnethercote marked this pull request as ready for review October 6, 2026 08:45
@rustbot rustbot added S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. and removed S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. labels Oct 6, 2026
@nnethercote
nnethercote force-pushed the overhaul-span-compression branch from 1596682 to 7d10c65 Compare October 6, 2026 22:05
Currently there are span representations if you have a context and no
parent, or a parent and no context. But if you have both a context *and*
a parent, the span is always interned. In incremental builds this is
something like 20-40% of all spans.

This commit uses a new compression scheme that gives small perf wins for
non-incremental and bigger perf wins for incremental. The following table
briefly summarizes the formats.

```
-----------------------------------------------------------------
              old                        new
-----------------------------------------------------------------
InlineCtxt    lo32,len15,ctxt16,parent0  lo32,len15,ctxt16,parent0
InlineParent  lo32,len15,ctxt0,parent16  lo32,len14,ctxt0,parent15
InlinePair    n/a                        lo24,len7,ctxt15,parent16
PartialInt.   ctxt16                     everything else
FullInt.      everything else            n/a
```

Things to note.

- The old scheme had a 32-bit field and two 16-bit fields. The new scheme has
  one 64-bit field and uses `repr(packed(4))`. (Keeping the alignment at
  4 is important to keep AST node sizes low.)

- The new scheme distinguishes the four formats with a single variable-length
  prefix in the high bits. This is simpler than the old combination of tag bits
  and special values.

- The new `SpanField` type is used to describe the various field widths
  for each format, and provide operations on them.

- The new `InlineParent` is very slightly worse than the old one, with one less
  bit for each of len and parent. This doesn't affect many spans.

- The new `InlinePair` format handles most cases where both a context and a
  parent are present. It has to compromise some on `lo` and `len`, but still
  gets a lot of them. This is the big performance win.

- The partially/fully interned cases are merged, because the Interned format
  can now always hold a full context inline, and the new `Interned` is
  equivalent to the old PartiallyInterned format. (This requires
  limiting ctxts to 29 bits, but we'll hit other problems, such as OOM,
  long before that limit is reached.) This also means we don't need to
  store the ctxt field in the interner, and we use the new
  `SpanDataNoCtxt` type for that.
@nnethercote
nnethercote force-pushed the overhaul-span-compression branch from 7d10c65 to e40a818 Compare October 7, 2026 03:07
@nnethercote

Copy link
Copy Markdown
Contributor Author

I made some small changes that shouldn't affect perf, let's check:

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Oct 7, 2026
@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Oct 7, 2026
@rust-bors

rust-bors Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

☀️ Try build successful (CI)
Build commit: 03aba7c (03aba7c9e40d02c0476b996cdc81cbe1bd24e247)
Base parent: 8d1a764 (8d1a76430406c877b35d0b627e7f796dcf0dfeca)

@rust-timer

This comment has been minimized.

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (03aba7c): comparison URL.

Overall result: ❌✅ regressions and improvements - please read:

Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf.

Next, please: If you can, justify the regressions found in this try perf run in writing along with @rustbot label: +perf-regression-triaged. If not, fix the regressions and do another perf run. Neutral or positive results will clear the label automatically.

@bors rollup=never rustc-perf
@rustbot label: -S-waiting-on-perf +perf-regression

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
0.9% [0.5%, 1.6%] 16
Improvements ✅
(primary)
-1.0% [-4.0%, -0.1%] 262
Improvements ✅
(secondary)
-0.8% [-5.5%, -0.1%] 250
All ❌✅ (primary) -1.0% [-4.0%, -0.1%] 262

Max RSS (memory usage)

Results (primary -1.3%, secondary -1.0%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
2.4% [2.4%, 2.4%] 1
Improvements ✅
(primary)
-1.3% [-2.5%, -0.5%] 34
Improvements ✅
(secondary)
-1.7% [-2.2%, -1.2%] 5
All ❌✅ (primary) -1.3% [-2.5%, -0.5%] 34

Cycles

Results (primary -3.6%, secondary -0.6%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
2.4% [2.0%, 2.9%] 2
Regressions ❌
(secondary)
5.3% [2.8%, 8.0%] 8
Improvements ✅
(primary)
-3.8% [-10.0%, -1.6%] 58
Improvements ✅
(secondary)
-3.1% [-6.7%, -1.9%] 19
All ❌✅ (primary) -3.6% [-10.0%, 2.9%] 60

Binary size

Results (primary -0.1%, secondary -0.1%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
- - 0
Improvements ✅
(primary)
-0.1% [-0.2%, -0.0%] 59
Improvements ✅
(secondary)
-0.1% [-0.2%, -0.0%] 53
All ❌✅ (primary) -0.1% [-0.2%, -0.0%] 59

Bootstrap: 489.893s -> 492.261s (0.48%)
Artifact size: 408.60 MiB -> 408.66 MiB (0.02%)

@rustbot rustbot removed the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Oct 7, 2026
@nnethercote

Copy link
Copy Markdown
Contributor Author

New perf results are basically identical to the old ones.


// A type used to work around `Span` not being visible in this crate. It is the same layout as
// A type used to work around `Span` not being visible in this crate. It is the same size as
// `Span`.

@joshtriplett joshtriplett Oct 8, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
// `Span`.
// `Span`, but not the same alignment.

View changes since the review

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

perf-regression Performance regression. S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants