Repository navigation
Overhaul span compression - #163857
Overhaul span compression#163857nnethercote wants to merge 1 commit into
Conversation
|
@bors try @rust-timer queue |
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Overhaul span compression
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
|
Finished benchmarking commit (d2adbd4): comparison URL. Overall result: ❌✅ regressions and improvements - please read:Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf. Next, please: If you can, justify the regressions found in this try perf run in writing along with @bors rollup=never rustc-perf Instruction countOur most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.
Max RSS (memory usage)Results (primary -1.5%, secondary -1.1%)A less reliable metric. May be of interest, but not used to determine the overall result above.
CyclesResults (primary -3.5%, secondary -3.5%)A less reliable metric. May be of interest, but not used to determine the overall result above.
Binary sizeThis perf run didn't have relevant results for this metric. Bootstrap: 489.693s -> 492.198s (0.51%) |
|
icounts results look good. Cycles results are even better, probably because max-rss/faults/cache-misses/branch-misses are all improved. |
|
LLM disclosure: I used an LLM for some of the ideas and analysis in this PR. I wrote the code and text myself. |
302cea2 to
1596682
Compare
1596682 to
7d10c65
Compare
Currently there are span representations if you have a context and no
parent, or a parent and no context. But if you have both a context *and*
a parent, the span is always interned. In incremental builds this is
something like 20-40% of all spans.
This commit uses a new compression scheme that gives small perf wins for
non-incremental and bigger perf wins for incremental. The following table
briefly summarizes the formats.
```
-----------------------------------------------------------------
old new
-----------------------------------------------------------------
InlineCtxt lo32,len15,ctxt16,parent0 lo32,len15,ctxt16,parent0
InlineParent lo32,len15,ctxt0,parent16 lo32,len14,ctxt0,parent15
InlinePair n/a lo24,len7,ctxt15,parent16
PartialInt. ctxt16 everything else
FullInt. everything else n/a
```
Things to note.
- The old scheme had a 32-bit field and two 16-bit fields. The new scheme has
one 64-bit field and uses `repr(packed(4))`. (Keeping the alignment at
4 is important to keep AST node sizes low.)
- The new scheme distinguishes the four formats with a single variable-length
prefix in the high bits. This is simpler than the old combination of tag bits
and special values.
- The new `SpanField` type is used to describe the various field widths
for each format, and provide operations on them.
- The new `InlineParent` is very slightly worse than the old one, with one less
bit for each of len and parent. This doesn't affect many spans.
- The new `InlinePair` format handles most cases where both a context and a
parent are present. It has to compromise some on `lo` and `len`, but still
gets a lot of them. This is the big performance win.
- The partially/fully interned cases are merged, because the Interned format
can now always hold a full context inline, and the new `Interned` is
equivalent to the old PartiallyInterned format. (This requires
limiting ctxts to 29 bits, but we'll hit other problems, such as OOM,
long before that limit is reached.) This also means we don't need to
store the ctxt field in the interner, and we use the new
`SpanDataNoCtxt` type for that.
7d10c65 to
e40a818
Compare
|
I made some small changes that shouldn't affect perf, let's check: @bors try @rust-timer queue |
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Overhaul span compression
This comment has been minimized.
This comment has been minimized.
|
Finished benchmarking commit (03aba7c): comparison URL. Overall result: ❌✅ regressions and improvements - please read:Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf. Next, please: If you can, justify the regressions found in this try perf run in writing along with @bors rollup=never rustc-perf Instruction countOur most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.
Max RSS (memory usage)Results (primary -1.3%, secondary -1.0%)A less reliable metric. May be of interest, but not used to determine the overall result above.
CyclesResults (primary -3.6%, secondary -0.6%)A less reliable metric. May be of interest, but not used to determine the overall result above.
Binary sizeResults (primary -0.1%, secondary -0.1%)A less reliable metric. May be of interest, but not used to determine the overall result above.
Bootstrap: 489.893s -> 492.261s (0.48%) |
|
New perf results are basically identical to the old ones. |
|
|
||
| // A type used to work around `Span` not being visible in this crate. It is the same layout as | ||
| // A type used to work around `Span` not being visible in this crate. It is the same size as | ||
| // `Span`. |
There was a problem hiding this comment.
| // `Span`. | |
| // `Span`, but not the same alignment. |
Currently there are span representations if you have a context and no parent, or a parent and no context. But if you have both a context and a parent, the span is always interned. In incremental builds this is something like 20-40% of all spans.
This commit uses a new compression scheme that gives small perf wins for non-incremental and bigger perf wins for incremental. The following table briefly summarizes the formats.
Things to note.
The old scheme had a 32-bit field and two 16-bit fields. The new scheme has one 64-bit field and uses
repr(packed(4)). (Keeping the alignment at 4 is important to keep AST node sizes low.)The new scheme distinguishes the four formats with a single variable-length prefix in the high bits. This is simpler than the old combination of tag bits and special values.
The new
SpanFieldtype is used to describe the various field widths for each format, and provide operations on them.The new
InlineParentis very slightly worse than the old one, with one less bit for each of len and parent. This doesn't affect many spans.The new
InlinePairformat handles most cases where both a context and a parent are present. It has to compromise some onloandlen, but still gets a lot of them. This is the big performance win.The partially/fully interned cases are merged, because the Interned format can now always hold a full context inline, and the new
Internedis equivalent to the old PartiallyInterned format. (This requires limiting ctxts to 29 bits, but we'll hit other problems, such as OOM, long before that limit is reached.) This also means we don't need to store the ctxt field in the interner, and we use the newSpanDataNoCtxttype for that.r? @petrochenkov