Repository navigation
Shrink MIR basic-block statement vectors to fit before caching them - #163915
Conversation
Every MIR basic block allocates a vector for its statements. These statement vectors grow via the usual heuristics, and never get re-shrunk to the correct size for their contents. This holds on to a fair bit of memory. Shrink them to fit right before we cache them. For an aws-sdk-ec2 check, this saves ~80 MiB of peak max-RSS (~1.45%), at the cost of 0.3% in instructions.
|
Some changes occurred to MIR optimizations cc @rust-lang/wg-mir-opt |
|
I tested aws-sdk-ec2 because it's not in the perf suite, and I tested a few other things directly, but let's see the whole suite: @bors try @rust-timer queue |
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Shrink MIR basic-block statement vectors to fit before caching them
This comment has been minimized.
This comment has been minimized.
|
Finished benchmarking commit (23ca711): comparison URL. Overall result: ❌✅ regressions and improvements - please read:Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf. Next, please: If you can, justify the regressions found in this try perf run in writing along with @bors rollup=never rustc-perf Instruction countOur most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.
Max RSS (memory usage)Results (primary -1.3%, secondary -2.0%)A less reliable metric. May be of interest, but not used to determine the overall result above.
CyclesResults (primary -2.8%, secondary 5.7%)A less reliable metric. May be of interest, but not used to determine the overall result above.
Binary sizeThis perf run didn't have relevant results for this metric. Bootstrap: 489.893s -> 491.828s (0.39%) |
|
The 0.1-0.2% performance regressions are expected, in exchange for the substantial max-RSS savings. |
|
@rustbot label: +perf-regression-triaged |
I believe this is a bit of a problem, because people often make the tradeoff in the other direction, so somebody might just do the reverse optimization in the future. MaxRSS is also super noisy so it's a bit difficult to judge that tradeoff very well. Nevertheless, I agree that we should prioritize memory a bit more. This has come up a few times recently (and on your |
| @@ -483,6 +483,10 @@ fn mir_promoted( | |||
| Some(MirPhase::Analysis(AnalysisPhase::Initial)), | |||
| ); | |||
|
|
|||
There was a problem hiding this comment.
Can you add a short comment explaining that the small time cost is worth the memory saving? r=me with that
|
@bors r=nnethercote |
This comment has been minimized.
This comment has been minimized.
Shrink MIR basic-block statement vectors to fit before caching them Every MIR basic block allocates a vector for its statements. These statement vectors grow via the usual heuristics, and never get re-shrunk to the correct size for their contents. This holds on to a fair bit of memory. Shrink them to fit right before we cache them. For an aws-sdk-ec2 check, this saves ~80 MiB of peak max-RSS (~1.45%), at the cost of 0.3% in instructions. r? nnethercote
|
@bors rollup=never rustc-perf |
|
The job Click to see the possible cause of the failure (guessed by this bot) |
|
💔 Test for 9dbebb6 failed: CI. Failed job:
|
|
@bors retry |
This comment has been minimized.
This comment has been minimized.
What is this?This is an experimental post-merge analysis report that shows differences in test outcomes between the merged PR and its parent PR.Comparing 76c9095 (parent) -> 69bccf0 (this PR) Test differencesShow 3 test diffs3 doctest diffs were found. These are ignored, as they are noisy. Test dashboardRun cargo run --manifest-path src/ci/citool/Cargo.toml -- \
test-dashboard 69bccf03c4733f57482f9691456efcd80f86a5f6 --output-dir test-dashboardAnd then open Job duration changes
How to interpret the job duration changes?Job durations can vary a lot, based on the actual runner instance |
|
Finished benchmarking commit (69bccf0): comparison URL. Overall result: ❌✅ regressions and improvements - please read:Our benchmarks found a performance regression caused by this PR. Next Steps:
@rustbot label: +perf-regression Instruction countOur most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.
Max RSS (memory usage)Results (primary -1.3%, secondary -3.4%)A less reliable metric. May be of interest, but not used to determine the overall result above.
CyclesResults (primary -2.8%)A less reliable metric. May be of interest, but not used to determine the overall result above.
Binary sizeThis perf run didn't have relevant results for this metric. Bootstrap: 490.84s -> 487.936s (-0.59%) |
View all comments
Every MIR basic block allocates a vector for its statements. These statement vectors grow via the usual heuristics, and never get re-shrunk to the correct size for their contents. This holds on to a fair bit of memory.
Shrink them to fit right before we cache them.
For an aws-sdk-ec2 check, this saves ~80 MiB of peak max-RSS (~1.45%), at the cost of 0.3% in instructions.
r? nnethercote