Skip to content

Remember a POLY parameter's hand-on answers per query in fwd_param_appends - #7390

Merged
matz merged 1 commit into
matz:masterfrom
makenowjust:fwd-poly-handed-on-memo
Oct 5, 2026
Merged

matz merged 1 commit into
matz:masterfrom
makenowjust:fwd-poly-handed-on-memo

Conversation

@makenowjust

@makenowjust makenowjust commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #7389.

fwd_poly_param_handed_on and fwd_param_appends follow a POLY parameter through the calls and supers it is handed to. A plain parameter's answer was not remembered, so the same (method, parameter, depth) was asked once per call path reaching it.

Change

src/analyze.c only. Each ask from outside (fwd_param_appends or fwd_poly_param_handed_on at depth 0, with no rest forwarder being asked) starts a fresh memo, an open-addressed table stamped with a generation. Within the ask, fwd_param_appends remembers, for each (method, parameter, depth), its answer and the cut-short flags (g_fwd_taint) its walk left, and on a repeat returns the answer and sets the same flags.

  • The depth is part of the key: the depth bound cuts a deeper ask shorter, so the same (method, parameter) can answer differently at different depths.
  • The flags are stored and replayed, so a remembered answer leaves g_fwd_taint exactly as asking again would; fwd_rest_bits' OPEN flag and the emitters' -1 answers depend on them.
  • Nothing is remembered while a rest forwarder is being asked (g_fwd_rest_depth > 0): there an answer leans on that forwarder's partial bits.
  • The memo lives for one ask only: the passes between asks change the types the answers read.

Results

Calls of fwd_param_appends for the issue's generator, spinel -c (raise mode; measured with the memo for an_subtree_hands_to_appender from a separate pull request applied -- the first two counts are the same without it):

program before after
gen.rb 8 16 16 (2,339 lines) 354,048 46,848
gen.rb 8 24 24 (4,851 lines) 1,745,856 156,096

The 141k-line machine-generated program from the issue (four other slow passes disabled for the measurement): 283 s -> 201 s in raise mode, 295 s -> 228 s in promote mode, whole compile. Timings are from a heavily loaded machine; the call counts are exact.

The generated C does not change

  • Test corpus: every test/*.rb (5,746 programs) compiled with spinel -c in both --int-overflow modes, on master (e5e8f79) and on this change, built in turn in the same checkout and fed the same relative paths: the generated C, stderr and exit status are identical for all 11,492 compilations except the 8 (4 programs, 2 modes each) that embed RUBY_DESCRIPTION, whose only difference is the revision string.
  • After the rebase onto 92510d6: 164 programs related to forwarding and appends, both modes, compiled the same way: no difference.
  • The 141k-line program and seven other machine-generated programs of 4k-56k lines: identical before and after, both modes.

Gate

make gate was run on 92510d6 with this change merged together with other fixes from the same investigation (#7375, #7377, #7378, #7384, and others):

scale-test: instance_eval forwarding work at 2x the wrappers is 1.71x (limit 2.50)
scale-test: work at 4x the program is 4.73x (linear 4.00, limit 5.20)
scale-test: work at 4x the program, compiled to C, is 6.08x (limit 6.90)
scale-test: call-shape work at 4x the units, compiled to C, is 4.22x (linear 4.00, limit 4.50)
Tests: 5928 pass, 0 fail, 0 error
gate: stamp for tree c190c8822c9a on master 92510d6c10a3; git commit --amend --no-edit adds the Gate: trailer
gate: ALL GREEN

Promote mode

SPINEL_INT_OVERFLOW=promote make test on the same combined branch: Tests: 5909 pass, 15 fail, 13 error. All 28 failing tests are among the 29 that master fails in promote mode at e5e8f79; none is new. The 29th (raise_rejects_invalid_arguments) no longer fails, which appears to come from master moving from e5e8f79 to 92510d6.

Summary by CodeRabbit

  • Performance
    • Repeated forwarding analysis results are now reused within a query, while preserving taint information.
  • Reliability
    • Analysis still respects the existing mutation check and depth limit; exceeding the limit returns false and sets taint.

…er method and depth

fwd_poly_param_handed_on and fwd_param_appends follow a POLY parameter
through the calls and supers it is handed to, four levels deep, to see
whether one of them appends to it. Only a rest's answer was remembered
(fwd_rest_bits); a plain parameter was asked afresh along every path, so
a value handed to many methods, each handing it on to many, cost the
number of paths up to the bound -- fan-out to the fifth power, with the
callee's whole call list walked at each leaf. A synthetic program of
eight layers of 32 methods, each handing its two POLY parameters to 32
methods of the next layer, spent 108 s here; a 141k-line
machine-generated program compiles 70-80 s faster with this change.

Each ask from outside now remembers, for that ask only, what each
(method, parameter, depth) answered and which bound-cut flags it left,
and replays both on a repeat. The depth stays in the key because the bound
cuts a deeper ask shorter, and nothing is remembered while a rest
forwarder is being asked, where an answer leans on that forwarder's
partial bits. The passes between asks change the types the answers read,
so each ask starts empty. The generated C is unchanged.
@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: a7ce89d4-f315-42bf-97e6-0252fe1931b6
📥 Commits

Reviewing files that changed from the base of the PR and between ab9b925 and 20e8283.

📒 Files selected for processing (1)
  • src/analyze.c

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

Forwarding analysis now memoizes POLY-parameter results by method, parameter, and recursion depth for each query. Cached results include taint. The analyzer bypasses memoization during rest-forwarder evaluation and retains the existing mutation check and depth limit.

Changes

Forwarding analysis memoization

Layer / File(s) Summary
Memoized forwarding checks
src/analyze.c
Forwarding queries start a memo generation when no rest forwarder is being evaluated. fwd_param_appends looks up results, checks uncached cases, and stores the result and taint. The table grows above a one-half load factor and exits if allocation fails.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Refactor · Severity of issue fixed: Medium

Sequence Diagram(s)

sequenceDiagram
  participant fwd_poly_param_handed_on
  participant fwd_param_appends
  participant ForwardingMemo
  participant an_param_mutated_in_place
  fwd_poly_param_handed_on->>ForwardingMemo: Start a query at depth 0
  fwd_poly_param_handed_on->>fwd_param_appends: Check the forwarded parameter
  fwd_param_appends->>ForwardingMemo: Look up method, parameter, and depth
  alt Cached result
    ForwardingMemo-->>fwd_param_appends: Return result and taint
  else Cache miss
    fwd_param_appends->>an_param_mutated_in_place: Check in-place mutation
    fwd_param_appends->>ForwardingMemo: Store result and taint
  end
Loading

Suggested reviewers: matz

Merge Risk: ⚪ Minimal · up to 20e82

This change speeds up forwarding analysis in the compiler by caching repeated results within a single query. No actionable merge-blocking risk was found. The author reports identical generated output across the corpus.

Architecture Summary

Architecture risk: 🔵 Low · up to 20e82

The change affects 1 system.

Changed systems: src

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — src (service) was modified; 1 changed file maps to changed impact.

Before / after behavior

  • observed — Modified behavior in src/analyze.c: Adds per-query forwarding memo state and lookup/insertion helpers keyed by method, parameter, and depth. Queries advance a generation and reset the entry count; generation wrap clears stamps. The table grows above a one-half load factor, rehashes current-generation entries, and exits on allocation failure. The forwarding query begins a new memo generation at depth 0 when no rest forwarder is being evaluated.
  • observed — Modified behavior in src/analyze.c: fwd_param_appends now checks the memo outside rest-forwarder evaluation, restoring cached taint on a hit. On a miss it isolates taint while checking in-place mutation, the depth-over-4 cutoff, or recursive forwarding, then caches the result and taint before restoring outer taint. The mutation and depth-limit outcomes remain unchanged.
  • observed — Modified behavior in src/analyze.c: The g_fwd_rest_depth declaration was removed from this location; it is now declared with the forwarding memo state.
  • observed — Modified behavior in src/analyze.c: fwd_poly_param_handed_on now starts a memo query at depth 0 when no rest forwarder is being evaluated.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: caching POLY-parameter answers per query in fwd_param_appends.
Linked Issues check ✅ Passed Issue #7389 asks to avoid repeating plain POLY-parameter forwarding walks across call paths. The reviewed change in src/analyze.c memoizes each (method, parameter, depth) answer for one query and …
Out of Scope Changes check ✅ Passed The change is limited to src/analyze.c and adds the per-query forwarding memo requested by issue #7389. The table management, generation reset, and taint replay support that objective. No unrelated …
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@matz
matz merged commit 458aa01 into matz:master Oct 5, 2026
4 checks passed
kieranklaassen pushed a commit to kieranklaassen/spinel that referenced this pull request Oct 5, 2026
… compiler

`h[:a1] = h[:a0]; h[:a2] = h[:a1]; ...` with a String of h changed in
place cost the compiler n! walks for n such stores: on f8abc3e eight
take 12 seconds, and nine do not finish in 90. The
same in an Array's slots, in a Hash an instance variable holds, in one
a parameter holds, and in two Hashes that store each other's elements.

The walk that demands a container's Strings as handles follows a boxed
stored value to the element read it is, and strbuf_demand_elem_arg
starts the walk of that read's container. Since matz#7369 a read whose walk
is under way answers 0, which ended the recursion. A second read of the
same container was still another read, so the stores were walked once
more for it with one read fewer left to meet, and again inside that.

The walk is of the stores into the container a local or an instance
variable holds, whichever element is read. A read of that container met
inside the walk now answers 0 as the same read does: the walk on the
stack reaches every store the inner one would. Each of the n stores
walks the others once.

An element of another element (`r[0][:b] = r[0][:a]`) is keyed by its
read as before; that walk did not grow this way.

The nearest changes upstream are matz#7369, which this completes, and the
walk-speed changes merged since (matz#7365's memo and the per-method one
beside it, matz#7384, matz#7386, matz#7388, matz#7390, matz#7392, matz#7396 to matz#7402). Those
make one walk cheaper or remember a walk that changed nothing; none
touches this guard, and here it is the number of walks that grows. A
walk that reaches a container moves the memo's generation whether it
changed anything or not, so the memo never holds one of these.

The generated C is unchanged for every other test.

Co-Authored-By: Claude Code <noreply@anthropic.com>
kieranklaassen pushed a commit to kieranklaassen/spinel that referenced this pull request Oct 5, 2026
… compiler

`h[:a1] = h[:a0]; h[:a2] = h[:a1]; ...` with a String of h changed in
place cost the compiler n! walks for n such stores: on c6bbdfb eight
take 13 seconds, and nine do not finish in 90. The
same in an Array's slots, in a Hash an instance variable holds, in one
a parameter holds, and in two Hashes that store each other's elements.

The walk that demands a container's Strings as handles follows a boxed
stored value to the element read it is, and strbuf_demand_elem_arg
starts the walk of that read's container. Since matz#7369 a read whose walk
is under way answers 0, which ended the recursion. A second read of the
same container was still another read, so the stores were walked once
more for it with one read fewer left to meet, and again inside that.

The walk is of the stores into the container a local or an instance
variable holds, whichever element is read. A read of that container met
inside the walk now answers 0 as the same read does: the walk on the
stack reaches every store the inner one would. Each of the n stores
walks the others once.

An element of another element (`r[0][:b] = r[0][:a]`) is keyed by its
read as before; that walk did not grow this way.

The nearest changes upstream are matz#7369, which this completes, and the
walk-speed changes merged since (matz#7365's memo and the per-method one
beside it, matz#7384, matz#7386, matz#7388, matz#7390, matz#7392, matz#7396 to matz#7402). Those
make one walk cheaper or remember a walk that changed nothing; none
touches this guard, and here it is the number of walks that grows. A
walk that reaches a container moves the memo's generation whether it
changed anything or not, so the memo never holds one of these.

The generated C is unchanged for every other test.

Co-Authored-By: Claude Code <noreply@anthropic.com>
kieranklaassen pushed a commit to kieranklaassen/spinel that referenced this pull request Oct 5, 2026
… compiler

`h[:a1] = h[:a0]; h[:a2] = h[:a1]; ...` with a String of h changed in
place cost the compiler n! walks for n such stores: on c6bbdfb eight
take 13 seconds, and nine do not finish in 90. The
same in an Array's slots, in a Hash an instance variable holds, in one
a parameter holds, and in two Hashes that store each other's elements.

The walk that demands a container's Strings as handles follows a boxed
stored value to the element read it is, and strbuf_demand_elem_arg
starts the walk of that read's container. Since matz#7369 a read whose walk
is under way answers 0, which ended the recursion. A second read of the
same container was still another read, so the stores were walked once
more for it with one read fewer left to meet, and again inside that.

The walk is of the stores into the container a local or an instance
variable holds, whichever element is read. A read of that container met
inside the walk now answers 0 as the same read does: the walk on the
stack reaches every store the inner one would. Each of the n stores
walks the others once.

An element of another element (`r[0][:b] = r[0][:a]`) is keyed by its
read as before; that walk did not grow this way.

The nearest changes upstream are matz#7369, which this completes, and the
walk-speed changes merged since (matz#7365's memo and the per-method one
beside it, matz#7384, matz#7386, matz#7388, matz#7390, matz#7392, matz#7396 to matz#7402). Those
make one walk cheaper or remember a walk that changed nothing; none
touches this guard, and here it is the number of walks that grows. A
walk that reaches a container moves the memo's generation whether it
changed anything or not, so the memo never holds one of these.

The generated C is unchanged for every other test.

Co-Authored-By: Claude Code <noreply@anthropic.com>
kieranklaassen pushed a commit to kieranklaassen/spinel that referenced this pull request Oct 5, 2026
… compiler

`h[:a1] = h[:a0]; h[:a2] = h[:a1]; ...` with a String of h changed in
place cost the compiler n! walks for n such stores: on c6bbdfb eight
take 13 seconds, and nine do not finish in 90. The
same in an Array's slots, in a Hash an instance variable holds, in one
a parameter holds, and in two Hashes that store each other's elements.

The walk that demands a container's Strings as handles follows a boxed
stored value to the element read it is, and strbuf_demand_elem_arg
starts the walk of that read's container. Since matz#7369 a read whose walk
is under way answers 0, which ended the recursion. A second read of the
same container was still another read, so the stores were walked once
more for it with one read fewer left to meet, and again inside that.

The walk is of the stores into the container a local or an instance
variable holds, whichever element is read. A read of that container met
inside the walk now answers 0 as the same read does: the walk on the
stack reaches every store the inner one would. Each of the n stores
walks the others once.

An element of another element (`r[0][:b] = r[0][:a]`) is keyed by its
read as before; that walk did not grow this way.

The nearest changes upstream are matz#7369, which this completes, and the
walk-speed changes merged since (matz#7365's memo and the per-method one
beside it, matz#7384, matz#7386, matz#7388, matz#7390, matz#7392, matz#7396 to matz#7402). Those
make one walk cheaper or remember a walk that changed nothing; none
touches this guard, and here it is the number of walks that grows. A
walk that reaches a container moves the memo's generation whether it
changed anything or not, so the memo never holds one of these.

The generated C is unchanged for every other test.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Asking whether a POLY parameter is handed on to an appending method costs one walk per call path

2 participants