feat!: deepen on the bound by default - #78
Draft
evanlinjin wants to merge 30 commits into
Draft
Conversation
…y_count Fixes CoinSelector::input_weight undercounting candidates that group multiple legacy inputs in a segwit transaction (where each legacy input serializes a 1 WU empty witness). Tracking segwit and legacy input counts separately also allows a single Candidate to mix legacy and segwit inputs.
…legacy Replaces the boolean is_segwit parameter in Candidate::new with explicit new_segwit and new_legacy constructors. Clarifies in doc comments that satisfaction_weight is the additional weight required beyond TXIN_BASE_WEIGHT (which already accounts for a 1-byte scriptSigLen).
…call
A selector was built for one target and evaluated against it throughout,
but every method took the target as a parameter, so nothing stopped
`cs.excess(target_a, drain)` being followed by `cs.is_funded(target_b)`.
The correctness arguments in the metrics are all stated at a fixed target
-- `LowestFee::bound`'s proof that a changeless superset always costs
more, `Changeless::change_unavoidable`'s assumption that the drain
decision is monotone in the excess -- and were held together by
convention rather than by types.
`CoinSelector::new` now takes the target and owns it. Twenty signatures
*lose* a parameter rather than gaining one: fifteen public methods
(`excess`, `implied_fee`, `is_funded`, `drain`, `select_until_target_met`,
the four `*_excess`, ...), plus `bnb_solutions` and `run_bnb`, plus all
three `BnbMetric` methods.
The crate had already reached this conclusion one layer down: `BnbIter`
stored the target as a field, took it once in `BnbIter::new`, and then
re-passed it into `metric.score` and `metric.bound` at every node. That
field and the re-threading are both gone.
This is a breaking change, and it reaches `BnbMetric`, so metrics
implemented outside this crate need their signatures updated:
fn score(&mut self, cs: &CoinSelector<'_>) -> Option<Ordf32>;
fn bound(&mut self, cs: &CoinSelector<'_>) -> Option<Ordf32>;
fn drain(&mut self, cs: &CoinSelector<'_>) -> Drain;
`CoinSelector::target()` exposes the target for metrics that need to read
it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Move the fixed target, candidates, and optional ancestor graph into one immutable problem object. CoinSelector now borrows that object, keeping all calculations tied to the same inputs and allowing ancestry metadata to remain separate from Candidate. Provide new_no_ancestors for prebuilt candidates and new for constructing candidates from input groups and their unconfirmed transaction graph.
Selecting an unconfirmed coin means paying to bump its ancestors. The feerate obligation includes the shortfall of the union of ancestors the selected candidates drag in (each charged once; weight and fee netted; saturates at 0). Score is still the child fee — the bump is already inside it. With ancestors, LowestFee falls back to a loose but admissible fee floor; tightening is a follow-up. BnB only batch-bans look-alikes with the same drags_in; Changeless disables its prune when ancestors are present.
Precompute ancestors reachable through exactly one candidate as summed private packages. Keep bitset de-duplication only for ancestors shared by multiple candidates, preserving exact union accounting while reducing the common-path work in every fee calculation. Add Criterion coverage for private and shared ancestry at 20, 50, and 100 candidates, plus exhaustive regressions for the optimized representation.
For funded nodes, subtract the ancestor surplus still reachable by a descendant. For unfunded nodes, derive a minimum added child weight from independent fractional relaxations of the target-rate, absolute-fee, and RBF constraints, then evaluate the fee floor at that weight. Candidate ancestry is deliberately represented only by the global bump lower bound: package surplus can absorb a later private deficit, so a per-candidate ancestor cost is not admissible. Keep infeasibility prunes off because ancestor funding is non-monotone. Add regressions for package subsidy, absolute/RBF double counting, and large-float cancellation, plus the existing exhaustive proptests.
Maintain aggregate selection state per branch and expose it through SelectionView so metric evaluation avoids repeatedly walking selected candidates. Track each branch's candidate cursor to skip repeated scans, and extend benchmarks across wallet- and exchange-scale pools.
Keep SelectionView's hypothetical updates set-like and synchronize ancestor reachability when branches exclude candidates. Remove unsound funding and changeless assumptions exposed by non-monotone ancestor debt, and preserve conservative fee rounding in the bound. Add regressions for public view updates, exclusion transitions, weight caps, mixed serialization overhead, and floating-point edge cases.
Separate deterministic solution-finding cases from larger pools expected to exhaust the fixed round cap. Assert each fixture's expected search outcome before measuring it so benchmark comparisons cannot silently time different paths.
Store private ancestor totals directly and allocate shared reference tracking only when the problem actually has shared ancestry. Preserve an explicit precision allowance for large floating-point ancestor fees so the smaller cache does not tighten the admissible bound.
Replace generic metric composition with a changeless metric that reuses LowestFee's funding, weight-cap, dust, and change decisions. Add a monotone selected-value bound for pools up to 24 candidates while retaining LowestFee's ordering for larger pools to avoid finite-round starvation. Cover the constrained objective with exhaustive and serialization-edge regressions, and document the migration from Changeless and tuple metrics.
Replace the best-first BinaryHeap frontier with depth-first search that visits the better-bound child first and backtracks in place. This drops per-branch selector/cache clones and, under a round cap, finds complete solutions on large pools where the old frontier often exhausted the budget without a selection.
`LowestFeeChangeless` only applied its selected-value bound to pools of at most 24 candidates. The cap existed because best-first search treats a bound as a priority: a bound that grows with the selection pushed funded branches to the back of the heap, so on a big pool the frontier starved before it reached one. Depth-first search reads a bound as a cut instead of a ranking — it finishes a branch's descendants before its siblings — so the bound can be applied at every pool size, where it prunes inclusion branches that have already overshot the incumbent.
Yield the greedy selection before expanding the first node, and adopt its score as the incumbent. The search is otherwise not anytime: a caller whose round budget runs out before the first complete selection gets `NoBnbSolution::RoundLimit` and falls through to whatever fallback it has, which on a large pool is far worse than the selection a single greedy pass would have handed it for free. Only the incumbent changes, not the bound, so the optimum stays reachable and the improving-solutions contract is unaffected. Metrics that reject the greedy prefix outright — `LowestFeeChangeless`, which will not score a selection that overshoots — are unchanged, and `RoundLimit` still means what it did for them. The two round-count assertions in `tests/bnb.rs` each move by one: the seed is a round.
Bitcoin Core's `SelectCoinsBnB` computes `is_feerate_high` once and lets it decide whether a prune that is only sometimes valid may fire; it does not drop the prune because the general case is unsound. `bound_with_ancestors` took the other route — "never returns `None`" — on the grounds that a fat private deficit can un-fund a prefix a subset would have funded, so infeasibility is not something it may claim. That argument covers "select everything and it is still unfunded". It does not cover the case this relaxation can prove outright: a fee constraint whose deficit the best input still available cannot close at *any* weight. Descendants only add, the deficit is already computed against the branch-wide `ancestor_bump_lower_bound`, and the gain already ignores whatever ancestors those inputs would drag in — so the estimate is optimistic on every axis, and a deficit it still cannot close belongs to an empty subtree. The scan that finds the best value-per-weight candidate already runs, so the test is free. It also prunes the unfunded leaves that had nothing left to add, which the old path could only rank.
Port Bitcoin Core's `SelectCoinsBnB` lookahead. Core keeps a running `curr_available_value` over the coins it has not decided on yet and backtracks as soon as that total cannot close the gap to the target; the cut needs no incumbent, so it fires from the very first descent. We had the same idea only in `LowestFee::bound`'s no-ancestor path, as an O(n) rescan that ran after the relaxation had already been set up, and not at all when the problem has ancestors. `SelectionCache` now carries the value and weight of the undecided candidates worth selecting, maintained by the same add/sub/ban/unban hooks that already track reachable ancestor surplus, so the test is O(1). Two one-sided relaxations keep it from pruning a branch that holds a solution: only candidates with positive standalone effective value count toward the total, and the current ancestor bump is swapped for `ancestor_bump_lower_bound`, which holds for the whole subtree. That second one is what lets the prune run with ancestors present, where funding is not monotone and "select everything and it is still unfunded" would have been an unsound claim.
LowestFee already decides for itself whether a selection should carry a change output, adding one only when it lowers the long-term fee, clears the dust threshold and fits max_weight. A separate changeless objective duplicates that decision and constrains it, and nothing in the crate needs the constraint. Removes LowestFeeChangeless along with the Changeless wrapper the unreleased changelog already retired, plus their tests and proptest regressions. BREAKING CHANGE: LowestFeeChangeless and Changeless are gone. Callers that required a changeless transaction should use LowestFee and inspect the Drain it returns.
`bound_with_ancestors` scanned every undecided candidate at each unfunded node to find the greatest value-per-weight and to notice weightless value. Branch and bound asks for that bound at every unfunded node, so an O(n) scan there made per-node cost grow with the pool: measured on shared_ancestry_*, 2389 ns/round at n=500 rising to 9384 at n=2000, against 385-2056 for the no-ancestry fixtures. The metric already requires candidates in descending value-per-weight order, and that order is keyed on f32. The exact f64 maximum can therefore only lie inside the run sharing the first undecided candidate's f32 key, which is why the old code scanned in f64 rather than taking the first: two exact ratios can tie in f32 and be ordered either way. Scanning just that run keeps the exact answer without touching the tail. Weightless value becomes a counter kept where the undecided aggregates already are. 5.9x to 8.8x faster per round at n=500 to 2000, and byte-identical results: across all 42 benchmark fixtures the score, selection, round count and exhausted flag are unchanged. A debug assertion checks the tie-run result against a full scan, so the ordering assumption is verified on every node the test suite searches.
`SelectionView` overrides these with cache-backed versions, so every call site in the crate and its tests already resolved to the view; the `CoinSelector` copies recomputed the same answers by iterating and had no callers left. Removes `effective_value`, `implied_feerate`, `rate_excess_wu`, `replacement_excess_wu` and `waste`, plus the two private helpers they were the last users of. `missing` and `drain` are deliberately kept even though the view also has them: the crate's own front-page example calls them on a bare `CoinSelector`, which is the case they exist for. The same argument keeps the rest of the overlap -- `weight`, `excess`, `is_funded` and friends all have live callers holding a selector rather than a view, and routing those through `compute_view` would cost an O(n) cache build to replace an O(n) method. BREAKING CHANGE: obtain a `SelectionView` with `CoinSelector::compute_view` and call the removed methods there.
`SelectionView` answered every one of these from its cache while the `CoinSelector` copy recomputed the same figure by iterating the selection. Keeping both meant two implementations of the weight model, the excess model and the ancestor bump, and the slower one was the default a caller reached for. Removes `absolute_excess`, `ancestor_bump`, `ancestor_bump_lower_bound`, `drain`, `drain_value`, `excess`, `fee`, `implied_fee`, `input_weight`, `is_funded`, `is_funded_with_drain`, `is_within_max_weight`, `missing`, `rate_excess`, `replacement_excess`, `selected_value` and `weight` from `CoinSelector`, along with the two private helpers they were the last users of. `select_until` now hands its predicate a `&SelectionView` and maintains that view's cache incrementally, so the greedy pass behind `select_until_target_met` -- which seeds every branch-and-bound search -- costs one cache build plus O(1) per step instead of rescanning the selection on every iteration. The crate's own front-page example now goes through `compute_view` too, which is what the removed methods were kept for. BREAKING CHANGE: obtain a `SelectionView` with `CoinSelector::compute_view` and call the removed methods there. `CoinSelector::select_until` takes a predicate over `&SelectionView` rather than `&CoinSelector`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HLiTkESMktypGJhFag2ZBM
`seed` carried `(CoinSelector, Ordf32)` while `best` separately held the same score. They are set together in `seed_greedy_incumbent` and nothing runs between construction and the first `next()`, so the score in the tuple was always exactly `best`. Store the selection alone and read the score from `best` when yielding. No behaviour change: identical score, selection, round count and exhausted flag on all 42 benchmark fixtures.
`drags_in` and `shared_drags_in` were one dense `Bitset` per candidate over every
ancestor, costing candidates x ancestors bits. On a 200,000-candidate pool with 26,666
ancestors that is 667 MB per array of very nearly nothing: a candidate drags in its
residing transactions and their unconfirmed parents, which measures mean 0.42 entries and
never more than two, so the sets are 0.002% full.
The cost is not only memory. Iterating a dense bitset is O(ancestors) per candidate
however few bits are set, so building the selection cache — which walks every candidate's
shared set — is O(candidates x ancestors) in time too. Setting up a search on 200,000
candidates took 464 ms before expanding a single node, which is enough to lose a
wall-clock budget outright: the benchmark harness reported "no solution" on that fixture
because the deadline expired during construction.
Stored flat instead: one `Vec<u32>` of indices with per-candidate offsets. Every read of
these sets is a full walk of one candidate's entries and they never change after
construction, so a slice is all they need to be. Construction reuses one scratch bitset
rather than allocating per candidate, so the old cost does not reappear while building.
200,000 candidates, 26,666 ancestors peak RSS setup
dense bitset 1,332 MB 464 ms
flat indices + offsets 58 MB 54 ms
`Bitset` is unchanged where it is used over candidates — the selected and banned sets are
dense and membership-tested constantly.
Breaking: `drags_in` and `shared_drags_in` now return `&[u32]` rather than `&Bitset`.
Byte-identical to the parent commit on all 42 benchmark fixtures — same selections,
scores, round counts and exhausted flags. 80 tests green on `--all-features` and
`--no-default-features`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HLiTkESMktypGJhFag2ZBM
The search order is descending `value / weight`, which the `LowestFee` bound depends on and which cannot see ancestry at all: a candidate whose unconfirmed parents cost more to bump than the next candidate is worth still sorts ahead of it. On a pool the search can work through, that is invisible, because the search fixes it. On a pool it cannot — a few hundred thousand candidates, where branch and bound returns the greedy prefix it started from — the ordering *is* the answer, and the blind one drags in parents it did not have to. So take a second greedy prefix, ordered by `(value - own bump) / weight`, and keep whichever of the two the metric scores better. `local_bump` overcounts a shared parent that some other selected candidate would have dragged in anyway, which is why this is an incumbent rather than the order the search runs in: the ordering the bound relies on is untouched, the optimum stays reachable, and the reordered prefix is adopted only when it actually scores better. Costs one greedy pass and one sort, and only when the problem has unconfirmed ancestors at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HLiTkESMktypGJhFag2ZBM
Depth-first traversal is linear in memory but expands whatever is under its feet, so it
prunes against whatever incumbent its dive order happened to find. On problems whose
candidates share unconfirmed ancestors that is far from the optimum, and the search
cannot recover: `subsidizing_ancestry_50` burns forty million nodes without improving on
an incumbent 2.5x worse than the answer a priority queue proves in 55,737.
This runs the same depth-first traversal in passes under a rising ceiling on the bound.
Pass k visits the nodes whose bound is at or below the threshold, which is the set a
priority queue expands before it first pops a node of that bound, so the passes
reconstruct best-first's expansion order without a frontier.
The incumbent carries across passes, and a pass ending with the incumbent at or below the
threshold proves it optimal: any better selection would have had every node on its path
bounded by its own score, so it could not have been pruned by either rule.
The threshold schedule is a speed knob and never a correctness one — raising the
threshold past the smallest rejected bound only ever adds nodes to a pass, never skips
one — so `eps` is free to trade re-expansion against how closely the queue's order is
followed.
`bnb_solutions` is unchanged and takes the plain dive; the new behaviour is opt-in
through `bnb_solutions_with_deepening`.
Measured on coinselect-benchmark's 42 fixtures at a wall-clock budget, eps=0.1:
subsidizing_ancestry_50 40,000,000 nodes, not exhausted, child fee 11,332
-> 64,544 nodes, exhausted, child fee 4,508
shared_ancestry_200 36,242 -> 21,069 nested_ancestry_200 30,203 -> 22,477
subsidizing_ancestry_100 30,140 -> 18,925 subsidizing_ancestry_200 27,281 -> 22,999
Exhausted rises from 31 to 34 of 42 and peak RSS stays flat at 3.5 MB. Where both
traversals exhaust they agree on all 32 fixtures, and the brute-force oracle confirms the
optimum on all 9 fixtures small enough to enumerate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HLiTkESMktypGJhFag2ZBM
Iterative deepening reconstructs a priority queue's node ordering, but it is not anytime:
under a ceiling low enough to be useful the early passes may not reach a complete
selection at all. On a pool too large to exhaust that is strictly worse than diving, which
reaches leaves immediately — measured at +70% on `wallet_mixed_2000`, and no threshold
step fixes both ends, because a step large enough to protect it is large enough to lose
`subsidizing_ancestry_50` outright.
So dive first and deepen after, carrying the incumbent across. The dive hands over once it
has gone as long without an improvement as it took to find the one it holds, which keeps
its budget on a pool that is still creeping downward and gives up quickly on one that is
stuck — the failure this exists to fix.
That rule needs a floor, because the greedy incumbent is set before the first node and so
leaves it nothing to measure against. The floor scales on candidate count rather than on
the budget, which is not visible here: a dive to a leaf costs at most one node per
candidate, so the floor is that depth times a constant. 200 was the best single value over
42 fixtures at three budgets and the metric is not sharply peaked around it.
Wallet track, against the plain dive, eps=0.1:
10 ms -0.48% 2 better, 0 worse
100 ms -3.81% 5 better, 0 worse exhausted 28 -> 31 of 42
1000 ms -3.89% 5 better, 1 worse exhausted 31 -> 33 of 42
`subsidizing_ancestry_50` reaches the optimum of 4,508 after a 10,001-node dive and 8
passes. Peak RSS stays at 3.6 MB. The default path is untouched: no flag, no behaviour
change, byte-identical to the parent commit on all 42 fixtures.
The one regression is opportunity cost, not a lost incumbent: handing over ends the dive,
so against a dive that keeps the whole budget the hybrid can come out behind. It cannot
come out behind a dive given the same dive budget.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HLiTkESMktypGJhFag2ZBM
This was referenced Aug 17, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft. The last three commits, stacked on
experiment/cheap-nodes(#77) — review that one first, since the
tuning here depends on it.
GitHub will not take a base branch that only exists in the fork, so this is opened against
masterand the diff shown includes everything below it. The change to review here is the top three commits.
#79 sits on top of this one.
This changes what every caller of
CoinSelector::bnb_solutionsgets, hencefeat!. The plaindive remains available as
bnb_solutions_dive_only.The problem
Depth-first search reaches complete selections immediately, then prunes against whatever its opening
dive found. On some inputs the dive stalls on a bad incumbent and never recovers. On
subsidizing_ancestry_50the incumbent sticks at 11,912 against a true optimum of 5,088 and doesnot move across 40 million further nodes — every prune it makes for the rest of the search is against
an incumbent two and a half times worse than the answer.
The root cause is in the ordering, not the bound: candidates are sorted by value-per-weight, which
cannot see what a coin's unconfirmed ancestors cost to bump. A best-first priority queue does not
have this problem, because it orders by the bound, which is ancestry-aware — but it pays for that
with a frontier that grows without limit.
What this does
Iterative deepening on the bound. Run the same depth-first traversal in passes under a rising
ceiling, so pass k visits the nodes a priority queue would expand before first popping that bound.
That recovers best-first's node ordering at depth-first's linear memory.
Dive first, then deepen. Deepening alone reaches complete selections late, which costs the large
pools that never exhaust — they spend their budget re-expanding instead of diving deep enough to find
a good complete selection. So: dive until the incumbent stops improving, then hand over to deepening
from the root, keeping the incumbent.
Note what that does not promise, because an earlier draft of this PR got it wrong. Against a dive
stopped at the same handover point the incumbent only improves, so the hybrid cannot be worse.
Against a dive that keeps spending the whole budget descending, it can be — the rounds spent
re-expanding are rounds the dive would have spent going deeper.
Make it the default for
bnb_solutionsandrun_bnb, when the problem has unconfirmedancestors. The plain dive stays available as
bnb_solutions_dive_only. Two tuned constants,DEFAULT_DEEPENING_EPS = 0.1andDEFAULT_DIVE_FLOOR_PER_CANDIDATE = 5, both documented with thesweep that chose them.
The ancestry gate is not a tuning choice. Deepening exists to escape a dive misled by a sort key that
cannot see what a coin's unconfirmed parents cost. With no ancestors there is no such blindness, the
dive does not stall, and re-expanding from the root is a straight loss: over 4,000 random
ancestor-free pools it measures worse on 54 and better on 1, worst case +88%, concentrated in the
2,000-20,000 round band a wallet with a few dozen UTXOs actually sits in. All seven
no_ancestrybenchmark fixtures are byte-identical to the plain dive.
Measured
42 fixtures, 100,000-round budget, against the parent commit.
Total package fee −6.63%. Seven fixtures improved by 1.8% to 48.4%; three regressed by 0.6% to
3.0% (see what it costs).
subsidizing_ancestry_50shared_ancestry_200subsidizing_ancestry_100nested_ancestry_200high_feerate_200shared_ancestry_100subsidizing_ancestry_200Against Bitcoin Core's wallet coin selection, fixtures where coin-select's package fee wins go from
38 of 42 to 41 of 42. The 42nd is not a loss: on
high_feerate_20the brute-force oracleenumerates all 2^20 subsets, and coin-select returns the exact optimum with the tree exhausted. Core
lands on a different selection with the same fee. It is a tie at a proven optimum.
An existing test asserting an exact round count is unchanged at 62,453: its fixture has no
unconfirmed ancestors, so the gate switches deepening off for it. Deepening would reach the same
exact-value solution there in 2,970 rounds — but the gate is set by what happens under a budget,
not by round count at exhaustion.
The two constants
Both are tuned to one benchmark and this should be weighed as such.
DEFAULT_DEEPENING_EPS = 0.1, the relative step between thresholds. A strict schedule — one pass perdistinct bound — costs up to 194x more nodes on the fixtures that already match a priority queue
node for node, because they pay for re-expansion and buy nothing. 0.1 holds that near 1.5x while
collapsing the pass count to single digits.
A trap worth knowing about:
eps = 0.4looks better on any "time to get ahead of Core" metric —40 of 42 instead of 38. It gets there by giving up on the hard fixtures entirely, so it never arrives
late, it just loses on fee (
subsidizing_ancestry_5044,990 against Core's 29,690). Timing metricsalone will select for quitting early.
DEFAULT_DIVE_FLOOR_PER_CANDIDATE = 5, how long the opening dive is protected before the stall rulecan fire. The dive needs a floor at all because the greedy incumbent is set before the first node,
which leaves the stall rule with nothing to measure against. Swept over 0, 5, 20, 50 and 200, at a
round budget and at 3 ms / 10 ms / 100 ms / 1000 ms:
5 wins where budgets are tight (2.59% at 3 ms) and loses by under 0.1% where they are loose. Floor 0
— handing over before the dive completes a single selection — costs 10.7%, so the floor is doing real
work; it just does not need to be large. That the right value is small follows from what the dive is
for: it has to reach one complete selection, which is one root-to-leaf path.
This was 200 in the earlier draft, measured when a node cost several times more. The floor is counted
in nodes, so the parent commit making nodes cheaper turned the same floor into a longer dive in
wall clock, and the tuning had to move with it. Whether a node-counted floor is the right unit at all
is a fair question to raise on this PR.
What it costs
The wins and losses split on pool size with no overlap: everything from n=50 to n=200 improves, and
wallet_mixed_500/_1000/_2000regress by 3.0% / 0.6% / 0.8%.Those three are not budget starvation. Given ten times the budget the dive keeps improving
(
wallet_mixed_1000: 56,633 -> 56,162 -> 56,140) while deepening sits at 56,950 at every budget from100,000 rounds to a full second. Deepening plateaus and the dive does not: on a pool that large no
pass can complete, so re-expansion buys nothing while a dive is still descending into unexplored
parts of the tree.
I tried twice to fix it by handing back to the dive when deepening stops improving — once with the
existing stall rule, once with a tighter one — and both produced byte-identical results on all 42
fixtures. The rule never fires. Both were reverted rather than shipped.
A pool-size gate would score better here. I have not added one: it would be a constant fitted to ten
data points from one benchmark, in a library that runs on wallets this benchmark has never seen. The
trade is stated so it can be argued with.
What it does not fix
Three fixtures —
subsidizing_ancestry_50,subsidizing_ancestry_100,nested_ancestry_200— nowovertake Core on fee by 17–35%, but need roughly 5–10x Core's wall clock to get there. Two things
were tried and did not help, both worth recording:
four of these. Deepening subsumes what it was buying.
variants) produces the identical prefix to today's key on all three. The answer on those
fixtures is not a prefix of any candidate ranking, so no static seed can reach it.
Test plan
cargo test,cargo test --release,cargo clippy --all-features,--no-default-featuresbuild:all pass.
tests/ancestor.rsgainsdeepening_escapes_a_dive_that_the_candidate_order_misleads, which pinsthe reason for the default rather than any round count: two groups, each a fat underpaying root
spent by one overpaying tip and one underpaying one, so coins on the same root share its bump and
what a coin costs depends on what else is selected. The dive commits to spreading across both roots;
deepening comes out 12.6% better at 600 rounds. The test asserts the size of the win, not the round
count, so a change that keeps deepening on but makes it converge later cannot pass silently.
A
deepening_statsmodule of process-global atomics was in an earlier draft for harnessinstrumentation. It has been removed:
AtomicU64does not exist onno_stdtargets without 64-bitatomics, and nothing read it.