I hit this trying to recover from #18. The node had derived ~6.3M blocks past the point where it silently diverged from canonical, so the documented fix is rewind --diverged-at 18864425 — and that rewind cannot complete. It allocates until the kernel kills it.
Reverting a 25,205,858 → 17,200,000 datadir (8.0M blocks) was OOM-killed after 1h48m on a 32 GB box (MemTotal 32,763,980 kB) that also had 16 GiB of swap. At the kill the process held 30.3 GiB resident and every byte of swap was gone (Free swap = 0kB), so the real demand was roughly 46 GiB of anonymous memory before the kernel stepped in. The command attempts the whole span as one aggregate computation instead of chunking it, so the cost scales with how far you are rewinding rather than staying flat.
The good news, and why this is cheap for you to reproduce: the datadir survives. MDBX rolls the transaction back, current_tip is unchanged afterwards and the static-file segment tips stay consistent. I ran it three times against the same directory with no damage.
The practical impact is that the documented recovery path stops working exactly when it is needed. rewind --diverged-at N exists for the case where the node has been running on a bad chain for a while — and the longer it ran, the more certain the recovery is to fail. It is also worth knowing that an uncapped run takes the host with it: each OOM made sshd stop answering for about two minutes (TCP connects, no banner, load average ~25) as the machine thrashed its swap dry — 6,547,350 major page faults on the failing run against 1,003,338 on the 1M-block run that succeeded. It recovers on its own, but it is a surprise if you are on the far end of that SSH session.
Reproduction
Run it against a throwaway copy — a rewind that spans a few million blocks is enough.
arb-reth rewind --datadir /path/to/data-test \
--chain-info chain-info.json --genesis genesis.json \
--chain-id 4663 --to 17200000
The binary is a release build of an unmodified ae3f91f checkout (locally named cleanbin-ae3f91f, which is the process name in the kernel log below). It ran for 1:48:01 and exited 137. A second run of the identical rewind reached the same point, with the two anon-RSS peaks 948 kB apart out of 31.8 million — 0.003% — so the failure is deterministic rather than a one-off. In both, all 16 GiB of swap was consumed at the kill, so the RSS figures are a floor on the real footprint rather than the peak.
Repeating it with a 1M-block span from the same tip succeeds:
| Span |
Blocks |
Changesets to revert |
Peak RSS |
Swap at end |
Result |
| 17,200,001 → 25,205,858 |
8.0M |
175,238,861 |
≥31,818,076 kB |
16 GiB, fully consumed |
OOM |
| 24,205,859 → 25,205,858 |
1.0M |
5,948,297 |
10,111,468 kB |
not measured, but no pressure at 10.1 GB of 31 GiB |
exit 0 in 7:04 |
One caveat on that second row, because it makes the code look better than it is: that window lies entirely inside the range my node had mis-executed under #18, and those blocks are much cheaper than real ones — 5.9 changesets per block against roughly 83 in the correctly-executed range below the divergence. So a 1M unwind of a healthy chain would carry something like 14x the rows I measured. I have not measured that case, and I am not claiming a scaling law from two points: 29.7x the changesets produced at least 3.1x the memory before the process died, which is steep but is a floor, not a curve. What the 1M run does establish is that 10 GB is needed to unwind a million of the cheapest possible blocks, so this is not a problem that only shows up at absurd spans.
The dry run reports the size of the job up front, which is a useful signal to quote:
INFO revert-range changesets (must be > 0 for state to revert)
account_changesets=51439827 storage_changesets=123799034
And the run itself shows where it goes:
INFO unwinding database (this may take a while) removing_above=17200000
WARN Changeset cache MISS in range, falling back to aggregate DB-based computation
start_block=17200001 end_block=25205858
Command terminated by signal 9
Elapsed (wall clock) time (h:mm:ss or m:ss): 1:48:01
Maximum resident set size (kbytes): 31821788
Thirty minutes elapse between the unwind starting and the cache-miss warning, then the process is killed with nothing further logged. The kernel confirms it, and shows that swap was gone too:
Node 0 active_anon:22514300kB inactive_anon:9387692kB active_file:1944kB inactive_file:2280kB ...
Free swap = 0kB
Total swap = 16777212kB
Out of memory: Killed process 2932978 (cleanbin-ae3f91) total-vm:8653885012kB,
anon-rss:31818076kB, file-rss:3072kB, shmem-rss:0kB, UID:0 pgtables:178428kB oom_score_adj:0
The 31.9 GB of system-wide anonymous memory in that first line is essentially all this one process, and the 16.8 GB of swap alongside it was released as soon as the process died, so the ~46 GiB total is attributable to the rewind rather than to anything else on the box.
I reproduced this on ae3f91f, which is current main as I write this, and the reth pin at Cargo.lock:9613 is paradigmxyz/reth rev ab273f171ad87248f0c3ea3e5f1d2600c581715f — unchanged by anything on main since — so it applies to current main rather than to some older state I happened to be on.
Why it happens: rewind uses reth's reorg unwind, not its deep-unwind pipeline
reth has two ways to unwind, and they are bounded very differently.
The one rewind uses is the reorg primitive, at crates/arb-reth-node/src/commands/rewind.rs:266:
let provider_rw = factory.database_provider_rw()?;
provider_rw.remove_block_and_execution_above(new_tip)?;
provider_rw.commit().map_err(|e| eyre::eyre!("commit unwind: {e}"))?;
One call, one transaction, whatever the span. In reth that lands at provider.rs:3489 and immediately delegates to unwind_trie_state_from (provider.rs:851), which collects three unbounded things in a row:
pub fn unwind_trie_state_from(&self, from: BlockNumber) -> ProviderResult<()> {
let changed_accounts = self.account_changesets_range(from..)?; // :852
...
let changed_storages = self.storage_changesets_range(from..)?; // :860
...
let trie_revert = self.changeset_cache.get_or_compute_range(self, from..=db_tip_block)?; // :879
Note from.. — open-ended to the tip, collected into a Vec. There is no chunking, no streaming and no row limit anywhere in that function. In my 8M case those two collections are 51.4M and 123.8M rows, and then get_or_compute_range computes the aggregate trie revert across the same 8,005,858 blocks in one pass. That path is entirely reasonable where reth calls it from — crates/engine/tree/src/persistence.rs:140, for live reorgs of a handful of blocks.
For deep unwinds reth uses Pipeline::unwind (crates/stages/api/src/pipeline/mod.rs:303) instead, which chunks at two levels: stages are unwound in reverse execution order (:320, self.stages.iter_mut().rev()), and each stage is called repeatedly in a loop until it reaches the target (:346, while checkpoint.block_number > to). Each call takes only a slice, via UnwindInput::unwind_block_range_with_threshold at crates/stages/api/src/stage.rs:184:
pub fn unwind_block_range_with_threshold(&self, threshold: u64)
-> (RangeInclusive<BlockNumber>, BlockNumber, bool) {
let mut start = self.unwind_to + 1;
let end = self.checkpoint;
start = max(start, end.block_number.saturating_sub(threshold));
let unwind_to = start - 1;
let is_final_range = unwind_to == self.unwind_to;
Seven stages use it, and the defaults in crates/config/src/config.rs are 10,000 blocks for the hashing stages and 100,000 for merkle, alongside commit_entries: 30_000_000 — a cap on rows rather than blocks, which my 175M-row unwind would exceed roughly sixfold on its own.
The clearest evidence that this is the intended path for the job is that reth's own equivalent command uses it. reth stage unwind — the direct analogue of arb-reth rewind — does this at crates/cli/commands/src/stage/unwind.rs:76:
pipeline.unwind(target, None)?;
So the short-span assumption behind remove_block_and_execution_above is real but implicit: nothing enforces it, and rewind passes it an operator-chosen span instead.
One smaller thing in the same shape, which is not what causes the OOM: rewind.rs:242-243 materializes the entire changeset range into two Vecs purely to log .len() and one sample. Those drop at :257, before the unwind starts, so they did not contribute to the peak — but they do mean the command loads 175M rows twice for a large span.
I could not point at a test that encodes the expectation either way, because rewind.rs currently has no #[test] or #[cfg(test)] at all.
How this might be fixed
The direction that matches reth is to chunk the unwind: loop over slices of the span, calling remove_block_and_execution_above and committing per chunk, rather than passing the whole range in one transaction. I can say this works in practice because it is what I ended up doing by hand — successive 1M-block rewinds, each peaking near 10 GB and taking about seven minutes — and reth's own chunk defaults are 10 to 100 times smaller than that, so there is plenty of headroom to pick something conservative.
Two things I want to be straight about rather than guess at.
I do not think Pipeline::unwind can simply be reused here. This crate depends on reth-stages-api only for MetricEvent (engine.rs:1070) and StageCheckpoint/StageId (snapshot.rs:89), and builds no staged pipeline at all, since blocks come from L1 derivation rather than staged sync. So the chunked loop would have to be written locally rather than delegated — you know far better than I do whether that is worth it versus some other approach.
And chunking would change behaviour in a way worth deciding deliberately: today the single transaction means a failure rolls the entire rewind back, which is exactly why my datadir survived three OOMs intact. Committing per chunk gives that up — an interrupted rewind would leave a consistent database at an intermediate tip instead of the original one. That is arguably better, since progress survives and the operator can resume, but it is a real change and not a free one.
I have not read enough of the surrounding unwind code to say which chunk size is safe with respect to the static-file segments and the resume log, so I have deliberately not proposed a patch.
I hit this trying to recover from #18. The node had derived ~6.3M blocks past the point where it silently diverged from canonical, so the documented fix is
rewind --diverged-at 18864425— and that rewind cannot complete. It allocates until the kernel kills it.Reverting a 25,205,858 → 17,200,000 datadir (8.0M blocks) was OOM-killed after 1h48m on a 32 GB box (
MemTotal32,763,980 kB) that also had 16 GiB of swap. At the kill the process held 30.3 GiB resident and every byte of swap was gone (Free swap = 0kB), so the real demand was roughly 46 GiB of anonymous memory before the kernel stepped in. The command attempts the whole span as one aggregate computation instead of chunking it, so the cost scales with how far you are rewinding rather than staying flat.The good news, and why this is cheap for you to reproduce: the datadir survives. MDBX rolls the transaction back,
current_tipis unchanged afterwards and the static-file segment tips stay consistent. I ran it three times against the same directory with no damage.The practical impact is that the documented recovery path stops working exactly when it is needed.
rewind --diverged-at Nexists for the case where the node has been running on a bad chain for a while — and the longer it ran, the more certain the recovery is to fail. It is also worth knowing that an uncapped run takes the host with it: each OOM madesshdstop answering for about two minutes (TCP connects, no banner, load average ~25) as the machine thrashed its swap dry — 6,547,350 major page faults on the failing run against 1,003,338 on the 1M-block run that succeeded. It recovers on its own, but it is a surprise if you are on the far end of that SSH session.Reproduction
Run it against a throwaway copy — a rewind that spans a few million blocks is enough.
The binary is a release build of an unmodified
ae3f91fcheckout (locally namedcleanbin-ae3f91f, which is the process name in the kernel log below). It ran for 1:48:01 and exited 137. A second run of the identical rewind reached the same point, with the two anon-RSS peaks 948 kB apart out of 31.8 million — 0.003% — so the failure is deterministic rather than a one-off. In both, all 16 GiB of swap was consumed at the kill, so the RSS figures are a floor on the real footprint rather than the peak.Repeating it with a 1M-block span from the same tip succeeds:
exit 0in 7:04One caveat on that second row, because it makes the code look better than it is: that window lies entirely inside the range my node had mis-executed under #18, and those blocks are much cheaper than real ones — 5.9 changesets per block against roughly 83 in the correctly-executed range below the divergence. So a 1M unwind of a healthy chain would carry something like 14x the rows I measured. I have not measured that case, and I am not claiming a scaling law from two points: 29.7x the changesets produced at least 3.1x the memory before the process died, which is steep but is a floor, not a curve. What the 1M run does establish is that 10 GB is needed to unwind a million of the cheapest possible blocks, so this is not a problem that only shows up at absurd spans.
The dry run reports the size of the job up front, which is a useful signal to quote:
And the run itself shows where it goes:
Thirty minutes elapse between the unwind starting and the cache-miss warning, then the process is killed with nothing further logged. The kernel confirms it, and shows that swap was gone too:
The 31.9 GB of system-wide anonymous memory in that first line is essentially all this one process, and the 16.8 GB of swap alongside it was released as soon as the process died, so the ~46 GiB total is attributable to the rewind rather than to anything else on the box.
I reproduced this on
ae3f91f, which is currentmainas I write this, and the reth pin atCargo.lock:9613isparadigmxyz/rethrevab273f171ad87248f0c3ea3e5f1d2600c581715f— unchanged by anything onmainsince — so it applies to currentmainrather than to some older state I happened to be on.Why it happens:
rewinduses reth's reorg unwind, not its deep-unwind pipelinereth has two ways to unwind, and they are bounded very differently.
The one
rewinduses is the reorg primitive, atcrates/arb-reth-node/src/commands/rewind.rs:266:One call, one transaction, whatever the span. In reth that lands at
provider.rs:3489and immediately delegates tounwind_trie_state_from(provider.rs:851), which collects three unbounded things in a row:Note
from..— open-ended to the tip, collected into aVec. There is no chunking, no streaming and no row limit anywhere in that function. In my 8M case those two collections are 51.4M and 123.8M rows, and thenget_or_compute_rangecomputes the aggregate trie revert across the same 8,005,858 blocks in one pass. That path is entirely reasonable where reth calls it from —crates/engine/tree/src/persistence.rs:140, for live reorgs of a handful of blocks.For deep unwinds reth uses
Pipeline::unwind(crates/stages/api/src/pipeline/mod.rs:303) instead, which chunks at two levels: stages are unwound in reverse execution order (:320,self.stages.iter_mut().rev()), and each stage is called repeatedly in a loop until it reaches the target (:346,while checkpoint.block_number > to). Each call takes only a slice, viaUnwindInput::unwind_block_range_with_thresholdatcrates/stages/api/src/stage.rs:184:Seven stages use it, and the defaults in
crates/config/src/config.rsare 10,000 blocks for the hashing stages and 100,000 for merkle, alongsidecommit_entries: 30_000_000— a cap on rows rather than blocks, which my 175M-row unwind would exceed roughly sixfold on its own.The clearest evidence that this is the intended path for the job is that reth's own equivalent command uses it.
reth stage unwind— the direct analogue ofarb-reth rewind— does this atcrates/cli/commands/src/stage/unwind.rs:76:So the short-span assumption behind
remove_block_and_execution_aboveis real but implicit: nothing enforces it, andrewindpasses it an operator-chosen span instead.One smaller thing in the same shape, which is not what causes the OOM:
rewind.rs:242-243materializes the entire changeset range into twoVecs purely to log.len()and one sample. Those drop at:257, before the unwind starts, so they did not contribute to the peak — but they do mean the command loads 175M rows twice for a large span.I could not point at a test that encodes the expectation either way, because
rewind.rscurrently has no#[test]or#[cfg(test)]at all.How this might be fixed
The direction that matches reth is to chunk the unwind: loop over slices of the span, calling
remove_block_and_execution_aboveand committing per chunk, rather than passing the whole range in one transaction. I can say this works in practice because it is what I ended up doing by hand — successive 1M-block rewinds, each peaking near 10 GB and taking about seven minutes — and reth's own chunk defaults are 10 to 100 times smaller than that, so there is plenty of headroom to pick something conservative.Two things I want to be straight about rather than guess at.
I do not think
Pipeline::unwindcan simply be reused here. This crate depends onreth-stages-apionly forMetricEvent(engine.rs:1070) andStageCheckpoint/StageId(snapshot.rs:89), and builds no staged pipeline at all, since blocks come from L1 derivation rather than staged sync. So the chunked loop would have to be written locally rather than delegated — you know far better than I do whether that is worth it versus some other approach.And chunking would change behaviour in a way worth deciding deliberately: today the single transaction means a failure rolls the entire rewind back, which is exactly why my datadir survived three OOMs intact. Committing per chunk gives that up — an interrupted rewind would leave a consistent database at an intermediate tip instead of the original one. That is arguably better, since progress survives and the operator can resume, but it is a real change and not a free one.
I have not read enough of the surrounding unwind code to say which chunk size is safe with respect to the static-file segments and the resume log, so I have deliberately not proposed a patch.