Summary
After arb-reth rewind, the node crash-loops with
Database error: state at block #<db_tip - persistence_threshold> is pruned
every few minutes, and never converges. The state at that block is not pruned — it is 128 blocks old on a node configured with --prune.account-history.distance 1000000.
The cause is that rewind unwinds the database without rolling back the AccountHistory / StorageHistory prune checkpoints. Those checkpoints are left above the new tip, and reth's "is this pruned?" test is a pure checkpoint comparison — it never looks at whether the changesets exist. So every historical state read below the stale checkpoint is refused, permanently.
This is the same class as #44: rewind moves the tip but does not fix up the derived metadata that is keyed to it. #44 was the L1 resume log; this is the prune checkpoints.
Root cause
rewind unwinds via the direct storage path, not the staged pipeline (crates/arb-reth-node/src/commands/rewind.rs:275):
provider_rw.remove_block_and_execution_above(new_tip)?;
remove_block_and_execution_above (provider.rs:3620) ends in update_pipeline_stages_after_unwind, which writes StageCheckpoints only (provider.rs:2418-2444) — PruneCheckpoints is never touched.
The only code in the tree that rolls PruneCheckpoints backward is PruneStage::unwind (crates/stages/stages/src/stages/prune.rs:103-124):
// We cannot recover the data that was pruned in `execute`, so we just update the checkpoints.
for (segment, mut checkpoint) in prune_checkpoints {
if let Some(block) = checkpoint.block_number && input.unwind_to < block {
checkpoint.block_number = Some(input.unwind_to);
...
rewind bypasses the pipeline, so that never runs.
Then the read side refuses on the checkpoint alone — crates/storage/provider/src/providers/consistent.rs:1427-1438:
let account_history_exists = self
.storage_provider
.get_prune_checkpoint(PruneSegment::AccountHistory)?
.and_then(|c| c.block_number.map(|c| database_start > c))
.unwrap_or(true);
if !account_history_exists {
return Err(ProviderError::StateAtBlockPruned(database_start))
}
Same shape at consistent.rs:1170-1188 for StorageHistory, and via LowestAvailableBlocks at state/historical.rs:178,206. Note there is no existence check: the changesets can be fully present and the read still fails.
Why the reported block number is the giveaway
The block in the error is database_start — the boundary between the in-memory chain and the database, i.e. db_tip - persistence_threshold. It is not a property of the pruned range at all.
That is directly observable: with persistence_threshold = 128 the node failed at db_tip - 128; setting it to 16 moved the failure to db_tip - 16. A genuine pruning boundary would not track a tuning knob.
Observed
Robinhood Chain (4663), on a datadir that had been rewound below its prune checkpoint:
db_tip 53,758,423, public tip 54,373,431
Database error: state at block #53758407 is pruned (= tip - 16)
- crash at
crates/arb-reth-engine/src/engine.rs:1628, roughly every 8 minutes
- configured
--prune.account-history.distance 1000000; the node's own startup line confirms history pruning enabled ... account_history: Some(Distance(1000000))
- purely local — an MDBX error, no L1/RPC involvement
No configuration reaches this. The retained window is already 1,000,000; --minimal prunes harder; and no setting rewrites an existing checkpoint.
The error is also invisible by default
engine.rs:261 wraps the inner error away:
builder.apply_pre_execution_changes().wrap_err("apply_pre_execution_changes failed")?;
so the journal shows only apply_pre_execution_changes failed. The real message appears only through the parallel path at prewarm.rs:664, which logs it at debug, i.e. you need RUST_LOG=info,arb-reth::prewarm=debug to see the cause at all. Propagating the inner ProviderError here would have made this a one-minute diagnosis instead of a multi-day one. (This is the same wrapper noted in #52.)
Suggested fix
snapshot import already handles exactly this hazard and can be reused:
commands/snapshot.rs:713 refuses to move a history checkpoint backward past the head
write_snapshot_history_boundaries (commands/snapshot.rs:~998) writes AccountHistory / StorageHistory checkpoints at the new head
rewind should do the equivalent after a successful unwind: for AccountHistory and StorageHistory, if checkpoint.block_number > new_tip, lower it to new_tip (matching PruneStage::unwind). Routing rewind through the pipeline unwind instead would also work, but #19 is a reason not to.
A startup assertion would be cheap insurance too: if any history prune checkpoint is above db_tip, refuse to boot with that as the message, rather than deriving for eight minutes and dying on a wrapped error.
Notes
Summary
After
arb-reth rewind, the node crash-loops withevery few minutes, and never converges. The state at that block is not pruned — it is 128 blocks old on a node configured with
--prune.account-history.distance 1000000.The cause is that
rewindunwinds the database without rolling back theAccountHistory/StorageHistoryprune checkpoints. Those checkpoints are left above the new tip, and reth's "is this pruned?" test is a pure checkpoint comparison — it never looks at whether the changesets exist. So every historical state read below the stale checkpoint is refused, permanently.This is the same class as #44:
rewindmoves the tip but does not fix up the derived metadata that is keyed to it. #44 was the L1 resume log; this is the prune checkpoints.Root cause
rewindunwinds via the direct storage path, not the staged pipeline (crates/arb-reth-node/src/commands/rewind.rs:275):remove_block_and_execution_above(provider.rs:3620) ends inupdate_pipeline_stages_after_unwind, which writesStageCheckpointsonly (provider.rs:2418-2444) —PruneCheckpointsis never touched.The only code in the tree that rolls
PruneCheckpointsbackward isPruneStage::unwind(crates/stages/stages/src/stages/prune.rs:103-124):rewindbypasses the pipeline, so that never runs.Then the read side refuses on the checkpoint alone —
crates/storage/provider/src/providers/consistent.rs:1427-1438:Same shape at
consistent.rs:1170-1188forStorageHistory, and viaLowestAvailableBlocksatstate/historical.rs:178,206. Note there is no existence check: the changesets can be fully present and the read still fails.Why the reported block number is the giveaway
The block in the error is
database_start— the boundary between the in-memory chain and the database, i.e.db_tip - persistence_threshold. It is not a property of the pruned range at all.That is directly observable: with
persistence_threshold = 128the node failed atdb_tip - 128; setting it to16moved the failure todb_tip - 16. A genuine pruning boundary would not track a tuning knob.Observed
Robinhood Chain (4663), on a datadir that had been rewound below its prune checkpoint:
db_tip53,758,423, public tip 54,373,431Database error: state at block #53758407 is pruned(= tip - 16)crates/arb-reth-engine/src/engine.rs:1628, roughly every 8 minutes--prune.account-history.distance 1000000; the node's own startup line confirmshistory pruning enabled ... account_history: Some(Distance(1000000))No configuration reaches this. The retained window is already 1,000,000;
--minimalprunes harder; and no setting rewrites an existing checkpoint.The error is also invisible by default
engine.rs:261wraps the inner error away:so the journal shows only
apply_pre_execution_changes failed. The real message appears only through the parallel path atprewarm.rs:664, which logs it atdebug, i.e. you needRUST_LOG=info,arb-reth::prewarm=debugto see the cause at all. Propagating the innerProviderErrorhere would have made this a one-minute diagnosis instead of a multi-day one. (This is the same wrapper noted in #52.)Suggested fix
snapshot importalready handles exactly this hazard and can be reused:commands/snapshot.rs:713refuses to move a history checkpoint backward past the headwrite_snapshot_history_boundaries(commands/snapshot.rs:~998) writesAccountHistory/StorageHistorycheckpoints at the new headrewindshould do the equivalent after a successful unwind: forAccountHistoryandStorageHistory, ifcheckpoint.block_number > new_tip, lower it tonew_tip(matchingPruneStage::unwind). Routingrewindthrough the pipeline unwind instead would also work, but #19 is a reason not to.A startup assertion would be cheap insurance too: if any history prune checkpoint is above
db_tip, refuse to boot with that as the message, rather than deriving for eight minutes and dying on a wrapped error.Notes
robinhood-ops@427dad9. Line numbers above are from that checkout.PruneCheckpointstable on the affected datadir to read the stale values back — the box is currently unreachable to me. The code path above is the claim; the crash is offered as consistent evidence (in particular thepersistence_thresholdcorrelation, which I do not think has another explanation). Happy to get the actual checkpoint values if useful.rewindOOMs on a large span because it drives reth's unchunked reorg unwind #19 (whyrewindavoids the pipeline unwind), bug: Stylus program call reverts on arb-reth but succeeds on Nitro (status 1 vs 0, ~half gas), silent stateRoot fork #52 (sameapply_pre_execution_changeswrapper hiding the real error; that report also hitstate at block #N is prunedfrombench-execon a datadir whoserewindhad just walked changesets across the same range — which in hindsight is this same bug).