Skip to content

bug: rewind leaves AccountHistory/StorageHistory prune checkpoints above the new tip, so every state read fails "state at block #N is pruned" #56

Description

@atulbansal1986

Summary

After arb-reth rewind, the node crash-loops with

Database error: state at block #<db_tip - persistence_threshold> is pruned

every few minutes, and never converges. The state at that block is not pruned — it is 128 blocks old on a node configured with --prune.account-history.distance 1000000.

The cause is that rewind unwinds the database without rolling back the AccountHistory / StorageHistory prune checkpoints. Those checkpoints are left above the new tip, and reth's "is this pruned?" test is a pure checkpoint comparison — it never looks at whether the changesets exist. So every historical state read below the stale checkpoint is refused, permanently.

This is the same class as #44: rewind moves the tip but does not fix up the derived metadata that is keyed to it. #44 was the L1 resume log; this is the prune checkpoints.

Root cause

rewind unwinds via the direct storage path, not the staged pipeline (crates/arb-reth-node/src/commands/rewind.rs:275):

provider_rw.remove_block_and_execution_above(new_tip)?;

remove_block_and_execution_above (provider.rs:3620) ends in update_pipeline_stages_after_unwind, which writes StageCheckpoints only (provider.rs:2418-2444) — PruneCheckpoints is never touched.

The only code in the tree that rolls PruneCheckpoints backward is PruneStage::unwind (crates/stages/stages/src/stages/prune.rs:103-124):

// We cannot recover the data that was pruned in `execute`, so we just update the checkpoints.
for (segment, mut checkpoint) in prune_checkpoints {
    if let Some(block) = checkpoint.block_number && input.unwind_to < block {
        checkpoint.block_number = Some(input.unwind_to);
        ...

rewind bypasses the pipeline, so that never runs.

Then the read side refuses on the checkpoint alone — crates/storage/provider/src/providers/consistent.rs:1427-1438:

let account_history_exists = self
    .storage_provider
    .get_prune_checkpoint(PruneSegment::AccountHistory)?
    .and_then(|c| c.block_number.map(|c| database_start > c))
    .unwrap_or(true);

if !account_history_exists {
    return Err(ProviderError::StateAtBlockPruned(database_start))
}

Same shape at consistent.rs:1170-1188 for StorageHistory, and via LowestAvailableBlocks at state/historical.rs:178,206. Note there is no existence check: the changesets can be fully present and the read still fails.

Why the reported block number is the giveaway

The block in the error is database_start — the boundary between the in-memory chain and the database, i.e. db_tip - persistence_threshold. It is not a property of the pruned range at all.

That is directly observable: with persistence_threshold = 128 the node failed at db_tip - 128; setting it to 16 moved the failure to db_tip - 16. A genuine pruning boundary would not track a tuning knob.

Observed

Robinhood Chain (4663), on a datadir that had been rewound below its prune checkpoint:

  • db_tip 53,758,423, public tip 54,373,431
  • Database error: state at block #53758407 is pruned (= tip - 16)
  • crash at crates/arb-reth-engine/src/engine.rs:1628, roughly every 8 minutes
  • configured --prune.account-history.distance 1000000; the node's own startup line confirms history pruning enabled ... account_history: Some(Distance(1000000))
  • purely local — an MDBX error, no L1/RPC involvement

No configuration reaches this. The retained window is already 1,000,000; --minimal prunes harder; and no setting rewrites an existing checkpoint.

The error is also invisible by default

engine.rs:261 wraps the inner error away:

builder.apply_pre_execution_changes().wrap_err("apply_pre_execution_changes failed")?;

so the journal shows only apply_pre_execution_changes failed. The real message appears only through the parallel path at prewarm.rs:664, which logs it at debug, i.e. you need RUST_LOG=info,arb-reth::prewarm=debug to see the cause at all. Propagating the inner ProviderError here would have made this a one-minute diagnosis instead of a multi-day one. (This is the same wrapper noted in #52.)

Suggested fix

snapshot import already handles exactly this hazard and can be reused:

  • commands/snapshot.rs:713 refuses to move a history checkpoint backward past the head
  • write_snapshot_history_boundaries (commands/snapshot.rs:~998) writes AccountHistory / StorageHistory checkpoints at the new head

rewind should do the equivalent after a successful unwind: for AccountHistory and StorageHistory, if checkpoint.block_number > new_tip, lower it to new_tip (matching PruneStage::unwind). Routing rewind through the pipeline unwind instead would also work, but #19 is a reason not to.

A startup assertion would be cheap insurance too: if any history prune checkpoint is above db_tip, refuse to boot with that as the message, rather than deriving for eight minutes and dying on a wrapped error.

Notes

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions