-
Notifications
You must be signed in to change notification settings - Fork 6.3k
[br] UnmarshalDir may panic after a cloud storage WalkDir error #70807
Copy link
Copy link
Closed
Labels
affects-25.10This bug affects the TiDB X 25.10.x versions.This bug affects the TiDB X 25.10.x versions.affects-26.3This bug affects the TiDB X 26.3.x versions.This bug affects the TiDB X 26.3.x versions.affects-8.5This bug affects the 8.5.x(LTS) versions.This bug affects the 8.5.x(LTS) versions.component/brThis issue is related to BR of TiDB.This issue is related to BR of TiDB.found-by-aiFound by AI-assisted testing or analysisFound by AI-assisted testing or analysisseverity/majortype/bugThe issue is confirmed as a bug.The issue is confirmed as a bug.
Description
Activity
Metadata
Metadata
Assignees
Labels
affects-25.10This bug affects the TiDB X 25.10.x versions.This bug affects the TiDB X 25.10.x versions.affects-26.3This bug affects the TiDB X 26.3.x versions.This bug affects the TiDB X 26.3.x versions.affects-8.5This bug affects the 8.5.x(LTS) versions.This bug affects the 8.5.x(LTS) versions.component/brThis issue is related to BR of TiDB.This issue is related to BR of TiDB.found-by-aiFound by AI-assisted testing or analysisFound by AI-assisted testing or analysisseverity/majortype/bugThe issue is confirmed as a bug.The issue is confirmed as a bug.
Bug Report
1. Minimal reproduce step (Required)
storage.UnmarshalDir.WalkDiroperation return a transient error after at least one metadata-read worker has already been scheduled. Examples include an object-storage 5xx response, a network timeout, a failed pagination request, or credentials being revoked during the listing.The v8.5.8 runtime path can be reproduced with a storage test double that schedules one worker, returns a
WalkDirerror, and delays the worker'sReadFile/unmarshal operation until afterWalkDirreturns.2. What did you expect to see? (Required)
BR should wait for all scheduled metadata workers, return the original object-storage error, and exit the restore operation cleanly. It must not panic while reporting an external-storage failure.
3. What did you see instead (Required)
The
UnmarshalDirreader closes its result channel whenWalkDirreturns an error, without first waiting for the workers that were already submitted. A worker can subsequently send a parsed metadata value to the closed channel, producing:This was directly reproduced in the v8.5.8 runtime path. The trigger is not limited to a synthetic invalid metadata format: a transient failure while listing cloud objects can provide the same interleaving. Depending on the call path, this can terminate BR during restore/PiTR metadata loading and leave the operation incomplete.
Relevant code:
br/pkg/storage/helper.go,UnmarshalDir.4. What is your TiDB version? (Required)
TiDB v8.5.8, commit
8b857efa20363d50a8fa2ea7dd9809a85a61b115.Related change: PR #64850, which introduced the current BR compact-log-restore storage helper implementation.