Search before asking
Apache SkyWalking Component
BanyanDB (apache/skywalking-banyandb)
What happened
The scheduled backup sidecar (/backup --schedule=@hourly --time-style=daily) logs an error every night even though the backup succeeds:
{"level":"error","module":"BACKUP-SCHEDULER.SCHEDULER.BACKUP","name":"backup","time":"2026-09-26T00:05:27Z","message":"action timed out"}
{"level":"error","module":"BACKUP-SCHEDULER.SCHEDULER.BACKUP","name":"backup","time":"2026-09-27T00:05:26Z","message":"action timed out"}
We saw this on the skywalking-showcase GKE cluster, on the cold data node (500Gi PVC, GCS destination). With --time-style=daily, the midnight run writes into a new date directory, so it re-uploads the whole node. On 2026-09-26 that was 28,393 objects, uploaded from 00:00:38Z to 01:05:21Z (~65 min). The GCS contents are complete, so the backup did finish.
The cause is in pkg/timestamp/scheduler.go (task.run): every scheduled action gets a hard-coded 5-minute timeout (t.clock.Timer(5 * time.Minute)). When it fires, the scheduler:
- logs
action timed out at error level and increments TotalTasksTimeout,
- stops waiting on the action but does not cancel it. The action keeps running detached, and its eventual success or failure is no longer tied to the scheduler.
Consequences:
- Every long backup (any initial or daily full upload of a non-trivial node) produces a false error and a timeout metric. Real failures are hard to tell apart, and alerts on the timeout metric are noisy. With the sidecar's usual
BYDB_LOGGING_LEVEL=error, this false timeout is the only backup-related line in the logs.
- The 5-minute value is not configurable for backup, lifecycle or any other scheduler user.
- The backup command needs its own
backupInFlight guard (banyand/backup/backup.go) because the scheduler starts the next tick while the abandoned action is still running.
What you expected to happen
A scheduled action that takes longer than 5 minutes but succeeds should not be reported as an error. For example:
- make the per-action timeout configurable per registration (or disable it for backup/lifecycle), and/or
- when the timeout fires, either cancel the action's context (a real timeout) or log it at warn/info as "still running", then log the real outcome when the action completes.
How to reproduce
Run banyand-backup --schedule=@hourly --time-style=daily --dest=... against a data node whose full upload takes more than 5 minutes. At the first run after midnight, action timed out is logged at error level while the upload keeps going and later completes.
Anything else
No response
Are you willing to submit a pull request to fix on your own?
Code of Conduct
Search before asking
Apache SkyWalking Component
BanyanDB (apache/skywalking-banyandb)
What happened
The scheduled backup sidecar (
/backup --schedule=@hourly --time-style=daily) logs an error every night even though the backup succeeds:We saw this on the skywalking-showcase GKE cluster, on the cold data node (500Gi PVC, GCS destination). With
--time-style=daily, the midnight run writes into a new date directory, so it re-uploads the whole node. On 2026-09-26 that was 28,393 objects, uploaded from 00:00:38Z to 01:05:21Z (~65 min). The GCS contents are complete, so the backup did finish.The cause is in
pkg/timestamp/scheduler.go(task.run): every scheduled action gets a hard-coded 5-minute timeout (t.clock.Timer(5 * time.Minute)). When it fires, the scheduler:action timed outat error level and incrementsTotalTasksTimeout,Consequences:
BYDB_LOGGING_LEVEL=error, this false timeout is the only backup-related line in the logs.backupInFlightguard (banyand/backup/backup.go) because the scheduler starts the next tick while the abandoned action is still running.What you expected to happen
A scheduled action that takes longer than 5 minutes but succeeds should not be reported as an error. For example:
How to reproduce
Run
banyand-backup --schedule=@hourly --time-style=daily --dest=...against a data node whose full upload takes more than 5 minutes. At the first run after midnight,action timed outis logged at error level while the upload keeps going and later completes.Anything else
No response
Are you willing to submit a pull request to fix on your own?
Code of Conduct