Skip to content

[Bug] BanyanDB scheduled backup reports false "action timed out" errors for long but successful runs #14111

Description

@hanahmily

Search before asking

  • I had searched in the issues and found no similar issues.

Apache SkyWalking Component

BanyanDB (apache/skywalking-banyandb)

What happened

The scheduled backup sidecar (/backup --schedule=@hourly --time-style=daily) logs an error every night even though the backup succeeds:

{"level":"error","module":"BACKUP-SCHEDULER.SCHEDULER.BACKUP","name":"backup","time":"2026-09-26T00:05:27Z","message":"action timed out"}
{"level":"error","module":"BACKUP-SCHEDULER.SCHEDULER.BACKUP","name":"backup","time":"2026-09-27T00:05:26Z","message":"action timed out"}

We saw this on the skywalking-showcase GKE cluster, on the cold data node (500Gi PVC, GCS destination). With --time-style=daily, the midnight run writes into a new date directory, so it re-uploads the whole node. On 2026-09-26 that was 28,393 objects, uploaded from 00:00:38Z to 01:05:21Z (~65 min). The GCS contents are complete, so the backup did finish.

The cause is in pkg/timestamp/scheduler.go (task.run): every scheduled action gets a hard-coded 5-minute timeout (t.clock.Timer(5 * time.Minute)). When it fires, the scheduler:

  • logs action timed out at error level and increments TotalTasksTimeout,
  • stops waiting on the action but does not cancel it. The action keeps running detached, and its eventual success or failure is no longer tied to the scheduler.

Consequences:

  1. Every long backup (any initial or daily full upload of a non-trivial node) produces a false error and a timeout metric. Real failures are hard to tell apart, and alerts on the timeout metric are noisy. With the sidecar's usual BYDB_LOGGING_LEVEL=error, this false timeout is the only backup-related line in the logs.
  2. The 5-minute value is not configurable for backup, lifecycle or any other scheduler user.
  3. The backup command needs its own backupInFlight guard (banyand/backup/backup.go) because the scheduler starts the next tick while the abandoned action is still running.

What you expected to happen

A scheduled action that takes longer than 5 minutes but succeeds should not be reported as an error. For example:

  • make the per-action timeout configurable per registration (or disable it for backup/lifecycle), and/or
  • when the timeout fires, either cancel the action's context (a real timeout) or log it at warn/info as "still running", then log the real outcome when the action completes.

How to reproduce

Run banyand-backup --schedule=@hourly --time-style=daily --dest=... against a data node whose full upload takes more than 5 minutes. At the first run after midnight, action timed out is logged at error level while the upload keeps going and later completes.

Anything else

No response

Are you willing to submit a pull request to fix on your own?

  • Yes I am willing to submit a pull request on my own!

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working and you are sure it's a bug!

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions