Should exponential backoff be... exponential? #73
Replies: 3 comments 1 reply
|
No, you're not missing a reason — the exponential curve exists because it's the right default for unknown-duration failures. It's the wrong tool for the five-hour limit, which is a known-duration failure the server explicitly tells us about. A fixed interval is strictly an improvement for the rate-limit case and a regression for transient errors; honouring Long answer to your questionI'll lay out exactly what the code does today and walk through why neither the current exponential curve nor a flat 15/30 min interval is actually ideal. What the code does todayFrom _backoff=30; _total_waited=0
while [ "$_total_waited" -lt "$MAX_RETRY_WAIT" ]; do
hlog_err "retriable error: ${FATAL_MSG}"
hlog "retrying in ${_backoff}s (waited ${_total_waited}/${MAX_RETRY_WAIT}s)"
printf '%s/%s\n' "$_total_waited" "$MAX_RETRY_WAIT" > "$RETRY_FILE"
sleep "$_backoff"
_total_waited=$((_total_waited + _backoff))
_backoff=$((_backoff * 2))
if [ "$_backoff" -gt 1800 ]; then
_backoff=1800
fiConcrete sleep schedule (start 30 s, double, cap at 1800 s = 30 min):
Worst-case wake-up lag once the backoff saturates: up to 30 min past the true reset time. With Why exponential is the defaultClassical reasons — which I'll argue don't fit the 5-hour case:
All three apply squarely to transient errors: Why exponential is wrong for the 5-hour rate limit specificallyReason 3 above is false here — the API does tell you when it's back. Every {"type":"rate_limit_event","rate_limit_info":{"status":"rejected","resetsAt":1775505600,"rateLimitType":"five_hour"}}(I saw the same in the live Opus 4.7 probe earlier — That's why your observation is sharp: the loop treats a hard-deadline wait (5 h) the same as an unbounded outage, and the doubling curve is increasingly wasteful as you approach the real reset. Two distinct kinds of waste:
Why a fixed 15-30 min interval isn't ideal eitherYou're right that it caps the oversleep to ≤ interval, but it has the mirror problem:
The fix that actually matches the problemBranch on what
Concretely: change |
|
Fantastic! Didn't know about the |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
The Claude usage window is (currently) 5 hours. As the window nears its end, the backoff grows, so the agent sleeps longest right when the new window is about to refresh.
Since
max_retry_waitalready caps things for the weekly-limit case, why is the shape exponential at all? A fixed interval (let's say 15/30 min) bounds the wake-up lag to a small fraction of the window and keeps behavior predictable.Am I missing a reason for the exponential curve?
All reactions