Conversation
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
Anty0
reviewed
Sep 9, 2026
Comment on lines
+23
to
+38
| fun <T> tryWithLocking( | ||
| name: String, | ||
| waitTime: Duration, | ||
| fn: () -> T, | ||
| ): T? { | ||
| val lock = getLock(name) | ||
| if (!lock.tryLock(waitTime.toMillis(), TimeUnit.MILLISECONDS)) { | ||
| return null | ||
| } | ||
| try { | ||
| return fn() | ||
| } finally { | ||
| lock.unlock() | ||
| } | ||
| } | ||
|
|
Member
Author
There was a problem hiding this comment.
@Anty0 which one?
tryLock is only in this function
bdshadow
force-pushed
the
bdshadow/rate-limit-concurrency-cap
branch
from
September 9, 2026 15:19
fee3fd9 to
2093ed9
Compare
Anty0
reviewed
Sep 14, 2026
Comment on lines
+34
to
+51
| override fun <T> tryWithLocking( | ||
| name: String, | ||
| waitTime: Duration, | ||
| fn: () -> T, | ||
| ): T? { | ||
| val lock = this.getLock(name) | ||
| if (!lock.tryLock(waitTime.toMillis(), TimeUnit.MILLISECONDS)) { | ||
| return null | ||
| } | ||
| try { | ||
| return fn() | ||
| } finally { | ||
| if (lock.isHeldByCurrentThread) { | ||
| lock.unlock() | ||
| } | ||
| } | ||
| } | ||
|
|
Member
There was a problem hiding this comment.
Thought this is basically the same function.
…ing threads When many requests share one rate limit bucket (e.g. a single API key used by a large fleet of shipped clients), every worker thread queues indefinitely on the bucket's distributed lock. The lock serves only a few dozen verdicts per second under contention, so the queue grows without bound and the whole thread pool can end up parked on a single bucket, making the node unresponsive for all tenants while it still passes health checks. RateLimitService now: - caps concurrent consumers per bucket per node (tolgee.rate-limits.max-concurrent-per-bucket, default 50) and rejects requests above the cap with 429 without touching the lock; the default leaves room for legitimate bursts such as apps fetching 40+ namespaces in parallel on startup, - waits for the bucket lock at most tolgee.rate-limits.lock-wait-ms (default 500) via the new LockingProvider.tryWithLocking instead of blocking forever, rejecting with 429 on timeout. Rejections are counted in the new tolgee.ratelimit.concurrency_rejections metric, tagged by reason (concurrency_cap / lock_timeout). No strike is recorded for these rejections since the bucket cannot be read safely without the lock.
bdshadow
force-pushed
the
bdshadow/rate-limit-concurrency-cap
branch
from
September 21, 2026 12:51
2093ed9 to
f8d22a1
Compare
The Redisson override of tryWithLocking repeated the whole interface default only to change the release call. The interface now exposes a releaseLock hook (default: plain unlock). Redisson overrides only the hook with releaseEvenIfInterrupted and uses it in all three locking methods.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
All rate limit verdicts for a bucket — including "you are rate limited, 429" — are serialized behind a distributed Redisson lock, and
RateLimitService.consumeBucketUnlessblocks on it indefinitely (lock.lock()).When many requests share one bucket (e.g. a single API key used by a large fleet of shipped clients), this becomes a self-inflicted denial of service: a turn on the lock costs ~4 Redis round-trips (acquire, bucket get, bucket/strike put, unlock + pub/sub), and Redisson's unfair wake-all handoff degrades to ~25–50 ms per turn with many waiters — so one bucket can only serve a few dozen verdicts per second. Once arrival on the bucket exceeds that, the queue grows without bound (wait = queue position × turn time), clients time out and retry — refilling the queue faster than it drains — and the entire worker pool ends up parked on a single bucket lock. The node stops serving all tenants while still passing health checks on the management port.
Per-IP rate limiting (proxy or app-level) cannot catch this shape: each client stays far below any per-IP threshold; it is the aggregate on one bucket that kills the node. And the failure mode is concurrency × duration, not request rate — so the fix is to stop queueing worker threads, not to tune limits.
Fix
RateLimitServiceno longer queues unboundedly on the bucket lock:tolgee.rate-limits.max-concurrent-per-bucket(default 50) requests are rejected with 429 immediately, without touching the lock. Node-local by design: the protected resource (worker threads) is node-local, and the cap keeps working even when Redis itself is slow. Set to 0 to disable.LockingProvider.tryWithLocking(name, waitTime, fn)withtolgee.rate-limits.lock-wait-ms(default 500 ms) instead of an infinitelock.lock(); timeout → 429.tryWithLockingis a default interface method;RedissonLockingProvideroverrides it only to keep itsisHeldByCurrentThreadunlock guard.Neither rejection records a strike (the bucket cannot be read or written safely without holding the lock); both use
retryAfter = lock-wait-ms.New metric:
tolgee.ratelimit.concurrency_rejectionscounter, taggedreason=concurrency_cap|lock_timeout.Behavior change
A client sending more than
max-concurrent-per-buckettruly simultaneous requests per node to one bucket now gets 429s for the excess instead of having them queue briefly. The default of 50 covers the largest known legitimate burst — apps with 40+ namespaces fetch them all in parallel on startup through one API key (one bucket), even if the whole burst lands on a single node — while still limiting a single hot bucket to ~25% of a node's worker pool (~200 threads). The cap only bites the flood shape described above.Testing
RateLimitServiceTest: rejection within the lock-wait bound when the lock is held; cap rejection happens without waiting for the lock (timing-asserted against a 10 s lock wait); cap on one bucket does not affect other buckets;max-concurrent-per-bucket = 0restores queueing. Metric counters asserted throughout.RateLimitServiceTest,:securityrate limit tests, andRateLimitsTest/RedisRateLimitsTestintegration suites pass unchanged.