Describe the bug
When a connection to api.revenuecat.com is established but the far end then goes silent (socket accepted, no bytes ever returned), the SDK's HTTP layer does not bound the request with any timeout. Purchases.getOfferings() pends indefinitely with no error, no retry, and no fallback-host attempt. We measured:
The SDK's own diagnostics stream proves it saw the hang without surfacing anything to the caller. After we killed the socket, it recorded:
http_request_performed {"host":"api.revenuecat.com","endpoint_name":"get_offerings","response_time_millis":403809,"successful":false,"response_code":-1,"connection_error_reason":"NO_NETWORK"}
http_request_performed {"host":"api.revenuecat.com","endpoint_name":"remote_config","response_time_millis":403830,"successful":false,"response_code":-1,"connection_error_reason":"NO_NETWORK"}
A 403 second response time against timeouts configured in the tens of seconds, and connection_error_reason is reported as NO_NETWORK even though the device's network was healthy the entire time.
Two dedupe mechanisms then amplify one hung socket into a full wedge of the offerings surface:
- Every subsequent
getOfferings() call coalesces onto the hung request: Same call already in progress, adding to callbacks map with key: BackgroundAwareCallbackCacheKey(cacheKey=[/v1/subscribers/<app_user_id>/offerings], appInBackground=false)
- Config refresh is skipped:
Remote config refresh already in progress. Skipping.
So no later call can ever succeed while the first one hangs, and nothing ever times the first one out.
Finally, in runs where we killed the hung socket while leaving the network path healthy, the SDK retried immediately and completed a successful offerings fetch within 29 to 400 ms of the kill. That confirms the network and the API were fine throughout; the only missing piece is a deadline on the hung request.
-
Environment
- Platform: Android. We consume the SDK from a Flutter app in production via purchases_flutter 10.8.0 / purchases-hybrid-common 18.29.0, but the behavior is entirely in the native Android SDK's HTTP layer.
- SDK version: reproduced on purchases-android 10.16.0 and 10.16.1 (BillingClient 8.3.0, not involved in this issue).
- OS version: Android 16 (stock emulator with a Google Play image).
- Android Studio version: n/a (Gradle CLI builds).
- How widespread: any user whose TCP connection to the API goes silent mid-request (NAT timeout, captive portal, flaky mobile network, middlebox) holds the entire offerings surface hostage for as long as the socket survives. In our tests the pend lasted 2.5 to 7 minutes and ended only because we severed the socket externally; nothing in the SDK's behavior suggests it would ever end it. Fleet telemetry shows roughly one silent paywall stall per week in this failure class, and the "late success the moment the socket dies" signature matches what we see in the wild.
-
Debug logs
Verbose SDK logs forwarded through our Flutter log handler; payloads are the SDK's own lines. During the pend, per additional getOfferings() call:
08-08 08:21:36.247 RC debug: Same call already in progress, adding to callbacks map with key: BackgroundAwareCallbackCacheKey(cacheKey=[/v1/subscribers/<app_user_id>/offerings], appInBackground=false)
08-08 08:27:42.451 RC debug: Remote config refresh already in progress. Skipping.
Zero API request started / error / retry lines for the hung endpoint for the entire pend. Then, immediately after we killed the hung socket (network healthy):
08-08 08:28:17.495 RC diagnostics: http_request_performed {"host":"api.revenuecat.com","endpoint_name":"get_offerings","response_time_millis":403809,"successful":false,"response_code":-1,"connection_error_reason":"NO_NETWORK"}
[new request attempts follow within ~30-80ms of the socket dying]
In the variant where the network path was healthy at socket death, those immediate retries completed a successful offerings fetch (10.16.0: success 0.4s after the kill following a 155.5s pend; 10.16.1: getOfferings completion 29ms and offerings delivered 38ms after the kill following a 192.0s pend).
-
Steps to reproduce, with a description of expected vs. actual behavior
-
Launch an Android emulator with -http-proxy http://127.0.0.1:<port> pointing at a small proxy that accepts the CONNECT for api.revenuecat.com (and the fallback host) and then forwards nothing, keeping the socket open and silent. This models an established connection whose far end stops responding.
-
Configure the SDK (verbose logs) and call Purchases.getOfferings().
-
Observe the pend: no callback, no timeout, no retry, no fallback host, indefinitely.
-
Call getOfferings() again several times; each logs "Same call already in progress" and joins the hung request. Trigger a foreground config refresh; it logs "Remote config refresh already in progress. Skipping."
-
Kill the hung socket at the proxy while switching the proxy to pass traffic through normally.
-
The SDK instantly retries and succeeds; the diagnostics entry for the hung request reports the full multi-hundred-second response_time_millis.
Expected: the request fails with a network error after the HTTP layer's timeout budget (tens of seconds at most), callers get an error callback, and a later getOfferings() can start a fresh request.
Actual: the request hangs for as long as the socket survives (402s / 155s / 192s measured), every concurrent caller is silently parked on it, and nothing is ever surfaced to the app.
- Other information
The connect/read timeout configuration appears to only bound the connection handshake and inter-byte gaps in some paths, not a fully silent established stream on this path; whatever the mechanism, the observed behavior on both 10.16.0 and 10.16.1 is that no deadline fires. The request-level coalescing and the config-refresh skip are reasonable dedupe designs on their own, but combined with an unbounded request they convert one bad socket into a permanent wedge of every future offerings call.
We have a fix in progress and intend to submit a PR (an overall deadline on dispatched requests so hung sockets fail loudly and coalesced callers get an error callback).
Related: #3922 covers an independent wedge we reproduced in the same investigation (billing connection ladder permanently stops after a disconnect lands during an in-flight reconnect). Either defect alone produces the same production symptom: getOfferings() pending silently and indefinitely at a paywall.
Additional context
Production impact for us: a paywall spinner that never resolves, with no signal the app can act on. The immediate-success-after-socket-death behavior means affected users sometimes see offerings load minutes later for no visible reason, which matches the diagnostics we quoted (response_time_millis=403809, successful=false) followed by an instant successful retry.
Describe the bug
When a connection to
api.revenuecat.comis established but the far end then goes silent (socket accepted, no bytes ever returned), the SDK's HTTP layer does not bound the request with any timeout.Purchases.getOfferings()pends indefinitely with no error, no retry, and no fallback-host attempt. We measured:The SDK's own diagnostics stream proves it saw the hang without surfacing anything to the caller. After we killed the socket, it recorded:
A 403 second response time against timeouts configured in the tens of seconds, and
connection_error_reasonis reported asNO_NETWORKeven though the device's network was healthy the entire time.Two dedupe mechanisms then amplify one hung socket into a full wedge of the offerings surface:
getOfferings()call coalesces onto the hung request:Same call already in progress, adding to callbacks map with key: BackgroundAwareCallbackCacheKey(cacheKey=[/v1/subscribers/<app_user_id>/offerings], appInBackground=false)Remote config refresh already in progress. Skipping.So no later call can ever succeed while the first one hangs, and nothing ever times the first one out.
Finally, in runs where we killed the hung socket while leaving the network path healthy, the SDK retried immediately and completed a successful offerings fetch within 29 to 400 ms of the kill. That confirms the network and the API were fine throughout; the only missing piece is a deadline on the hung request.
Environment
Debug logs
Verbose SDK logs forwarded through our Flutter log handler; payloads are the SDK's own lines. During the pend, per additional
getOfferings()call:Zero
API request started/ error / retry lines for the hung endpoint for the entire pend. Then, immediately after we killed the hung socket (network healthy):In the variant where the network path was healthy at socket death, those immediate retries completed a successful offerings fetch (10.16.0: success 0.4s after the kill following a 155.5s pend; 10.16.1: getOfferings completion 29ms and offerings delivered 38ms after the kill following a 192.0s pend).
Steps to reproduce, with a description of expected vs. actual behavior
Launch an Android emulator with
-http-proxy http://127.0.0.1:<port>pointing at a small proxy that accepts the CONNECT forapi.revenuecat.com(and the fallback host) and then forwards nothing, keeping the socket open and silent. This models an established connection whose far end stops responding.Configure the SDK (verbose logs) and call
Purchases.getOfferings().Observe the pend: no callback, no timeout, no retry, no fallback host, indefinitely.
Call
getOfferings()again several times; each logs "Same call already in progress" and joins the hung request. Trigger a foreground config refresh; it logs "Remote config refresh already in progress. Skipping."Kill the hung socket at the proxy while switching the proxy to pass traffic through normally.
The SDK instantly retries and succeeds; the diagnostics entry for the hung request reports the full multi-hundred-second
response_time_millis.Expected: the request fails with a network error after the HTTP layer's timeout budget (tens of seconds at most), callers get an error callback, and a later
getOfferings()can start a fresh request.Actual: the request hangs for as long as the socket survives (402s / 155s / 192s measured), every concurrent caller is silently parked on it, and nothing is ever surfaced to the app.
The connect/read timeout configuration appears to only bound the connection handshake and inter-byte gaps in some paths, not a fully silent established stream on this path; whatever the mechanism, the observed behavior on both 10.16.0 and 10.16.1 is that no deadline fires. The request-level coalescing and the config-refresh skip are reasonable dedupe designs on their own, but combined with an unbounded request they convert one bad socket into a permanent wedge of every future offerings call.
We have a fix in progress and intend to submit a PR (an overall deadline on dispatched requests so hung sockets fail loudly and coalesced callers get an error callback).
Related: #3922 covers an independent wedge we reproduced in the same investigation (billing connection ladder permanently stops after a disconnect lands during an in-flight reconnect). Either defect alone produces the same production symptom:
getOfferings()pending silently and indefinitely at a paywall.Additional context
Production impact for us: a paywall spinner that never resolves, with no signal the app can act on. The immediate-success-after-socket-death behavior means affected users sometimes see offerings load minutes later for no visible reason, which matches the diagnostics we quoted (
response_time_millis=403809,successful=false) followed by an instant successful retry.