Retry transient inference proxy unavailability - #1994
Conversation
A transient router<->worker breakdown surfaces as httpx.RequestError (stale/dead
cached proxy URL) or RuntimeError('...no proxy URL published...') (router dropped its
worker, not yet republished). The engine/worker stays alive across these, so
_forward_with_retry now poll-retries with a re-read proxy URL (SKYRL_PROXY_RETRY_TIMEOUT_SEC
default 900s, backoff 5s) instead of failing the whole run on the first blip. Genuine
vLLM 4xx/5xx still surface. Fixes 27B PiSSA drivers dying on a recoverable router blip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request introduces a retry loop with configurable backoff and timeout in _forward_with_retry to handle transient vLLM proxy and network errors, along with corresponding unit tests. The review feedback suggests optimizing the retry mechanism to prevent a database read storm when multiple concurrent requests attempt to force-refresh a stale proxy URL simultaneously.
|
Hey Neil, thanks so much for putting your PR up: https://github.com/NovaSky-AI/SkyRL/pull/1994/changes! If I understand it correctly (🙏) I think my last PR (keeping multi-tenant LoRA runtime warm on last unload: #2001) might eliminate the need for proxy retries because the proxy should no longer change even when the number of active tenants hit zero. Let me know if you think differently! |
Testing
Verified 3 retry tests pass and all configured hooks pass.