feat(inference): compare sticky least-loaded routing - #3553
Closed
mikasenghaas wants to merge 7 commits into
Closed
mikasenghaas wants to merge 7 commits into
mikasenghaas wants to merge 7 commits into
Conversation
mikasenghaas
commented
Sep 15, 2026
mikasenghaas
left a comment
Member
Author
There was a problem hiding this comment.
bump vf to latest main to grab the latest fix to bash harness
Member
Author
|
Addressed the review in caceead: bumped verifiers to current main (4cd87df5), including the latest bash-harness fixes. The lock now carries prime-runs 0.1.2, both configs dry-run successfully, and the focused suite passes (12 tests). |
mikasenghaas
added a commit
that referenced
this pull request
Sep 17, 2026
## Summary Makes `sticky_least_loaded` the default routing algorithm on `vllm-router` (ported from upstream into our fork in PrimeIntellect-ai/router#55). New session are routed to the engine with _least active sessions_, any requests in an active sessions are sticky-routed to maximize prefix cache hits. So this is a mix of least loaded and consistent hashing which should be strictly better for RL workloads. - pin vllm-router v0.2.1 release wheels for x86_64 and aarch64 - make sticky_least_loaded the default vllm-router policy while preserving explicit policy overrides - release rollout sessions on success, failure, and cancellation for managed routed inference, retaining discarded retry trace IDs - call /finish_session without capability probing or 404/405 caching; log individual failures at debug level and continue future cleanup - document the new default and lifecycle behavior ## Context - Router implementation: PrimeIntellect-ai/router#55 - Router release: https://github.com/PrimeIntellect-ai/router/releases/tag/v0.2.1 - Side-by-side experiment: #3553 ## Verification We ran matched 100-step GLM-4.5-Air ScaleSWE jobs on six H200 nodes (2 trainer + 4 TP=8 inference replicas), changing only the vllm-router policy: [sticky least-loaded](https://wandb.ai/primeintellect/routing/runs/b44ce2f5354a416bb37788a4db9d12c3) versus [consistent hashing](https://wandb.ai/primeintellect/routing/runs/de4b5c72f95c4c5ebbc7e9614690af52). They reached trainer step 100 in 8h46m and 8h44m respectively, with nearly identical mean step time (309s vs. 308s). <img width="1239" height="762" alt="Screenshot 2026-09-17 at 10 11 13 AM" src="https://github.com/user-attachments/assets/a3204c40-49b0-41f0-b1ae-9d85391f893d" /> | Metric | Sticky least-loaded | Consistent hashing | Difference | | --- | ---: | ---: | ---: | | Decode throughput | 13.85k tokens/s | 13.11k tokens/s | +5.6% | | Prefill throughput | 702.6k tokens/s | 665.0k tokens/s | +5.7% | | Decode-token imbalance (CV) | 1.2% | 6.8% | 82% lower | | Running-request imbalance (CV) | 5.1% | 23.9% | 79% lower | | KV-usage imbalance (CV) | 8.2% | 25.1% | 67% lower | Sticky processed about 6% more tokens in essentially the same wall time. The higher useful throughput and markedly lower engine imbalance, without an end-to-end regression, make it the better default. ## Validation - uv sync --all-extras --all-packages - uv lock --check - uv run pytest tests/unit/orchestrator tests/unit/test_configs.py -q (276 passed, 1 skipped) - uv run ruff check . - verified the installed vllm-router CLI exposes sticky_least_loaded - 10 focused local cleanup checks: successful and failed episodes, cancellation, overload/stale/superseded drops, shutdown, max steps, cancellation during cleanup, and continued cleanup after a 404 <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes default multi-replica routing and adds best-effort session cleanup on the rollout hot path; failures are logged but could leave stale router session state on older routers. > > **Overview** > Switches the default **vllm-router** policy from `consistent_hash` to **`sticky_least_loaded`**, pins **vllm-router v0.2.1**, and documents that new `X-Session-ID` sessions land on the least-loaded replica while later turns stay sticky for KV reuse. > > For managed deployments where `admin_base_url` bypasses the router for engine admin, the orchestrator now **releases router sessions when rollouts finish**: after each successful env episode it calls **`/finish_session`** per trace id via the client-facing router URL, probes support once, and disables cleanup on 404/405 so external routers are unaffected. Explicit `[inference.router] policy` overrides still work. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 571b095. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
vllm-routerto an experiment-only, checked-in Linux x86_64 wheel built from the exact head commit of feat(router): add sticky least-loaded session routing vllm-project/router#235, which addssticky_least_loaded4cd87df5) for the latest bash-harness streaming and retry fixes1 trainer + 3 one-node inference replicas; model, training, evaluation, deployment, and routing-replay settings are identical, while job and sandbox identifiers keep the two arms isolatedvllm-router-routingW&B project asconsistent-hashandsticky-least-loaded/finish_sessionendpoint, automatically enabled for sticky least-loaded routingSession lifecycle
The pinned verifiers clients send each rollout trace UUID as
X-Session-IDon every train and eval generation request. PrimeRL launches vLLM Router with--request-id-headers x-session-id, so router#235 can distinguish a new session from later turns in the same session.This change additionally POSTs each completed trace ID to
/finish_session. Cleanup is best-effort and bounded so router cleanup cannot discard a completed episode; cancelled or failed environment requests without returned trace IDs fall back to the router TTL.Router pin
This experiment does not require an upstream or Prime router release. The repository vendors an 8.3 MB CPython abi3 Linux x86_64 wheel built directly from router#235 commit
6422afa, anduvinstalls it fromdeps/wheels. The wheel SHA-256 is7c3047e87ff5a0a02f3dfadd6f8012ceb85aabd8f0c2a5fe25ba88ee61b19f83.The pin temporarily moves from PrimeIntellect-ai/router v0.2.0 wheels to the upstream PR package version 0.1.15; this experiment uses the regular, non-P/D router path.
Validation
uv lock --checkuv sync --all-extras --all-packages --dry-runresolves the local wheelsticky_least_loadedand--request-id-headersuv run ruff format ...anduv run ruff check ...on changed Python and test filesuv run --no-sync pytest tests/unit/orchestrator/test_clients.py tests/unit/test_configs.py::test_sticky_router_auto_enables_session_release(12 passed)uv run --no-sync rl @ <config> --dry-run--request-id-headers x-session-id, andfinish_sessions=falsefor consistent hash /truefor sticky least-loaded