feat: add sticky least-loaded session routing - #55
Conversation
4be8b9e to
f9c96d7
Compare
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — This PR adds a substantial stateful session-routing policy, session-release API, PD routing integration, and production container package upgrades rather than a small isolated change. Its session lifecycle and cross-scope routing behavior require human review, including an unresolved identifier-collision risk. Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more. |
Balance new sessions by active-session count while preserving existing affinity, with explicit release and idle expiry. Preserve legacy session identifiers and wire the policy through regular and prefill/decode routing. Squashes vllm-project#235 commits a06430d and 6422afa. The implementation was adapted upstream from SumanthRH/router commit b50d926. Signed-off-by: bvolpato <brunocvcunha@gmail.com>
f9c96d7 to
9704756
Compare
9704756 to
fc9ad53
Compare
fc9ad53 to
3819532
Compare
554d57e to
8cdb9f0
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 8cdb9f0. Configure here.
Pass the fork-specific optional run_id argument in the upstream prefill/decode regression test. Upgrade Debian runtime packages so the container image includes current security fixes and passes the Trivy scan. Bound sticky session state with graceful stateless fallback and capacity-triggered expiry, balance stateless requests across tied replicas, keep per-request headers from masking stable session IDs, resolve model-specific routing requirements before serializing typed requests, and preserve supported legacy session fields.
8cdb9f0 to
f64492f
Compare
## Summary Makes `sticky_least_loaded` the default routing algorithm on `vllm-router` (ported from upstream into our fork in PrimeIntellect-ai/router#55). New session are routed to the engine with _least active sessions_, any requests in an active sessions are sticky-routed to maximize prefix cache hits. So this is a mix of least loaded and consistent hashing which should be strictly better for RL workloads. - pin vllm-router v0.2.1 release wheels for x86_64 and aarch64 - make sticky_least_loaded the default vllm-router policy while preserving explicit policy overrides - release rollout sessions on success, failure, and cancellation for managed routed inference, retaining discarded retry trace IDs - call /finish_session without capability probing or 404/405 caching; log individual failures at debug level and continue future cleanup - document the new default and lifecycle behavior ## Context - Router implementation: PrimeIntellect-ai/router#55 - Router release: https://github.com/PrimeIntellect-ai/router/releases/tag/v0.2.1 - Side-by-side experiment: #3553 ## Verification We ran matched 100-step GLM-4.5-Air ScaleSWE jobs on six H200 nodes (2 trainer + 4 TP=8 inference replicas), changing only the vllm-router policy: [sticky least-loaded](https://wandb.ai/primeintellect/routing/runs/b44ce2f5354a416bb37788a4db9d12c3) versus [consistent hashing](https://wandb.ai/primeintellect/routing/runs/de4b5c72f95c4c5ebbc7e9614690af52). They reached trainer step 100 in 8h46m and 8h44m respectively, with nearly identical mean step time (309s vs. 308s). <img width="1239" height="762" alt="Screenshot 2026-09-17 at 10 11 13 AM" src="https://github.com/user-attachments/assets/a3204c40-49b0-41f0-b1ae-9d85391f893d" /> | Metric | Sticky least-loaded | Consistent hashing | Difference | | --- | ---: | ---: | ---: | | Decode throughput | 13.85k tokens/s | 13.11k tokens/s | +5.6% | | Prefill throughput | 702.6k tokens/s | 665.0k tokens/s | +5.7% | | Decode-token imbalance (CV) | 1.2% | 6.8% | 82% lower | | Running-request imbalance (CV) | 5.1% | 23.9% | 79% lower | | KV-usage imbalance (CV) | 8.2% | 25.1% | 67% lower | Sticky processed about 6% more tokens in essentially the same wall time. The higher useful throughput and markedly lower engine imbalance, without an end-to-end regression, make it the better default. ## Validation - uv sync --all-extras --all-packages - uv lock --check - uv run pytest tests/unit/orchestrator tests/unit/test_configs.py -q (276 passed, 1 skipped) - uv run ruff check . - verified the installed vllm-router CLI exposes sticky_least_loaded - 10 focused local cleanup checks: successful and failed episodes, cancellation, overload/stale/superseded drops, shutdown, max steps, cancellation during cleanup, and continued cleanup after a 404 <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes default multi-replica routing and adds best-effort session cleanup on the rollout hot path; failures are logged but could leave stale router session state on older routers. > > **Overview** > Switches the default **vllm-router** policy from `consistent_hash` to **`sticky_least_loaded`**, pins **vllm-router v0.2.1**, and documents that new `X-Session-ID` sessions land on the least-loaded replica while later turns stay sticky for KV reuse. > > For managed deployments where `admin_base_url` bypasses the router for engine admin, the orchestrator now **releases router sessions when rollouts finish**: after each successful env episode it calls **`/finish_session`** per trace id via the client-facing router URL, probes support once, and disables cleanup on 404/405 so external routers are unaffected. Explicit `[inference.router] policy` overrides still work. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 571b095. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->

Summary
sticky_least_loadedrouting policy and explicit/finish_sessionlifecycle endpointrun_idargumentMotivation: Better load-balancing across engines compared to previous default
consistent_hashing(green)Note: The drop on the purple curve was a node failure and job restart, and not caused by the routing scheme.
Validation
cargo +1.95.0 fmt --checkcargo +1.95.0 clippy --all-targets -- -D warningscargo +1.95.0 testuv run --isolated --extra dev pytest py_test/unit -q(105 passed, 4 skipped)GitHub Actions: all checks pass, including image build and Trivy scan
Context
prime-rl#3553 compares consistent hashing with sticky least-loaded routing. The experiment needs published manylinux wheels from this fork so cluster
uv sync --all-extras --all-packagesdoes not compile the Rust extension on compute nodes.Upstream: vllm-project#235
Note
Add
StickyLeastLoadedPolicyfor session routingStickyLeastLoadedPolicyto route requests to the least-loaded replica using a session identifier, preserving affinity until sessions expire or workers become unhealthy.extract_session_idinsrc/policies/hash_key.rsto read session IDs from specific headers and request bodies, and updates the HTTP router to pass full request bodies and headers to policies that require them.POST /finish_sessionendpoint insrc/server.rsto release session assignments, and extendsLoadBalancingPolicywithfinish_sessionandneeds_request_bodydefault methods.PolicyFactory,PolicyRegistry, config enums, CLI arguments, and Python bindings to accept thesticky_least_loadedpolicy.Router::route_typed_requestinsrc/routers/http/router.rsnow serializes the full request for body-dependent policies; serialization failures return HTTP 500.Dockerfile.routernow upgrades Debian packages and installsca-certificateswithout recommendations.Macroscope summarized f64492f.
Note
Medium Risk
Changes core load-balancing and routing input selection (full-body serialization, mutex-backed session state); misconfigured session IDs or missing
finish_sessioncan skew load until TTL, and multi-router deployments need coordinated ingress and release.Overview
Introduces
sticky_least_loaded, a session-aware policy that pins ongoing sessions to one worker while assigning new sessions to the replica with the fewest active sessions (ties broken with rendezvous hashing). Session IDs come from the same headers as consistent hashing plus JSON fields (session_params.session_id,user,session_id,user_id); requests without an ID still load-balance but are not tracked.Adds
POST /finish_session?session_id=...(auth-gated) and registry fan-out so default, per-model, and PD prefill/decode policy instances can release reservations. Idle sessions expire (default 2h, env-configurable) with bounded state (max sessions / ID length, stateless fallback when limits are hit).Routing plumbing changes: policies can require the full serialized typed body (not just prompt text), PD selection passes headers into policies, and request schemas gain optional legacy session fields. CLI, config, Python bindings, docs, tests, and the runtime Dockerfile (
apt upgrade) are updated accordingly.Reviewed by Cursor Bugbot for commit 8cdb9f0. Bugbot is set up for automated code reviews on this repo. Configure here.