Skip to content

feat: add sticky least-loaded session routing - #55

Merged
mikasenghaas merged 2 commits into
mainfrom
feat/sticky-least-loaded
Sep 16, 2026
Merged

mikasenghaas merged 2 commits into
mainfrom
feat/sticky-least-loaded

Conversation

@mikasenghaas

@mikasenghaas mikasenghaas commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

Summary

  • squash the two commits from feat(router): add sticky least-loaded session routing vllm-project/router#235 into one upstream-authored feature commit
  • keep fork-specific compatibility and image fixes in a separate integration commit
  • add the sticky_least_loaded routing policy and explicit /finish_session lifecycle endpoint
  • preserve legacy session identifiers while accepting configured request-ID headers
  • include the upstream policy, API, parser, and routing coverage
  • adapt the discovered P/D test to the fork-specific optional run_id argument
  • upgrade runtime Debian packages to pick up current security fixes
  • bound sticky session state and session-ID size with stateless fallback, sweeping expired entries before degrading affinity
  • resolve the effective model policy before choosing typed-body routing input
  • preserve supported legacy session fields across typed request schemas
  • keep per-request headers available to consistent hashing without letting them mask sticky session IDs
  • round-robin stateless requests across equally least-loaded workers

Motivation: Better load-balancing across engines compared to previous default consistent_hashing (green)

Screenshot 2026-09-15 at 6 18 24 PM

Note: The drop on the purple curve was a node failure and job restart, and not caused by the routing scheme.

Validation

  • cargo +1.95.0 fmt --check

  • cargo +1.95.0 clippy --all-targets -- -D warnings

  • cargo +1.95.0 test

  • uv run --isolated --extra dev pytest py_test/unit -q (105 passed, 4 skipped)

  • GitHub Actions: all checks pass, including image build and Trivy scan

Context

prime-rl#3553 compares consistent hashing with sticky least-loaded routing. The experiment needs published manylinux wheels from this fork so cluster uv sync --all-extras --all-packages does not compile the Rust extension on compute nodes.

Upstream: vllm-project#235

Note

Add StickyLeastLoadedPolicy for session routing

  • Introduces StickyLeastLoadedPolicy to route requests to the least-loaded replica using a session identifier, preserving affinity until sessions expire or workers become unhealthy.
  • Adds extract_session_id in src/policies/hash_key.rs to read session IDs from specific headers and request bodies, and updates the HTTP router to pass full request bodies and headers to policies that require them.
  • Adds a POST /finish_session endpoint in src/server.rs to release session assignments, and extends LoadBalancingPolicy with finish_session and needs_request_body default methods.
  • Updates PolicyFactory, PolicyRegistry, config enums, CLI arguments, and Python bindings to accept the sticky_least_loaded policy.
  • Behavioral Change: Router::route_typed_request in src/routers/http/router.rs now serializes the full request for body-dependent policies; serialization failures return HTTP 500. Dockerfile.router now upgrades Debian packages and installs ca-certificates without recommendations.

Macroscope summarized f64492f.


Note

Medium Risk
Changes core load-balancing and routing input selection (full-body serialization, mutex-backed session state); misconfigured session IDs or missing finish_session can skew load until TTL, and multi-router deployments need coordinated ingress and release.

Overview
Introduces sticky_least_loaded, a session-aware policy that pins ongoing sessions to one worker while assigning new sessions to the replica with the fewest active sessions (ties broken with rendezvous hashing). Session IDs come from the same headers as consistent hashing plus JSON fields (session_params.session_id, user, session_id, user_id); requests without an ID still load-balance but are not tracked.

Adds POST /finish_session?session_id=... (auth-gated) and registry fan-out so default, per-model, and PD prefill/decode policy instances can release reservations. Idle sessions expire (default 2h, env-configurable) with bounded state (max sessions / ID length, stateless fallback when limits are hit).

Routing plumbing changes: policies can require the full serialized typed body (not just prompt text), PD selection passes headers into policies, and request schemas gain optional legacy session fields. CLI, config, Python bindings, docs, tests, and the runtime Dockerfile (apt upgrade) are updated accordingly.

Reviewed by Cursor Bugbot for commit 8cdb9f0. Bugbot is set up for automated code reviews on this repo. Configure here.

@mikasenghaas
mikasenghaas force-pushed the feat/sticky-least-loaded branch 2 times, most recently from 4be8b9e to f9c96d7 Compare September 16, 2026 04:23
@mikasenghaas
mikasenghaas marked this pull request as ready for review September 16, 2026 04:40
Comment thread src/policies/sticky_least_loaded.rs
Comment thread src/policies/registry.rs
Comment thread src/policies/sticky_least_loaded.rs
Comment thread src/policies/sticky_least_loaded.rs
Comment thread src/protocols/spec.rs

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread src/routers/http/router.rs
@macroscopeapp

macroscopeapp Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR adds a substantial stateful session-routing policy, session-release API, PD routing integration, and production container package upgrades rather than a small isolated change. Its session lifecycle and cross-scope routing behavior require human review, including an unresolved identifier-collision risk.

Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more.

Balance new sessions by active-session count while preserving existing affinity, with explicit release and idle expiry. Preserve legacy session identifiers and wire the policy through regular and prefill/decode routing.

Squashes vllm-project#235 commits a06430d and 6422afa.

The implementation was adapted upstream from SumanthRH/router commit b50d926.

Signed-off-by: bvolpato <brunocvcunha@gmail.com>
@mikasenghaas
mikasenghaas force-pushed the feat/sticky-least-loaded branch from f9c96d7 to 9704756 Compare September 16, 2026 04:47
samsja
samsja previously approved these changes Sep 16, 2026
Comment thread src/policies/sticky_least_loaded.rs Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread src/policies/sticky_least_loaded.rs
@mikasenghaas
mikasenghaas force-pushed the feat/sticky-least-loaded branch from fc9ad53 to 3819532 Compare September 16, 2026 19:26
Comment thread src/policies/sticky_least_loaded.rs
@mikasenghaas
mikasenghaas force-pushed the feat/sticky-least-loaded branch 2 times, most recently from 554d57e to 8cdb9f0 Compare September 16, 2026 19:40

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 8cdb9f0. Configure here.

Comment thread src/policies/hash_key.rs
Pass the fork-specific optional run_id argument in the upstream prefill/decode regression test.

Upgrade Debian runtime packages so the container image includes current security fixes and passes the Trivy scan.

Bound sticky session state with graceful stateless fallback and capacity-triggered expiry, balance stateless requests across tied replicas, keep per-request headers from masking stable session IDs, resolve model-specific routing requirements before serializing typed requests, and preserve supported legacy session fields.
@mikasenghaas
mikasenghaas force-pushed the feat/sticky-least-loaded branch from 8cdb9f0 to f64492f Compare September 16, 2026 20:04
Comment thread src/policies/hash_key.rs
@mikasenghaas
mikasenghaas merged commit f9e6f35 into main Sep 16, 2026
9 checks passed
mikasenghaas added a commit to PrimeIntellect-ai/prime-rl that referenced this pull request Sep 17, 2026
## Summary

Makes `sticky_least_loaded` the default routing algorithm on
`vllm-router` (ported from upstream into our fork in
PrimeIntellect-ai/router#55). New session are
routed to the engine with _least active sessions_, any requests in an
active sessions are sticky-routed to maximize prefix cache hits. So this
is a mix of least loaded and consistent hashing which should be strictly
better for RL workloads.

- pin vllm-router v0.2.1 release wheels for x86_64 and aarch64
- make sticky_least_loaded the default vllm-router policy while
preserving explicit policy overrides
- release rollout sessions on success, failure, and cancellation for
managed routed inference, retaining discarded retry trace IDs
- call /finish_session without capability probing or 404/405 caching;
log individual failures at debug level and continue future cleanup
- document the new default and lifecycle behavior

## Context

- Router implementation:
PrimeIntellect-ai/router#55
- Router release:
https://github.com/PrimeIntellect-ai/router/releases/tag/v0.2.1
- Side-by-side experiment:
#3553

## Verification

We ran matched 100-step GLM-4.5-Air ScaleSWE jobs on six H200 nodes (2
trainer + 4 TP=8 inference replicas), changing only the vllm-router
policy: [sticky
least-loaded](https://wandb.ai/primeintellect/routing/runs/b44ce2f5354a416bb37788a4db9d12c3)
versus [consistent
hashing](https://wandb.ai/primeintellect/routing/runs/de4b5c72f95c4c5ebbc7e9614690af52).
They reached trainer step 100 in 8h46m and 8h44m respectively, with
nearly identical mean step time (309s vs. 308s).

<img width="1239" height="762" alt="Screenshot 2026-09-17 at 10 11
13 AM"
src="https://github.com/user-attachments/assets/a3204c40-49b0-41f0-b1ae-9d85391f893d"
/>

| Metric | Sticky least-loaded | Consistent hashing | Difference |
| --- | ---: | ---: | ---: |
| Decode throughput | 13.85k tokens/s | 13.11k tokens/s | +5.6% |
| Prefill throughput | 702.6k tokens/s | 665.0k tokens/s | +5.7% |
| Decode-token imbalance (CV) | 1.2% | 6.8% | 82% lower |
| Running-request imbalance (CV) | 5.1% | 23.9% | 79% lower |
| KV-usage imbalance (CV) | 8.2% | 25.1% | 67% lower |


Sticky processed about 6% more tokens in essentially the same wall time.
The higher useful throughput and markedly lower engine imbalance,
without an end-to-end regression, make it the better default.

## Validation

- uv sync --all-extras --all-packages
- uv lock --check
- uv run pytest tests/unit/orchestrator tests/unit/test_configs.py -q
(276 passed, 1 skipped)
- uv run ruff check .
- verified the installed vllm-router CLI exposes sticky_least_loaded
- 10 focused local cleanup checks: successful and failed episodes,
cancellation, overload/stale/superseded drops, shutdown, max steps,
cancellation during cleanup, and continued cleanup after a 404

<!-- CURSOR_SUMMARY -->
---

> [!NOTE]
> **Medium Risk**
> Changes default multi-replica routing and adds best-effort session
cleanup on the rollout hot path; failures are logged but could leave
stale router session state on older routers.
> 
> **Overview**
> Switches the default **vllm-router** policy from `consistent_hash` to
**`sticky_least_loaded`**, pins **vllm-router v0.2.1**, and documents
that new `X-Session-ID` sessions land on the least-loaded replica while
later turns stay sticky for KV reuse.
> 
> For managed deployments where `admin_base_url` bypasses the router for
engine admin, the orchestrator now **releases router sessions when
rollouts finish**: after each successful env episode it calls
**`/finish_session`** per trace id via the client-facing router URL,
probes support once, and disables cleanup on 404/405 so external routers
are unaffected. Explicit `[inference.router] policy` overrides still
work.
> 
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit
571b095. Bugbot is set up for automated
code reviews on this repo. Configure
[here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants