Skip to content

feat(inference): compare sticky least-loaded routing - #3553

Closed
mikasenghaas wants to merge 7 commits into
mainfrom
feat/sticky-least-loaded
Closed

mikasenghaas wants to merge 7 commits into
mainfrom
feat/sticky-least-loaded

Conversation

@mikasenghaas

@mikasenghaas mikasenghaas commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

Summary

  • pin vllm-router to an experiment-only, checked-in Linux x86_64 wheel built from the exact head commit of feat(router): add sticky least-loaded session routing vllm-project/router#235, which adds sticky_least_loaded
  • bump verifiers to current main (4cd87df5) for the latest bash-harness streaming and retry fixes
  • add matched 200-step, four-node GLM-4.5-Air SWE configs using 1 trainer + 3 one-node inference replicas; model, training, evaluation, deployment, and routing-replay settings are identical, while job and sandbox identifiers keep the two arms isolated
  • retain the single-trainer OOM-avoidance recipe: stateless sign-SGD, full CPU offload, 1,024-token LM-head chunking, activation checkpointing and offloading, and no optimizer checkpoint state
  • isolate Triton, vLLM, and FlashInfer caches by user and SLURM job on node-local storage
  • expand configured inference environment variables before starting in-process workers, preserving the concrete per-job paths from the launcher
  • disable the optional fused allreduce/RMS compilation pass in both experiment arms to avoid concurrent FlashInfer JIT compilation during startup
  • log both runs to the vllm-router-routing W&B project as consistent-hash and sticky-least-loaded
  • release completed rollout sessions through the router /finish_session endpoint, automatically enabled for sticky least-loaded routing
  • start both comparison arms from scratch; the experiment configs do not enable checkpoint resume

Session lifecycle

The pinned verifiers clients send each rollout trace UUID as X-Session-ID on every train and eval generation request. PrimeRL launches vLLM Router with --request-id-headers x-session-id, so router#235 can distinguish a new session from later turns in the same session.

This change additionally POSTs each completed trace ID to /finish_session. Cleanup is best-effort and bounded so router cleanup cannot discard a completed episode; cancelled or failed environment requests without returned trace IDs fall back to the router TTL.

Router pin

This experiment does not require an upstream or Prime router release. The repository vendors an 8.3 MB CPython abi3 Linux x86_64 wheel built directly from router#235 commit 6422afa, and uv installs it from deps/wheels. The wheel SHA-256 is 7c3047e87ff5a0a02f3dfadd6f8012ceb85aabd8f0c2a5fe25ba88ee61b19f83.

The pin temporarily moves from PrimeIntellect-ai/router v0.2.0 wheels to the upstream PR package version 0.1.15; this experiment uses the regular, non-P/D router path.

Validation

  • uv lock --check
  • full frozen uv sync --all-extras --all-packages --dry-run resolves the local wheel
  • installed the wheel into an isolated target and verified the router CLI exposes sticky_least_loaded and --request-id-headers
  • uv run ruff format ... and uv run ruff check ... on changed Python and test files
  • uv run --no-sync pytest tests/unit/orchestrator/test_clients.py tests/unit/test_configs.py::test_sticky_router_auto_enables_session_release (12 passed)
  • both corrected configs pass fresh-run uv run --no-sync rl @ <config> --dry-run
  • resolved plans use 200 steps, one trainer node, three one-node inference replicas, the memory-safe trainer settings, their intended router policy, --request-id-headers x-session-id, and finish_sessions=false for consistent hash / true for sticky least-loaded
  • verified concrete per-job cache expansion and staggered startup for both comparison arms

@mikasenghaas mikasenghaas left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bump vf to latest main to grab the latest fix to bash harness

Comment thread configs/experiments/vllm-router-routing/consistent-hash.toml Outdated
Comment thread configs/experiments/vllm-router-routing/consistent-hash.toml Outdated
Comment thread configs/experiments/vllm-router-routing/consistent-hash.toml Outdated
@mikasenghaas

Copy link
Copy Markdown
Member Author

Addressed the review in caceead: bumped verifiers to current main (4cd87df5), including the latest bash-harness fixes. The lock now carries prime-runs 0.1.2, both configs dry-run successfully, and the focused suite passes (12 tests).

mikasenghaas added a commit that referenced this pull request Sep 17, 2026
## Summary

Makes `sticky_least_loaded` the default routing algorithm on
`vllm-router` (ported from upstream into our fork in
PrimeIntellect-ai/router#55). New session are
routed to the engine with _least active sessions_, any requests in an
active sessions are sticky-routed to maximize prefix cache hits. So this
is a mix of least loaded and consistent hashing which should be strictly
better for RL workloads.

- pin vllm-router v0.2.1 release wheels for x86_64 and aarch64
- make sticky_least_loaded the default vllm-router policy while
preserving explicit policy overrides
- release rollout sessions on success, failure, and cancellation for
managed routed inference, retaining discarded retry trace IDs
- call /finish_session without capability probing or 404/405 caching;
log individual failures at debug level and continue future cleanup
- document the new default and lifecycle behavior

## Context

- Router implementation:
PrimeIntellect-ai/router#55
- Router release:
https://github.com/PrimeIntellect-ai/router/releases/tag/v0.2.1
- Side-by-side experiment:
#3553

## Verification

We ran matched 100-step GLM-4.5-Air ScaleSWE jobs on six H200 nodes (2
trainer + 4 TP=8 inference replicas), changing only the vllm-router
policy: [sticky
least-loaded](https://wandb.ai/primeintellect/routing/runs/b44ce2f5354a416bb37788a4db9d12c3)
versus [consistent
hashing](https://wandb.ai/primeintellect/routing/runs/de4b5c72f95c4c5ebbc7e9614690af52).
They reached trainer step 100 in 8h46m and 8h44m respectively, with
nearly identical mean step time (309s vs. 308s).

<img width="1239" height="762" alt="Screenshot 2026-09-17 at 10 11
13 AM"
src="https://github.com/user-attachments/assets/a3204c40-49b0-41f0-b1ae-9d85391f893d"
/>

| Metric | Sticky least-loaded | Consistent hashing | Difference |
| --- | ---: | ---: | ---: |
| Decode throughput | 13.85k tokens/s | 13.11k tokens/s | +5.6% |
| Prefill throughput | 702.6k tokens/s | 665.0k tokens/s | +5.7% |
| Decode-token imbalance (CV) | 1.2% | 6.8% | 82% lower |
| Running-request imbalance (CV) | 5.1% | 23.9% | 79% lower |
| KV-usage imbalance (CV) | 8.2% | 25.1% | 67% lower |


Sticky processed about 6% more tokens in essentially the same wall time.
The higher useful throughput and markedly lower engine imbalance,
without an end-to-end regression, make it the better default.

## Validation

- uv sync --all-extras --all-packages
- uv lock --check
- uv run pytest tests/unit/orchestrator tests/unit/test_configs.py -q
(276 passed, 1 skipped)
- uv run ruff check .
- verified the installed vllm-router CLI exposes sticky_least_loaded
- 10 focused local cleanup checks: successful and failed episodes,
cancellation, overload/stale/superseded drops, shutdown, max steps,
cancellation during cleanup, and continued cleanup after a 404

<!-- CURSOR_SUMMARY -->
---

> [!NOTE]
> **Medium Risk**
> Changes default multi-replica routing and adds best-effort session
cleanup on the rollout hot path; failures are logged but could leave
stale router session state on older routers.
> 
> **Overview**
> Switches the default **vllm-router** policy from `consistent_hash` to
**`sticky_least_loaded`**, pins **vllm-router v0.2.1**, and documents
that new `X-Session-ID` sessions land on the least-loaded replica while
later turns stay sticky for KV reuse.
> 
> For managed deployments where `admin_base_url` bypasses the router for
engine admin, the orchestrator now **releases router sessions when
rollouts finish**: after each successful env episode it calls
**`/finish_session`** per trace id via the client-facing router URL,
probes support once, and disables cleanup on 404/405 so external routers
are unaffected. Explicit `[inference.router] policy` overrides still
work.
> 
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit
571b095. Bugbot is set up for automated
code reviews on this repo. Configure
[here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant