Conversation
ChatCompletionRequest::extract_text_for_routing returned session_params. session_id, or an empty string when the request had none. cache_aware feeds that value into its radix tree, so every chat request looked identical to every other one: the tree could never match a prefix, the match rate stayed at zero, and the policy fell through to its minimum-load branch. Chat completions is the most used endpoint on this router, so in practice cache_aware was plain load balancing there. Return the conversation text instead, matching the CompletionRequest and ResponsesRequest implementations of the same trait method. Session affinity is unaffected: hash policies resolve their key through hash_key::extract_hash_key, which reads request headers first. Measured on a 3-node replay of recorded Codex agent traces (128 sessions, 5280 requests, ~100K-token contexts) with --policy cache_aware. The router's own debug counters, per turn of one session: before matched_chars=0 input_chars=0 after matched_chars=0 input_chars=5100 after matched_chars=5100 input_chars=6702 after matched_chars=6702 input_chars=8304 Each turn now matches the whole of the previous turn, which is the prefix reuse cache_aware exists to exploit. A 20-session consistent_hash run over the same workload kept every session on one worker, unchanged. AI assistance was used for this change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: aab30fab32
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // policy silently degrades to min-load. Session affinity is a separate | ||
| // concern, resolved from headers and the body by `hash_key::extract_hash_key`. | ||
| let mut buffer = String::new(); |
There was a problem hiding this comment.
Preserve body session IDs for hash-based routing
When a chat client uses the documented session_params.session_id body field without a session header, route_typed_request now passes only this joined conversation to ConsistentHashPolicy and RendezvousHashPolicy. Consequently, hash_key::extract_hash_key cannot see the session ID, and a growing conversation produces a different fallback hash on each turn, potentially sending continual-prompting requests to different backends. The previous implementation supplied a stable session ID, so the conversation text and body routing metadata need to be exposed separately rather than discarding the latter.
Useful? React with 👍 / 👎.
Problem
ChatCompletionRequest::extract_text_for_routingreturnssession_params.session_id, or an empty string when the request carries none:https://github.com/vllm-project/router/blob/main/src/protocols/spec.rs#L544-L558
cache_awarefeeds that value into its radix tree. With an empty string every chat request looks identical to every other one, so the tree never matches a prefix,match_ratestays at 0, and selection falls through to the minimum-load branch. Chat completions is the most used endpoint here, socache_awareis effectively plain load balancing on it.The sibling implementations of the same trait method already return real content —
CompletionRequestreturns the prompt,ResponsesRequestwalks the input items. Chat is the odd one out.Fix
Return the conversation text: system, user (including text parts of multimodal content), and assistant messages joined in order.
Session affinity is unaffected. Hash policies resolve their key through
hash_key::extract_hash_key, which checksx-session-idand the other session headers first.Behaviour change
A caller that sends
session_params.session_idin the body and no session header previously got a stable routing key, because the session id happened to be the routing text and fell through to therequest:{text}fallback. That key now becomes a hash of the conversation, which grows every turn, so such a caller loses session affinity until it sends a session header. Callers using headers — the documented path — are unaffected.Testing
cargo test --lib— 485 passed, 0 failed. Two tests added alongside the existingGenerationRequesttrait tests: one asserts a continued conversation extends the previous turn's routing text, one assertssession_paramsno longer displaces the conversation.End to end on 3 nodes, replaying recorded Codex agent traces (128 sessions, 5280 requests, contexts to ~100K tokens) against 4 vLLM replicas with
--policy cache_aware. The router's own debug counters for consecutive turns of one session:matched_chars=0 input_chars=0matched_chars=0 input_chars=5100matched_chars=0 input_chars=0matched_chars=5100 input_chars=6702matched_chars=0 input_chars=0matched_chars=6702 input_chars=8304Each turn now matches the entirety of the turn before it, which is the prefix reuse
cache_awareexists to exploit. A 20-sessionconsistent_hashrun over the same workload kept every session pinned to one worker, unchanged by this patch.Duplicate check
gh pr list --search cache_awareand an issue search for the routing-text behaviour found nothing covering this. The open cache-aware PRs are #211 (load double-decrement), #176 (PD stream load tracking) and #130 (KV-event routing), all unrelated.AI assistance was used for this change; the diff and the test results above were reviewed and reproduced by the submitter.
🤖 Generated with Claude Code