Skip to content

Single-node vLLM op-attribution profiling (NVIDIA + AMD): per-kernel op/module, clocks, MoE routing - #3639

Draft
hbarclay wants to merge 49 commits into
mainfrom
hbarclay/profiling
Draft

hbarclay wants to merge 49 commits into
mainfrom
hbarclay/profiling

Conversation

@hbarclay

@hbarclay hbarclay commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Opt-in profiling for single-node vLLM srt-slurm runs, NVIDIA and AMD. Set profile on e2e-tests.yml ("{}" takes the defaults). Without it nothing changes.

Each profiled run collects:

  • Per device activity (kernel, memcpy, memset): the launching op, op chain, input shapes and dtypes, module call stack, and the Python launcher with its vLLM caller frames (for Triton, TileLang, CuTe, DeepGEMM and FlashInfer).
    • Eager kernels join to their CPU launch by correlation id.
    • CUDA graph replays join by node id to the launches recorded when that graph was captured.
    • HIP graph replays carry no node id. Each replay stream keeps capture order, so they join through the interleaving of the per-stream sequences that best matches the kernel families each op ran eagerly.
    • CPU KV-offload memcpys (cuMemcpyBatchAsync, which Kineto does not record) join through a launch_copy log.
  • Per kernel: graphics/SM, memory and video clocks and throttle reasons over the kernel's lifetime. NVML (or amdsmi gpu_metrics on AMD) is polled every 250 µs per GPU; B200 reports clock changes on a 100 ms grid.
  • Per step: batch composition and cudagraph mode, plus MoE routing (expert token counts per layer, and per-token top-k ids for steps of ≤1024 tokens). Routing comes from a capturer the patch binds to each MoE layer before graph capture, so graph replays record it too.
  • Windows: prefill at the warmup start; decode once ≥90% of steps replay FULL graphs.

Two load modes:

  • "mode": "agentic" (default): the AgentX trace replay, ended with SIGINT to aiperf after the last window. 30–60 min per point.
  • "mode": "synthetic": random-token prompts of isl tokens (default 8192) at the point's concurrency; prefill-only requests until the prefill window is captured, then osl-token (default 4096) decodes. No KV-cache reuse, no trace download.

Artifacts per run:

  • profile_<result>: raw torch traces, capture trace and logs.
  • profile_steps_<result>: one JSON file per step, plus index.json and report.json.

Schema and method: inferencex-e2e/benchmarks/profiling/vllm/README.md.

DSV4-Pro FP4, B200 TP8 c4. One decode step is a single FULL graph replay; every layer-20 kernel is shown with its host op or launcher and its module:

decode step

Prefill step with per-kernel SM clock. The SW power cap engages at 243 ms, and SM steps 1965 → 1927 → 1942 → 1882 MHz:

prefill step

Validation. Every run below: 0 unattributed activities, clocks on every kernel, routing on every scheduled step with every MoE layer bound.

DSV4-Pro FP4, B200, DSpark, agentic:

point run prefill steps decode steps (FULL graph)
TP8 c1 36944666980 30 31
TP8 c4 36934474004 32 32
TP8 c16 36949870363 32 31
DEP8 c32 36934477240 32 30 on each of 5 busy ranks
DEP8 c96 36944669332 32 30
DEP8 c192 36958365816 / 36952776032 32 27

Other GPUs and models (op-or-launcher = share of device time with an op or launcher):

model GPU point mode run op-or-launcher
DSV4-Pro FP4 B300 TP8 c4 agentic 37025911392 100%
DSV4-Pro FP4 B300 DEP8 c128 agentic 37255757725 100%
DSV4-Pro FP4 B200 TP8 c4 synthetic 37025761564 100%
DSV4-Pro FP4 MI355X TP8 c4 agentic 37264449533 100%
MiniMax-M3 FP4 B200 TP4 c10 synthetic 37052301061 94.8%
MiniMax-M3 FP4 B300 TP4 c10 synthetic 37052286762 94.7%
MiniMax-M3 FP8 H200 TP8 c4 synthetic 37052329448 100%
DSV4.1-Flash FP4 B200 TP2, TP4 c16 synthetic 37052314988 100%
DSV4.1-Flash FP4 H200 TP4, TP8 c8 synthetic 37052344486 100%
DSV4.1-Flash FP4 GB200 TP2, TP4 c16 synthetic 37052392521 100%
DSV4.1-Flash FP4 GB300 TP2, TP4 c16 synthetic 37052406745 100%
Kimi-K3 FP4 MI355X TP8 c4 synthetic 37264462279 100%
DSV4.1-Flash FP4 MI300X TP4, TP8 c8 synthetic 37264474498 100%

Idle DP ranks run dummy forwards only, so they carry no routing.

Known issues:

  • DEP8 c192 (B200): a warmup window was followed by CUBLAS_STATUS_EXECUTION_FAILED in the attention compressor in 2 of 3 runs. Profile prefill and decode in separate runs.
  • FlashInfer JIT kernels (sparse top-k, plan) launch outside torch ops and the launcher hooks; they carry module paths only. This is the MiniMax gap.
  • AMD module paths: the DSV4 recipes compile at mode 3, which bypasses module hooks; ops and launchers still attribute.
  • AMD images: the pinned DSV4 MI355X and DSV4.1-Flash MI300X/MI325X nightly tags were pruned from Docker Hub; those rows ran on vllm/vllm-openai-rocm:v0.31.0 from a dispatch-only branch.
  • Not covered: Kimi-K3 on B300 (its Mooncake segments fail RDMA registration with ENOMEM even when shrunk to 217 GiB per rank), multi-node configs, SGLang (Qwen) and H100 (no runner available).
  • HIP graph join evidence: report.json reports the share of positions matched by kernel-family evidence (stream_order_evidence / stream_order_positions); the rest follow device start order.
  • Throughput: windows pause the engine for trace export, so a profiled run's throughput is not a result.
  • Clock staleness: NVML polls occasionally block for 10–55 ms. Each kernel carries prior_us, the age of its clock state.

🤖 Generated with Claude Code

hbarclay and others added 30 commits September 28, 2026 02:49
A `profile` input (JSON) on e2e-tests and benchmark-tmpl turns on
op-attribution profiling for single-node vLLM srt-slurm recipes:

- benchmarks/profiling/vllm/sitecustomize.py (loaded via PYTHONPATH, inert
  unless INFX_PROF_DIR is set) gives every CUDA graph an ordinal, marks its
  capture and replays, profiles the V2 runner's capture_model with shapes
  and stacks on the selected ranks, and logs every step's batch composition
  (per-request scheduled/computed/spec tokens, cudagraph dispatch).
- single_node.py adds the vLLM torch profiler config (shapes, stacks,
  bounded iterations), the patch env, and a run-local readable compile cache.
- srt_agentic.sh opens the configured torch windows on every vLLM server
  during the replay (profile_windows.py).
- The launcher tars /logs/infx_profile into its own profile_* artifact.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
extract.py reads an unpacked profile artifact and writes one row per device
activity: eager launches join through the CUDA correlation id; graph-replayed
nodes, ordered by graph node id, join the launches recorded while that graph
was captured. Rows carry the op, op chain, input shapes, kernel file,
module path, stream and device timing; per-step totals carry the logged
batch composition. report.json records the count and kernel-name checks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first DEP8 probe hung: 48 eager steps traced with Python stacks took
over ten minutes to export per rank, and a worker blocked past the RPC
timeout took the engine down. Replay windows now record ops and shapes
without stacks; eager modules mark themselves (infx_mod#<qualified name>)
through global module hooks registered only while a window is open and only
when nothing is compiled, and compiled pieces mark themselves through the
piecewise backend (infx_piece#<index>). Capture keeps full stacks: it runs
once at startup.

Step and graph logs are line-buffered (teardown kills the engines), windows
default to 32 iterations, and VLLM_RPC_TIMEOUT is raised for profiled runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…g as they need

A profiled AgentX point held its node for the full 20-minute fast replay to
collect two short windows, and its first window landed in the client's
warmup (dataset configuration and warmup precede measurement). Windows now
count from the moment the servers have completed the warmup's requests
(vllm:request_success_total), falling back to wall time if no server reports
it, and INFX_PROFILE_DURATION ends the replay 240 s after the last window
starts. It is set only for profiled runs; agentx-fast is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The second DEP8 probe wrote every rank's trace, then lost an engine core
while all eight workers post-processed their windows: DEP8 pins ~227 GB of
host offload per rank, and vLLM's default cuda-time summary table walks
every event in Python. The summary is off; the raw trace is unchanged.

Windows now start from aiperf's own measured-phase log line (vLLM's success
counter is not exported by this frontend), with the counter and wall time as
fallbacks. Profiled replays tolerate failed requests, since an export pauses
the engine. Idle DP ranks' dummy forwards, which keep EP collectives in
step, are marked (infx_dummy#k) and logged like scheduled steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
About two thirds of DEP8's device time is kernels launched straight from
Python (DeepGEMM Mega-MoE, TileLang, Triton) with no torch op and so no
recorded shapes. Each eager module marker now carries its tensor inputs'
shapes and dtypes, and the extractor reports them per kernel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… kernel

Kernels launched straight from Python have no torch op. The launch entry
points of the libraries DSV4 uses (Triton JITFunction.run, TileLang and
CuTe DSL kernel calls, vLLM's DeepGEMM and FlashInfer wrappers) mark
themselves while profiling with the launcher and the vLLM frames above it,
bypassing themselves under torch.compile. Each extracted kernel carries its
module stack (with input shapes), torch op chain and launcher with callers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
After a profiled run uploads its full profile (profile_<result>), the job
extracts it into profile_steps_<result>: window<w>/<rank>/<kind><k>.json.gz
holds one step's kernels, each with its module stack, op chain, launcher and
device timing, next to the step's logged batch composition; index.json lists
every step file with its summary and report.json the join checks. The step
is best effort, so the full profile stays the record.

The extractor now streams traces (the TP8 capture trace is 4 GB of JSON:
7.4 GB peak, 7 minutes) and checks kernel names once across all traces,
reporting graph-replayed kernels never seen eagerly instead of per-row flags
learned from one rank.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
vLLM's per-step execute_context_* annotation was reported as the op of the
kernels launched directly under it; user annotations now go to their own
field and op is the innermost torch op. DSV4's B200 recipes run with
compilation mode NONE, so the Inductor provenance and run-local compile
cache bought nothing and re-JITted DeepGEMM kernels on every run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
DEP8's recipe sizes its CPU KV-offload pool to nearly the whole host
(8 x 243 GB), and the node ran out of memory while its first profile window
was collecting. Profiled runs shrink the pool by host_headroom_gib (default
128 GiB, 16 GiB per rank); a short profiling run fills a small fraction of
the pool, and the override is part of the run's recorded config.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The top-level reorganization (#3525) put the benchmark job's tree under
inferencex-e2e/, which srt-slurm mounts as /infmax-workspace. The profiling
directory joins it, so the worker PYTHONPATH and the client's window script
resolve unchanged, and the workflow reads and uploads the profile from
INFERENCEX_E2E_ROOT.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
vLLM advances its profiler only for scheduled steps, so an idle DP rank's
window never reached max_iterations and recorded dummy forwards until the
window was stopped: DEP8's idle ranks wrote 250-280 MB traces against
~30 MB elsewhere. Dummy forwards now step the profiler, so every rank's
window covers the same 32 engine steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Kineto writes event names into the trace JSON unescaped, so module and
launcher markers carrying JSON made every trace unparseable; the streaming
parser then buffered each trace to its end, yielded nothing, and the CI step
reported an empty success after 37 minutes. Markers now use quote-free
encodings (infx_mod#<name>#7x7168:bfloat16;..., infx_py#<launcher>#<frame>|...),
and launcher callers skip the patch's own frames, whose container path
contains /vllm/. The parser re-escapes quotes in name lines, so traces from
the first runs still parse, raises on undecodable or truncated input, and
the extractor exits non-zero when no device activity is attributed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Markers are quote-free, so the extractor no longer re-escapes name lines or
parses the first runs' JSON marker names. A trace it cannot decode still
fails the extraction instead of yielding nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The window trigger now takes each vLLM server from SRT_AGG_ENDPOINTS, the
endpoint list srt-slurm gives custom benchmarks (the leader's public port
for direct vLLM, each pool's port behind vllm-router), instead of deriving
URLs from the metrics list; it falls back to the client's server URL.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On DEP8 c32 every rank ran eager mixed steps for the first six minutes after
warmup (the lanes' long first turns in chunked prefill); full-graph decode
steps took over from minute seven (~75-80% of steps from minute eight). With
windows at 60 s and 240 s both landed in the ramp, so the main model's decode
graphs never replayed in a window. The second window now opens at 540 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…st one

Windows were timed from the servers reaching CONC completed requests, but
AgentX warmup is larger (DEP8 c32: 34 mandatory primers + 32), so both windows
fired inside warmup and the measured phase then ran its full length for
nothing. A window is now [aiperf phase, delay, iterations], anchored on
aiperf logging that phase's start: by default one 60 s into warmup (the
lanes' long first turns, prefill-heavy) and one at the measured phase's
start (decode-dominated), and the measured replay lasts only 150 s past the
last measured window. The completed-request signal is gone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…indows

Under DP8+EP, vLLM agrees one CUDA graph mode across DP ranks, so a single
prefilling rank keeps every rank eager. Each measured-phase AgentX turn
re-prefills first: at DEP8 c32 no step replayed a FULL graph during warmup,
and FULL decode passed 90% of steps only ~130 s into the measured phase, so
a window at its start profiled eager steps and 150 s of measured replay was
not enough.

The decode window now anchors on engine state: the window client tails the
engines' step logs and opens it once 90% of the last 20 s of steps replayed
FULL graphs. The measured phase is capped at 600 s, and once the last
window closes (after at least 60 s of measured replay, so aiperf has results)
the client SIGINTs aiperf, a user cancel that exports and exits zero.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AgentX warmup sends every lane's long first turn at its start. At TP8 c4 on
B200 those prefills had drained by 60 s in: the window there held 2 prefill
steps and 28 decode steps. Opening it as warmup starts catches them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
aiperf services retitle themselves "aiperf <service_id>" with setproctitle,
so no process's argv still named bin/aiperf by the time the windows closed:
DEP8 c32 found none to signal and replayed the full 600 s cap. SIGINT goes to
"aiperf system_controller", the process whose handler cancels gracefully.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SimpleCPUOffloadConnector copies KV blocks with cuMemcpyBatchAsync from its
copy thread. Kineto records neither that driver call nor record_function
ranges on that thread, so those memcpys had no CPU launch: at DEP8 c96 they
were 259 unattributed device-to-pinned copies, 181 ms. The patch logs every
DmaCopyBackend.launch_copy (direction, blocks, bytes, step, vLLM callers);
each is one memcpy on its direction's stream, in queue order. The extractor
pairs them per direction by least total issue lag, bytes checked, after
putting the log's wall clock on the trace clock through the step markers. On
the DEP8 c32 traces a synthetic log with decoys before and after the window
pairs 32/32 correctly.

aiperf exports no server metrics after the window client's user cancel, and
the DSV4 recipe requires them (AIPERF_REQUIRED_SERVER_METRIC_PREFIX), which
failed DEP8 c96 after its profile was taken; profile runs clear it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
B200 profile points held their node 30 to 60 minutes. At DSV4 DEP8 c192 a
warmup window preceded an engine fault in both runs that had one, while the
decode-only run and unprofiled production runs completed; the README says to
profile such a point's prefill and decode separately.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The only clock record was the AgentX power monitor's 1 s nvidia-smi CSV: one
or two samples per profile window, most of them taken while the engine idles
through the trace export. The window client now polls NVML itself (graphics,
SM, memory and video clocks and the clock event reasons, every GPU) as fast
as NVML answers, from just before each window opens until every engine has
logged its iterations, into clocks/. The env record keeps each rank's GPU
UUID, and the extractor gives every kernel its GPU's clock ranges, event
reasons and sample count over its lifetime, on the trace clock through the
step markers, plus per-trace coverage in report.json.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…poll

In TP8 c1's first clock run, polling ran at ~32 kHz, but single NVML polls
blocked for 10 to 55 ms: 487 stalls, 8.6 s of the 45 s prefill window (16 in
decode). One thread per GPU keeps a blocked call on one GPU from holding up
the rest; each poll records when it began and returned, so the extractor
takes the last poll that returned before a kernel started as its state and
every overlapping poll for its range. Polls throttle to one per 250 us per
GPU (SM clock changes were tens of ms apart) and are kept in memory and
written when the window's sampling stops. Old single-stamp files still load.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On TP8 c1 the per-GPU sampler brought decode kernels' state staleness to
p99 0.5-1 ms, but prefill kept p99 347 ms: 1788 of the last prefill step's
3057 kernels ran after sampling stopped, since the engine logs a step before
the GPU finishes it. Sampling now continues for twice the longest of each
rank's window steps past the last one (capped at 10 s: a DEP8 step waiting on
peers' trace export lasted 45 s).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
At DEP8 c32 three idle DP ranks filled their 32 steps with dummy forwards
while the busy ranks' profilers started 1.4 s after the window's POST, so
counting logged steps stopped sampling 2.25 s before the busy ranks' last
profiled kernels ran. The patch now logs each rank's WorkerProfiler start and
stop (profiler/<rank>.jsonl, the stop before its trace export); the window
client samples until every rank that started has stopped, plus twice the
longest step in the window. Engines without the log fall back to step counts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The re-collected prefill-only c192 run completed; the fault followed a warmup
window in two of three runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Expert routing is dynamic: which experts each token picks decides the work
of every MoE kernel in a step. The patch binds vLLM's RoutedExpertsCapturer
to every MoE layer of the target model before CUDA graph capture, so each
layer's top-k expert ids land in a device buffer from eager steps and graph
replays alike. Inside a window each step's rows are copied to pinned host
memory asynchronously and written at the window's stop as routing/<rank>/
step<k>.npz: per-layer expert token counts for every step, per-token ids for
steps of at most 1024 tokens. The extractor (still stdlib-only) reads them
and gives each step file routing.expert_tokens keyed by MoE module name,
joinable to the MoE kernels' module_stack, and ships the .npz files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hbarclay and others added 2 commits October 1, 2026 19:26
main replaced runners/slurm_utils.sh with infx.launch (#3576) and runs one
AgentX concurrency per srt_agentic.sh invocation. infx-profile.tar is now
staged by the srt driver's finish_single_node, before the server logs are
bundled, with a driver test; the window client wraps the single replay.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Rendered from the per-step artifact of TP8 c4 run 36654381484: a decode
step's FULL graph replay with every layer-20 kernel joined to its host op,
launcher and module, and a prefill step's per-kernel SM clock under the SW
power cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@hbarclay hbarclay changed the title Op-attribution profiling for single-node vLLM runs: per-kernel ops, modules, clocks and MoE routing Single-node vLLM op-attribution profiling: per-kernel op/module, clocks, MoE routing Oct 1, 2026
hbarclay and others added 8 commits October 1, 2026 22:20
On the B200 image (vLLM 0.17.2rc1.dev6310+g591bb95e7) RoutedExpertsCapturer
also takes a kv_cache_config, so binding failed on every rank of DEP8 c32
(run 36914904333) and no routing was recorded. The patch now binds its own
capturer to the capture_fn hooks: each MoE layer's top-k ids are copied into
a [max tokens, layers, topk] device buffer (-1 where unwritten), which CUDA
graphs capture and replay. MoE runners are checked before plain capture
sources, since a runner's hook is its router's.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
INFX_PROFILE mode=synthetic replaces the AgentX replay with fixed shapes
(synthetic_load.py): CONC random-token prompts of isl tokens to
/v1/completions, one-token outputs while the prefill window is open, then
osl-token generations until the decode window closes. Its phase starts go to
a log in aiperf's phase-line form, so the window client anchors its windows
unchanged. No trace dataset, aiperf install or AgentX warmup; the driver
writes the point's result JSON, and the workflow skips AgentX validation.

The clock sampler falls back to amdsmi's gpu_metrics on AMD (graphics clock
into graphics/SM columns, throttle status into event_reasons; fields chosen
from what the part reports). Ranks join to sampler GPUs by UUID or, as
amdsmi and torch format UUIDs differently, PCI address.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only DSV4's recipes and Kimi K3 on B300 set VLLM_USE_V2_MODEL_RUNNER; other
models' recipes may run V1's vllm.v1.worker.gpu_model_runner. The patch now
wraps that runner the same way (steps, capture, sampling, routing), registers
its draft model (runner.drafter), and logs its CudagraphDispatcher.dispatch
calls during V1 steps, where the last one (after the DP batch sync) is the
step's mode; the decode trigger reads each step's last dispatch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
B300 DEP8 captures dozens of CUDA graph sizes; its capture trace (2.1 GB with
Python stacks) got the extractor OOM-killed on the runner, leaving run
37025915873 without a per-step artifact. The extractor now streams the replay
traces for the graph ordinals they replay, finds those graphs' capture ranges,
and loads only events overlapping them. On B200 synthetic TP8 c4 rank 0 that
is 66 graphs, peak RSS 2.7 GB, and all 158 replays join; whole-trace loading
had reported 61 PIECEWISE replays per rank as count mismatches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Kimi K3 on B300 keeps KV in an embedded Mooncake store: each of 8 TP ranks
registers a 281 GB global segment for RDMA. Its synthetic profile run
(37052272356) failed at engine start-up with "Failed to register memory:
Cannot allocate memory" from the transfer engine. host_headroom_gib now also
shrinks a mooncake-master service's per-rank segment by headroom/ranks
(281 GB -> 265 GB at 128 GiB over TP8), as it already did for SimpleCPU pools.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
srtctl applies each --set to base and to every override_* section, so the
indexed services[0] path created a mapping in variants without services and
failed golden acceptance ("cannot index a non-list with [0]"). Set the
selected recipe's services list, with the shrunk segment, in one override.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ROCm's tracer gives graph-replayed kernels no node id, so every MI355X
decode kernel fell through to the eager path and took the hipGraphLaunch
context (no op). Device start order is not capture order across the two
replay streams (347 of 2665 positions differ between replays of the DSV4
decode graph), but each stream keeps capture order. Pick the interleaving
of the per-stream sequences that best matches the kernel families each op
ran eagerly, ties to start order, once per graph shape.

MI355X DSV4 TP8 rank 0: 96/96 replays joined, 0 count mismatches; decode
moves from 94k eager rows to 90k graph rows. B200 node-id joins unchanged
(158 joined, 0 mismatch).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The MI355X DSV4 agentic decode graph replays on three streams, which the
two-stream DP left unjoined (256 replays). Replace it with a beam search over
the streams' heads, and score a memset/memcpy against a kernel launch (or the
reverse) as a mismatch. Matches the exact DP on every two-stream graph;
beam 64 and 256 agree on the three-stream graph, with 0 kernel/copy
mismatches against 112 in device order. Rank 0: 96/96 replays joined.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@hbarclay hbarclay changed the title Single-node vLLM op-attribution profiling: per-kernel op/module, clocks, MoE routing Single-node vLLM op-attribution profiling (NVIDIA + AMD): per-kernel op/module, clocks, MoE routing Oct 5, 2026
hbarclay and others added 9 commits October 5, 2026 18:00
The launcher hooks named vllm.utils.flashinfer and DeepGEMM's mega module
only, so kernels launched from other vendored packages had no op and no
launcher: on MiniMax-M3 (B200/B300) the sparse attention forward/combine,
FlashInfer sparse top-k and the k2q index kernels from
vllm/third_party/fmha_sm100/api.py, and FlashInfer's plan kernel, ~5% of
device time. Hook every module under vllm.third_party and flashinfer on
import: public functions and public methods of the classes it defines.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On ROCm, keep each captured graph (instantiated at first replay instead of
at capture end) and record its kernel/memcpy/memset nodes in creation order
with each kernel node's function name (hipGraphGetNodes,
hipGraphKernelNodeGetParams, hipKernelNameRef[ByPtr]). The extractor
attaches those names to the captured launches when the node and launch
sequences agree, and the stream alignment then scores each position on the
node's own kernel instead of on eager evidence. Missing APIs degrade to the
previous evidence-based alignment.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
At compilation mode 3 Dynamo inlines module calls, so module hooks never run
and compiled pieces carried only infx_piece markers (DSV4 MI355X: 1-9% of
device time had a module path). Inductor's scheduler now writes one host
line per scheduler node into the wrapper code, naming the deepest module
that holds all of the node's source FX nodes (nn_module_stack); at run time
it opens an infx_mod marker resolved to the registered module name, closed
by the next line or the piece's exit. Host-only, so kernels and fusion are
unchanged. Profiled runs disable vLLM's and Inductor's compile caches so
the wrapper code is regenerated with the lines.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… wrappers

The MI355X probe (torch 2.13) showed both AMD additions inert:
raw_cuda_graph() refused every graph because the pybind base __init__ runs
with the caller's keep_graph=False after the patched __new__, and the module
lines were never written because compiled regions use
SubgraphPythonWrapperCodegen, which an exact type check skipped. Override
__init__ too, accept any Python wrapper (C++ wrappers stay excluded), and
record one note per compile path taken in compile/<pid>.jsonl.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The node lists did not help: HIP lists a graph's nodes in its own
interleaving of the streams, not the host's launch order (MI355X DSV4
decode graph, first difference at node 27), so they cannot be tied to the
captured launches, and forcing keep_graph changed when graphs instantiate.
Remove them. Instead mark torch.cuda.stream blocks (infx_stream#<handle>):
each replay stream runs its captured launches in order, so a graph whose
capture-side stream groups pair one to one with its replay streams (same
count and kernel/memcpy/memset sequence) joins in order per stream. Other
graphs keep the beam alignment; report.json counts joined_by_stream_marks
and joined_by_alignment. Also read internal-linkage mangled names
(_ZN4vllmL16...) as one family, and note which FX meta compiled nodes carry
when they lack nn_module_stack.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d runs

vLLM 0.31 no longer reads VLLM_RPC_TIMEOUT (it warns it is unknown); its
worker RPC limit is VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS (300 s). And with the
compile caches off, start-up recompiles: a profiled DSV4.1-Flash MI300X
start-up was cancelled right after a 597 s model load against the 600 s
VLLM_ENGINE_READY_TIMEOUT_S. Set them to at least 1800 s and 3600 s,
keeping any larger value the recipe sets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On vLLM's compile path the FX nodes Inductor sees carry no nn_module_stack
(MI355X probe: meta has stack_trace, source_fn_stack, from_node only), so
no module lines were written. Inductor's placeholders keep the order of the
graph handed to compile_fx, whose Dynamo names spell each weight's owner
(l_self_modules_layers_modules_3_..._parameters_weight_): record those names
per compile_fx call and give each scheduler node the deepest module common
to the weights it reads. Nodes that read no weights get no module (their
line closes the previous marker) instead of inheriting one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The module hooks were limited to compilation mode 0 so Dynamo would never
trace them. On the AMD recipes (mode 3) the model itself runs eagerly
around a few small compiled helpers (MI300X DSV4.1-Flash: compiled graphs of
4-7 inputs such as l_input_, l_x_, l_self_alpha), so 90-98% of device time
had no module path. The hooks now return at once while Dynamo traces (no
graph breaks, no profiler ops in compiled graphs), so they are enabled in
every mode; compiled helpers keep their Inductor wrapper lines.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant