Skip to content
This repository was archived by the owner on Oct 4, 2026. It is now read-only.

ROCm on Windows: keep VRAM resident and bound attention and VAE memory - #304

Merged
Pfannkuchensack merged 38 commits into
upstream-mergefrom
feat/rocm-windows-memory
Oct 1, 2026
Merged

Pfannkuchensack merged 38 commits into
upstream-mergefrom
feat/rocm-windows-memory

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Summary

Generations on a ROCm build of PyTorch under Windows ran 4-5x slower than the hardware allows, and Krea-2 fp8 did not run at all. Windows never fails an allocation that does not fit: it moves it into shared system memory and keeps it there, so nothing in the log says why a run crawls. Measured on an RX 9060 XT (gfx1200, torch 2.12+rocm7.14.1): Z-Image Turbo nvfp4 at 1024px denoised in 144/139/186 s over three runs; it now takes 34 s, with pixel-identical images. Krea-2 fp8 produced no step in two minutes; it now takes 60 s per image. Z-Image at 1536px is possible for the first time.

Four causes, addressed in order:

  1. The allocator fragments. An allocation that finds no contiguous VRAM is placed in system memory as a whole, even with the budget and plenty of room free, and it stays there. A ROCm build on Windows now defaults to expandable_segments:True, set before torch is imported and only when no allocator variable is configured. Any explicit pytorch_cuda_alloc_conf still wins.
  2. The model cache plans against memory Windows will not keep resident. torch.cuda.mem_get_info is the device total minus this process's own usage there: it ignores other processes, and Windows starts paging once the process passes its WDDM budget (15.09 of 15.92 GiB alone; next to another GPU process it stays there while this process holds little and steps down to this process's share once the two together pass ~12 GiB -- 7.58 GiB next to an 8 GiB holder). The cache's free figure is now capped by that budget through D3DKMTQueryVideoMemoryInfo, and a worker says so when memory stays in system RAM across two sessions (PDH GPU Process Memory\Shared Usage). Both are best-effort and silent everywhere else.
  3. gfx1200 has no working fused SDPA kernel on this build, so attention runs on the math kernel and materializes the whole score matrix: 8 GiB per call for Z-Image at 1024px, 12.9 GiB for Krea-2. PyTorch enables AOTriton's flash and memory-efficient kernels on gfx1200 only with TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1, which InvokeAI does not set. On torch 2.12+rocm7.14.1, every call that reaches them with the flag set fails with hipErrorInvalidValue (Z-Image, FLUX.1, CLIP and VAE shapes, masked and unmasked), and a full CLIP text encoder forward does not finish. torch 2.13.0+rocm10.0.0 runs them correctly with the flag, the VAE's 512-wide heads included; turning the flag on belongs to the ROCm 10 upgrade, not to this PR. The existing ROCm SDPA guard now computes such a call in chunks of at most 1 GiB of scores - head groups first, query rows otherwise - and the working-memory estimates are capped to match. Chunking is exact in exact arithmetic and lands within one bf16 ulp of the unchunked call; head groups are bitwise identical where the kernel's arithmetic does not depend on the batch it runs over, which holds on CPU but not for a rocBLAS batched GEMM that re-tiles with the group count. Krea-2's denoise reservation prices the score matrix, which it did not before.
  4. VAE decodes only tiled after an out-of-memory error, which Windows never raises. Large decodes are now tiled up front when the untiled estimate would claim more than 90% of the VAE's device (Anima keeps its measured 70%), switchable with the new auto_tiled_decode setting. On Windows the line is the WDDM budget, overridden only by what this process keeps resident (CurrentUsage minus its paged bytes). For the Qwen-Image VAE the decision uses the measured peak curve (the chord between measured points) instead of the reservation constant; reservations keep the constant. On ROCm/16 GiB that moves 197 sizes from tiled to untiled; none become newly tiled. The FLUX.1 autoencoder's working-memory constants follow the convolution backend, as the FLUX.2 ones already did: MIOpen needs 3600 B/pixel-byte to decode where cuDNN needs 2200.

Items 3 and 4 change behaviour on every platform: math-kernel attention runs in chunks on Linux ROCm too, and decodes that would take most of the card are tiled everywhere. On CUDA the estimates and the fused path are untouched.

Documentation says only what is true: AMD remains supported on Linux only; the Windows ROCm notes are hints.

Related Issues / Discussions

None.

QA Instructions

Gates

  • uv run --no-sync pytest -n 6 (after merging origin/main): 11,787 passed, 184 skipped, 9 xfailed.
  • uv tool run ruff@0.11.2 check / format --check on the changed paths: clean.
  • pnpm -C docs build: 326 pages, link validation clean. Generated openapi.json, schema.ts and docs/src/generated/settings.json were regenerated with their generators; generate_docs_json.py now emits Path defaults with .as_posix(), so regenerating on Windows no longer flips the separators and fails check-docs-data.
  • ROCm hardware tests (-m slow, RX 9060 XT): 5 passed, 1 skipped. The skip is test_a_fused_kernel_is_still_wrong_for_the_wide_head: no fused backend runs that shape on gfx1200 with this build, so there is nothing to compare.
  • In the ROCm venv, 4-5 tests in tests/app/invocations/test_flux2_working_memory.py fail depending on the build (5 on gfx1200 / torch 2.12+rocm7.14 here); they fail identically on the base commit, as they assume a CUDA/Linux backend.

End-to-end, from this branch's code with no runtime patches

RX 9060 XT, 16 GB (port 9091):

Run Result
Z-Image nvfp4 1024px x3 50.7 / 36.3 / 36.3 s (denoise 34.3 s), pixel-identical to the pre-change baseline
Z-Image nvfp4 1536px x3 180.4-180.9 s, decode 2.5-2.7 s, no overflow
Krea-2 fp8 1024px x3 78.3 / 59.6 / 60.2 s; runs 0 and 2 bit-identical. The same image takes 400.7 s at the merge-base
Shared system memory, both 0.28 GiB peak
Paging warning exactly one, after a second process held 8 GiB; silent in clean runs

RTX 4090, 24 GB (port 9090), for the CUDA regression:

Run Result
Z-Image nvfp4 1024px x3 identical to the stored reference images; denoise 4.84 s, decode 0.25 s, unchanged
Z-Image nvfp4 2560px tiles up front (26.9 GiB estimated > 90% of 24 GiB), decode 3.1 s, no OOM

Measurements behind two design decisions

  • The cap is the budget minus this process's live allocations, not minus CurrentUsage. HIP under Windows keeps 1-2 GiB after a free plus empty_cache() and hands it to the next allocation, while CurrentUsage still counts it: 4 GiB allocated, then freed, leaves CurrentUsage at 1.14 GiB with memory_reserved at 0 and torch's free figure back at its starting value; the next 2 GiB cost only 1 GiB of new usage. Capping against CurrentUsage would hide what an offload just freed.
  • Under expandable segments the offload loop empties the allocator after each model it moves out. Freed pages are invisible until then (del freed 0.00 GiB, empty_cache() freed 2.99 GiB), so the loop saw no progress and unloaded every unlocked model.

Not verified

Linux ROCm (the query-row chunks in the VAE and the FLUX.1 MIOpen constant), multi-GPU under Windows ROCm, other Windows ROCm builds, MPS/XPU pre-tiling, and small CUDA cards on real hardware. The Qwen-Image VAE's ROCm constants are unchanged (measurable only on Windows here), and auto_detect_slice_size still uses torch.cuda.mem_get_info directly.

Review

Material findings resolved:

  • A settings.json regenerated on Windows flipped two path defaults and would have failed the docs job; the generator now emits POSIX paths.

  • The allocator test fixture leaked PYTORCH_CUDA_ALLOC_CONF into later tests, which made a model-cache test fail depending on file order.

  • The offload loop over-unloaded under expandable segments (see the measurement above), now covered by a regression test that fails without the fix.

  • The allocator default silently did nothing when torch was already imported; it now warns instead.

  • The paging warning could fire on an overflow that returns to VRAM by itself, and advised a restart; it now requires two consecutive readings and names settings to change.

  • Chunking sliced 3-D and wider-batch K/V wrongly; those calls stay whole.

  • The FLUX.1 decode node reserved nothing for a diffusers-layout AutoencoderKL, which is the exact failure this PR targets; it is now priced and pre-tiled like the Z-Image node.

  • The up-front tiling log was at INFO on every Anima decode and suggested turning the setting off, which is bad advice for Wan (no OOM retry) and Anima (tiling measured faster).

  • Test gaps closed: masked query-row chunking, Krea-2 with CFG and a regional mask, auto_tiled_decode=false for Wan and Anima, a Wan cpu_only test that passed regardless of the code, PDH struct offsets plus a hardware test, and a test that the worker loop calls the warning at all.

  • The pre-tiling ceiling over-trusted CurrentUsage: next to a foreign process, or while over-committed, it counts what Windows has already paged out. The override is now bounded by residency, with tests for both states and area-monotonicity tests for the Qwen pricing on both backends.

Compatibility / Rollout

  • New setting auto_tiled_decode (default true); generated OpenAPI, frontend types and docs settings regenerated.
  • No persisted-data migrations. The allocator default only applies to a ROCm build on Windows and only when nothing else is configured.
  • wddm.py is new and best-effort: every failure answers "unknown" and leaves the previous behaviour in place.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

With pytorch_cuda_alloc_conf unset, a ROCm build on Windows now runs with expandable_segments:True,
so fragmented VRAM no longer pushes allocations into shared system memory.
An explicit allocator setting or allocator env var still wins; docs and generated config artifacts updated.
On ROCm under Windows, free VRAM is now capped by the WDDM budget (gdi32 D3DKMTQueryVideoMemoryInfo),
past which Windows pages allocations into system memory instead of failing them.
torch's figure there ignores other processes and runs up to the physical total; other platforms are unchanged.
After each session, a worker on a Windows ROCm device reads the process's shared GPU memory (PDH) and warns
once per episode when more than 512 MiB of it sits in system memory, where every generation that touches it slows down.
Adds a low-VRAM docs section on AMD GPUs under Windows.
The ROCm SDPA guard now splits any math-kernel call over 1 GiB of scores into head groups (bitwise identical)
or query rows (one oversized head), and working-memory estimates price one chunk instead of the whole matrix.
Z-Image 1024px peaks at 0.5 GiB instead of 8 GiB per attention on an RX 9060 XT; renamed to install_rocm_sdpa_guard.
The Krea-2 working-memory estimate now adds the score matrix where the build materializes it (no fused kernel),
priced from the loaded model's heads and the attended sequence; one 1 GiB chunk on ROCm, nothing on CUDA.
Before, Krea-2 fp8 on an RX 9060 XT loaded the whole transformer and paged each 12.9 GiB attention into system RAM.
On ROCm the FLUX.1 VAE (also Z-Image's) now budgets 3600 decode / 2750 encode bytes per pixel·byte instead of 2200/1100,
measured 3451/2688 on an RX 9060 XT: the same convolution stack and numbers as the FLUX.2 VAE, so both share one table.
The estimate follows the VAE's compute device, so a cpu_only VAE keeps the cuDNN column.
…code

The FLUX.1, Z-Image, Qwen-Image (Krea-2) and Wan decodes now tile when the untiled estimate exceeds 90% of the VAE's
GPU memory (Anima keeps its measured 70%), instead of relying on an OOM that Windows paging never raises.
New auto_tiled_decode setting (default on) turns it off; force_tiled_decode and the tiled field still win.
- Cap free VRAM at the WDDM budget minus live allocations: HIP keeps freed memory that CurrentUsage still counts
- Empty the allocator cache after each offload under expandable segments so the cache stops once enough is free
- Hardware tests for the headroom after a free and for the PDH lookup
- Skip the allocator default once torch is imported; stop the test fixture leaking it
- Warn only about memory paged across two sessions, with actionable advice
- Leave K/V broadcasts the chunks cannot slice whole; cover masked row chunks and Krea-2 CFG/regional pricing
- Price and pre-tile FLUX.1 decodes with a diffusers-layout VAE
- Log up-front tiling at debug; document what auto_tiled_decode off means
- Emit POSIX path defaults in the generated docs settings
@lstein

lstein commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

I can verify this on Linux ROCm, but I don't have a Windows boot on my AMD rig.

@lstein

lstein commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Passes all the ROCm tests, including test_a_fused_kernel_is_still_wrong_for_the_wide_head on my W7900 rig.

These tests fail on purpose. They pin behaviour the current up-front
tiling rule does not have, and should go green with the fix.

- A Qwen-Image/Krea-2 decode whose measured peak fits the card is tiled
  anyway, because the gate is fed the padded reservation figure (a flat
  5500 B/pixel-byte) instead of the expected peak (3273 measured at
  1536px). Tiling is not pixel-identical, so this silently changes output
  for images that would have decoded in a single pass.
- should_pretile_vae_decode compares against the card's nameplate total,
  not the Windows video-memory budget this PR added
  TorchDevice.cuda_mem_get_info for, so a decode Windows will page into
  system memory is left untiled.
@lstein

lstein commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Adversarial review

Reviewed against c2e4cf6ee6 with three independent read-only passes (attention chunking; WDDM/allocator/cache; VAE tiling and working memory), each verified against the code afterwards. A lot of this holds up well — the ctypes work is correct down to the struct offsets, the chunking survived ~450 differential shape cases, and both generated artifacts regenerate byte-identical. The findings below are what survived verification.

I've pushed 33e16af with three intentionally-failing tests for the two findings I think are blockers. CI will be red by design — they pin behaviour the current rule does not have and should go green with the fix. Happy to drop that commit if you'd rather not carry red tests.


1. Reservation headroom is reused as a tiling trigger, so decodes that fit are silently tiled

invokeai/backend/util/vae_working_memory.py:22-38

The estimator constants are deliberately over-provisioned ("max observed + ~8% headroom"). Over-reserving used to cost only cache eviction. Comparing that padded figure against a 90% line converts the conservatism into a non-pixel-identical output change, with no error and nothing above DEBUG in the log.

Qwen-Image is the sharp case, because its ROCm decode constant is a flat 5500 while the curve in your own table is not flat — it falls to 3273 at 1536². Computed by calling the real estimator:

px estimate (5500) measured peak outcome
1024 10.74 GiB 8.93 GiB untiled everywhere
1536 24.17 GiB 14.38 GiB tiles on 16 and 24 GiB — fits both
1792 32.90 GiB 22.34 GiB tiles ≤24 GiB (genuinely too big)
2048 42.97 GiB 37.60 GiB tiles (correct)

On a 24 GiB ROCm card a 1536² Qwen-Image decode that previously ran untiled and succeeded now tiles. Krea-2 decodes through this node too. On CUDA FLUX.1 the crossovers are benign (shipped 2200 vs measured 2185), so this is specifically Qwen/Krea-2 on MIOpen.

The reservation is right as it is — it's the tiling decision that shouldn't be carrying reservation headroom. Test: TestPretilingDoesNotFireOnReservationHeadroom in tests/app/invocations/test_qwen_image_working_memory.py, written at the node level so any fix shape satisfies it.

One knock-on: test_a_decode_too_large_for_its_gpu_is_tiled_up_front_unless_switched_off asserts pretile.assert_called_once_with(compute_device, 20 * 2**30) — the raw padded estimate — so it pins the current behaviour and will need updating alongside.

2. The gate measures total VRAM on the platform this PR taught the cache not to trust

invokeai/backend/util/vae_working_memory.py:33-38

should_pretile_vae_decode uses torch.cuda.get_device_properties(device).total_memory. The docstring justifies the whole feature with "on Windows, drivers page an allocation that does not fit into system memory instead of failing it (always for ROCm)" — and this same PR adds TorchDevice.cuda_mem_get_info capping free VRAM by wddm.video_memory_budget for exactly that reason. The gate consults neither that budget, nor free memory, nor max_cache_vram_gb.

Trigger: 16 GiB RX 9060 XT, a browser holding ~3 GiB so the budget is ~12.5 GiB. FLUX.1/Z-Image at 1408² estimates 13.29 GiB — under the 14.4 GiB threshold, over the budget. No pre-tile, no OOM (Windows pages instead), so the retry at flux_vae_decode.py:114 never fires and the generation just crawls. That is the failure the helper exists to prevent. Your own measurement says Windows lowers the budget within about a second of another GPU process starting.

Also reachable on Linux/CUDA: max_cache_vram_gb: 6 on a 24 GiB card means nothing below 21.6 GiB ever pre-tiles while the cache budget is 6 GiB. Qwen has no OOM retry, so it hard-fails.

Test: test_the_pretile_gate_uses_the_memory_the_device_will_keep_resident in tests/backend/util/test_vae_pretile.py. It patches both wddm.video_memory_budget and the devices binding, so a fix routing through either is covered.


Other material findings (no tests pushed)

3. The expandable-segments offload fix is inert on multi-GPU. model_cache.py:2558, 2576-2577, 2584. TorchDevice.empty_cache() is peer-aware — when another registered generation device holds its lock it sets _empty_cache_deferred and returns without freeing. So the new per-offload call is a no-op in that window, _get_reclaimable_allocator_bytes() returns 0 under expandable segments, the loop sees no progress, and every unlocked model is offloaded. Line 2584 now reads if vram_bytes_freed > 0 and not empty_cache_per_offload, so the end-of-loop release is skipped too. Net vs base is "no better", not "worse". test_offloading_under_expandable_segments_stops_once_enough_is_free can't see it because it replaces TorchDevice.empty_cache wholesale with a fake, so the deferral never runs. Single-GPU is unaffected (both callers run on the session thread); the trigger is multi-GPU + expandable segments, which is the default on Windows ROCm.

4. The "bitwise identical" claim for head-group chunking is false. attention.py:135, :232, and the PR description. Reproduced on gfx1100 / torch 2.13+rocm7.2 with Z-Image's own shape (1,30,4608,128) bf16 under sdpa_kernel([MATH]): the math kernel is run-to-run deterministic, yet chunked == whole is False, max abs diff 0.00048828125. Splitting the batched GEMM's 30 head-batches into groups of 2 changes rocBLAS's tiling. Under one bf16 ulp, so it's a documentation fix rather than a code one — but test_heads_are_split_first_and_the_result_is_bitwise_unchunked asserts torch.equal on CPU fp32, where the claim does hold, so it validates the guarantee on the one platform it isn't about. The head-groups-vs-query-rows distinction is actually inverted for real shapes here: the FLUX VAE mid-block (row chunking) was bitwise identical in the same run.

5. paged_bytes can raise, breaking the module's stated contract. wddm.py:337. _pdh_query = _open_shared_usage_query(pdh) sits outside the try/except guarding the collect, and that helper checks return codes but catches nothing — on Windows ctypes turns an SEH fault into OSError. The module header promises "any failure yields None … callers must keep their existing behaviour." Contained today because the only caller wraps it.

6. auto_tiled_decode: false doesn't restore previous behaviour. Anima's 0.7 rule was unconditional before this PR, and its documented purpose isn't OOM avoidance — it's keeping the ~4 GB transformer resident ("7s+ observed on 8GB, vs ~1s tiled"). With the switch off on an 8 GiB card at 1024², you get the 7x slow path and the setting's own promise ("still retry tiled after running out of memory") doesn't cover it, because no OOM occurs. The switch also doesn't undo the reservation changes in the FLUX.1 and Z-Image nodes.

7. SD1/SDXL, SD3 and CogView4 left on the cuDNN constant. vae_working_memory.py:89, :135, :635. The module header groups "the diffusers AutoencoderKL (SD1/SDXL, SD3, CogView4) and the FLUX.1 AutoEncoder" as the same network, and this PR's evidence is that MIOpen needs 3600 where cuDNN needs 2200 for exactly that stack. Those three still hardcode 2200 if decode else 1100 and don't take a device, so on ROCm they under-reserve ~1.64x, and the pre-tiling gate inherits the under-estimate. The backend-branch pattern already exists two functions away in estimate_vae_working_memory_qwen_image. Pre-existing, but this PR quantifies it. Relatedly flux2_vae_decode.py — the node these constants were fitted on — has no tiling code and no OOM retry, so "tiled everywhere" is 5 of 12 decode nodes.

Smaller

  • sdpa_score_matrix_bytes caps at SDPA_MATH_CHUNK_BYTES whenever rocm_sdpa_chunks_math() (attention.py:530-532), which only tests ROCm + sentinel. The guard declines to chunk for is_causal, dropout_p > 0, non-4-D query, and 3-D/wider-batch K/V. No current caller prices such a shape, but a future one would be reserved 1 GiB against a matrix up to ~20x larger. The predicate name reads stronger than it is.
  • cuda_mem_get_info calls video_memory_budget unguarded (devices.py:471-473) on the hot path of every generation. I traced the escape surface — index-less devices, _resolve_adapter and the D3DKMT call are all caught — so the only hole is _load_gdi32 raising outside (AttributeError, OSError). Defensive gap, not a live bug.
  • Adapter handle leak if an exception escapes the enumeration loop (wddm.py:222-244); bounded to ~1 per device per process.
  • Contradictory docstrings on empty_cache() under expandable segments: model_cache.py:2344-2347 says it "reclaims nothing", :2555-2557 says an offload "only shows up once empty_cache() unmaps its pages". The newer one is right; the older is the stated justification for returning 0 there.
  • Two ROCm detectors disagree: torch_cuda_allocator.py:36-38 tests "+rocm" in the packaging metadata, wddm._supported() tests torch.version.hip. A locally-built Windows ROCm wheel gets the budget cap but not expandable segments, silently — and the docs point these users at self-installed torch.
  • rocm_causal_conv3d.py:143 still names install_rocm_sdpa_head_dim_guard.
  • flux_vae_encode.py:47 / z_image_image_to_latents.py:56 now price against compute_device, but lines 53/67 still place the tensor on choose_torch_device(). Pre-existing; the partial threading makes it visible.
  • auto_detect_slice_size (attention.py:34) and diffusers_pipeline.py:223 still use raw torch.cuda.mem_get_info.
  • In qwen_image_latents_to_image.py the tiled local goes stale after the pre-tile branch (only effective_tile_size is updated, which is what actually drives tiling — I confirmed tiling does engage). A later if tiled: added below would silently be wrong.

Verified as claimed

  • ruff check / format --check clean on all 37 changed Python files.
  • The test_flux2_working_memory.py failures are genuinely pre-existing — I ran the base tree from git archive and got the identical set. It's 4, not 5.
  • docs/src/generated/settings.json and openapi.json both regenerate byte-identical; schema.ts matches config_default.py.
  • "CUDA estimates untouched" holds: the new cudnn column is exactly the old hardcoded FLUX.1 literals, and install_rocm_sdpa_guard returns early off ROCm.
  • The docs' "~1850×1850 on a 16GB Nvidia GPU" matches my computed 1874px.
  • Deleted TestUseTiledDecode coverage migrated cleanly into test_vae_pretile.py.
  • Wan and Anima already pre-tiled at 0.9/0.7; switching them to vae_info.compute_device is a correctness improvement on multi-GPU.
  • Chunking held up: row chunking slices the query axis (softmax is over keys); mask broadcasting is right for 2/3/4-D, bool and additive, including the B == H == L == S trap; GQA grouping forces a multiple of heads // kv_heads; is_causal and dropout_p > 0 are excluded before any slicing, so the classic causal-chunk bug is unreachable; zero-extent inputs never enter the chunked path. Krea-2's score-matrix pricing matches the chunk budget, CFG sequencing and the GQA expansion.

I could not complete a full-suite run — this box was saturated and -n logical is not safe here. Changed-file tests are green apart from the three I added and the four pre-existing FLUX.2 failures.

…droom

- Compare against the Windows video-memory budget where there is one, not the card's total
- Qwen-Image ROCm constants now follow the bounded math attention this branch ships: measured 2650-2772 decode and
  1541-1552 encode across 512-2048px, against 5500/6300 fitted while a call built its whole score matrix
- Calibrate with the attention guard installed, so the script measures the path the app runs
…ld up

- Credit freed bytes while a peer device defers empty_cache, so multi-GPU stops over-unloading
- Never let the budget or paged-bytes lookups raise, and close adapter handles if enumeration throws
- Detect a ROCm build from torch's version.py too, so a locally built wheel gets the allocator default
- Correct the bitwise-identical claim for head-group chunking
…kend

- auto_tiled_decode no longer disables Anima's own rule, which is a speed optimization, not an OOM fallback
- SD1/SDXL, SD3 and CogView4 take the MIOpen constant on ROCm, like the FLUX.1 autoencoder they share a stack with
- Encode nodes place their tensors on the VAE's device, matching the decode nodes
…eiling

- Keep the padded Qwen-Image reservations; the tiling decision reads the measured curve instead, so a decode whose
  real peak fits the card is no longer tiled by a reservation's headroom
- Compare against the video-memory budget plus what this process can release: Windows halves the budget once the
  process passes ~12 of 16 GiB, and the bare figure tiled a 7.0 GiB decode that fits
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

Thanks — this was a genuinely useful review, and the two tests made the first finding much easier to act on. All three are green now, and everything else you raised is either fixed or answered below. One of my first attempts at finding 1 was wrong in a way worth recording, so I've left that in rather than quietly dropping it.

1. Reservation headroom as a tiling trigger — fixed the way you framed it

Implemented as you described it: the reservation keeps its headroom, and the tiling decision no longer carries it. qwen_image_untiled_decode_peak_bytes prices that decision from the measured curve in the comment above the constants (the per-point W7900 column), so 1536² is compared as 14.4 GiB rather than 24.2 GiB and stays untiled on a 24 GiB card, while 1792² still tiles there and not on 32 GiB. Both your tests run exactly as you wrote them.

I went the wrong way first and want to record why, since it nearly shipped: I re-measured on an RX 9060 XT (gfx1200, torch 2.12+rocm7.14, fp16) and got a flat 2650-2772 decode / 1541-1552 encode across 512²..2048², attributed that to the score-matrix chunking this branch adds, and made the constants conditional on the guard. That attribution does not survive your own table. At 512² the gap between the two cards is 1.24 GB while the entire score matrix there is ~0.3 GB, and at 1536² an unbounded matrix would be ~23 GB against a measured 15.4 GB total — so the W7900 run was never dominated by an unbounded score matrix in the first place. The spread is the card, not the guard, and shipping the lower pair would have under-reserved by 1.9x (decode) to 3.8x (encode) on the hardware the shipped figures were measured on, with Qwen Image Edit's encode having no tiled retry to fall back on. Reverted; the constants are unchanged.

The gfx1200 numbers stay in the comment, because that spread is the actual argument for your finding: a reservation that has to cover a 1.9x card difference is not a figure to decide non-pixel-identical output with.

One thing I did keep: scripts/calibrate_qwen_vae_working_memory.py never installed the ROCm attention guard, so it measured a path the app does not run. It does now.

2. The gate measured the nameplate total — fixed

The gate consults the budget now. My first version compared against the bare budget, and an end-to-end run on the card caught what that does mid-session: Windows holds the budget at 15.09 GiB until the process passes about 12 of 16 GiB and then halves it to 7.62 — below what we already hold. With models resident, the gate read 7.6 GiB, drew a 6.9 GiB line, and tiled a 7.0 GiB Z-Image decode that fits. Output stopped being pixel-identical (PSNR 42 dB) and the decode went from 0.89 s to 1.5 s: your finding 1, reintroduced at the other end.

So the ceiling is the budget plus what this process itself holds, capped at the card: the cache evicts models to honour the reservation, so that memory is available to the decode. Measured:

own allocations 0.14 2.15 6.15 10.15 12.15 GiB
budget reported 15.09 15.09 15.09 15.09 7.62 GiB

Your test passes unchanged (it patches mem_get_info to a process holding nothing, so the ceiling is the budget). A second test pins the mid-session case with these numbers, and the E2E is pixel-identical again.

The max_cache_vram_gb half of that finding I did not implement: that cap bounds model residency, not the working-memory reservation, so a 6 GiB cap on a 24 GiB card still leaves the decode the rest of the card. If you meant something else by it, say so and I'll look again.

3. Offload fix inert on multi-GPU — fixed

TorchDevice.empty_cache() now reports whether it ran. When a peer device defers it, _offload_unlocked_models credits the bytes it just freed to the measurement instead — they are in this process's allocator, which reuses them for the load being made room for. The regression test is parametrized over alone / peer-device-busy; the second case fails without the credit, exactly as you described.

4. "Bitwise identical" — corrected, thank you

You are right, and the CPU test was validating the claim on the one platform it isn't about. The comment and the docstring now say what actually holds: exact in exact arithmetic, bitwise where the kernel's arithmetic does not depend on the batch it runs over (CPU), and within a bf16 ulp on a card whose batched GEMM re-tiles with the group count — with your 0.0005 measurement named. The test keeps its exact assertion and says why it is exact there. The PR description is updated too.

5. paged_bytes could raise — fixed

The query-open moved inside the try, and video_memory_budget is wrapped as a whole, so the module keeps its "any failure yields None" contract at both entry points rather than relying on its callers. The adapter enumeration now closes what it has opened if anything escapes the loop.

6. auto_tiled_decode: false didn't restore previous behaviour — fixed

Anima's 0.7 rule is no longer gated on the setting. It predates it and is a speed optimization, not an OOM fallback; switching it off only bought the 7s path. The setting description and the docs now say so.

7. SD1/SDXL, SD3, CogView4 on the cuDNN constant — fixed

They take a device and use _flux_vae_scaling_constant, so MIOpen gets 3600/2750 like the FLUX.1 autoencoder they share a stack with; the six call sites pass vae_info.compute_device. Also in that area: the FLUX.1 and Z-Image encode nodes now place their tensors and generator on the VAE's device, which is what the decode nodes already did.

flux2_vae_decode.py having no tiling and no OOM retry is real and out of scope here; worth its own issue.

Smaller ones

Done: the rocm_causal_conv3d reference to the old guard name, the contradictory empty_cache docstrings, the adapter-handle leak, the unguarded video_memory_budget call on the generation path, the stale tiled local in the Qwen node, and the estimator docstring now states which call shapes the 1 GiB cap assumes (4-D, non-causal, dropout-free, K/V of the query's batch — what every caller prices).

Two ROCm detectors disagreeing was the interesting one: torch_cuda_allocator now also reads hip = ... out of the installed torch's version.py without importing torch, so a locally built wheel gets the allocator default as well as the budget cap. Verified against both venvs here (ROCm: true, CUDA: false).

Not changed: auto_detect_slice_size and diffusers_pipeline.py still use raw mem_get_info — pre-existing, and out of this PR's scope.

Verification after the fixes

  • Full suite -n 6: 8978 passed. One timing test (test_loop_scheduler_overhead_is_linear[iterate]) failed under six workers and passes alone; it is load-sensitive on this box and unrelated to the change. ruff check and format --check clean, pnpm -C docs build clean.
  • End-to-end on the RX 9060 XT from this branch's code: Z-Image nvfp4 1024px three times, pixel-identical to the pre-change baseline, decode 0.89 s; 1536px still pre-tiled at 177 s; Krea-2 fp8 1024px 59 s, run-to-run bit-identical; shared system memory 0.28 GiB peak.
  • Krea-2 output moved against my older reference by 10.5/255, which is upstream, not this branch: one image at the merge-base (c2e4cf6ee6) differs from the branch by 0.124 and from the old reference by 10.54. The Qwen3-VL encoder path changed on main in between. The same image takes 400.7 s at the merge-base and 59 s on the branch.
  • ROCm hardware slow tests on the RX 9060 XT: 5 passed, 1 skipped (test_a_fused_kernel_is_still_wrong_for_the_wide_head now skips where no fused backend runs the shape at all, as on gfx1200 — it asserted on wrong > 0 and had no kernel to find).
  • On the count of pre-existing test_flux2_working_memory.py failures we measure differently: this box reports 5 in the ROCm venv, identically on the base commit and on the branch (gfx1200, torch 2.12+rocm7.14). Yours reports 4. Either way they are pre-existing and backend-dependent; the PR description now says "4-5 depending on the ROCm build" rather than a single number.

@Pfannkuchensack

Copy link
Copy Markdown
Member Author

@lstein can you test again?

@lstein lstein left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up review

Thanks — all seven findings from the last round are genuinely addressed, and splitting the tiling decision from the reservation is a better shape than what I suggested. The 4 pre-existing FLUX.2 failures are fixed too, and both generated artifacts still regenerate byte-identical.

Two of the fixes overshoot, though: R1 and R2 each make a specific case worse than your own immediately preceding commit, and neither is caught by CI or by the new tests. Those plus the slow-test ImportError are what I'd like resolved before merge.

Verified on this box (torch 2.13.0+rocm7.2) by running the shipped functions, not by reading.


R1 — the budget ceiling cancels the budget out, and regresses the case it was written for

invokeai/backend/util/vae_working_memory.py:47 (feb4d8e)

total_bytes = min(total_bytes, budget_bytes + (total_bytes - free_bytes))

total_bytes - free_bytes is this process's live allocations — which the budget already bounds — so adding them back double-counts. Once our usage reaches total - budget, the budget has zero effect:

ours= 0.00 GiB -> effective ceiling  7.62 GiB
ours= 4.00 GiB -> effective ceiling 11.62 GiB
ours= 8.29 GiB -> effective ceiling 15.91 GiB   <-- budget no longer participates

The credit assumes releasing our memory raises the budget. That holds when we depressed it, but not when another process did — and that is the case the gate exists for. RX 9060 XT 16 GB, a second GPU program takes 8 GiB so Windows trims the budget to 7.62 GiB, Invoke holds 3.0 GiB:

budget=7.62G ours=3.00G -> ceiling=10.62G line=9.56G
Qwen 1024px peak = 8.93G
NEW code tiles?  False
OLD (min(total,budget)) would tile?  True
TRUE ceiling after evicting everything = 7.62G -> decode fits? False

At 40a78452df this tiled. Now it does not, Windows pages the decode, no OOM is raised, and the Qwen node has no retry — the generation just crawls, which is the exact failure mode this PR exists to remove.

The intent is right and the measurement behind it is real (15.09 GiB un-trimmed vs 7.62 trimmed). It just needs to be expressed as a restored budget rather than a credit for our own usage — something closer to min(total, max(budget, untrimmed_budget)), remembering the un-trimmed figure from when the process was small.

Note also that test_the_gate_counts_memory_this_process_can_still_release sits in the clamped regime (ceiling pinned to the full card), so "ignore the budget entirely" passes it too. It discriminates against the old code but not against this fix.

R2 — the CUDA tiling constant has zero headroom and extrapolates through an OOM

invokeai/backend/util/vae_working_memory.py:697

_QWEN_VAE_MEASURED_DECODE_PEAK_CUDA = 2690 is exactly max(2660, 2519, 2690, 2671, 2281) — the maximum observation, not a bound — and the calibration table's 2048² cell for CUDA decode reads (OOM), so the flat constant extrapolates straight through the point the calibration could not complete. On a 24 GiB 4090:

   px  gate peak  reservation     line  tiled?
 1792     17.28G       18.63G   22.85G  new=False  prev-commit=False
 1920     19.83G       21.38G   22.85G  new=False  prev-commit=False
 2048     22.57G       24.33G   22.85G  new=False  prev-commit=True
 2176     25.47G       27.46G   22.85G  new=True   prev-commit=True

At 2048² the gate now leaves it untiled by 280 MB while the reservation (24.33 GB) exceeds the whole card, so the cache logs the shortfall and streams, and the decode then peaks at ≥22.6 GB alongside the context, VAE weights and latents. This is not worse than main (which never pre-tiled), but the PR added the protection and this commit removes it at the one CUDA resolution its own data says fails, on a node with no OOM retry.

Simplest fix: give CUDA the same per-point table as ROCm, or at minimum pad the constant and stop extrapolating past 1792².

R3 — slow LTX-2 tests are broken by the guard rename

tests/backend/ltx2/test_ltx2_real_files.py:197-199 still imports and calls install_rocm_sdpa_head_dim_guard, renamed here to install_rocm_sdpa_guard:

ImportError: cannot import name 'install_rocm_sdpa_head_dim_guard' from 'invokeai.backend.util.attention'

Four tests call _accelerator(). The file is pytestmark = pytest.mark.slow, so neither CI nor the default suite sees it — it only fires on an accelerator box under -m slow, which is the lane this PR's own hardware tests use. main added the caller, this branch did the rename, and the merge did not reconcile them.

R4 — the allocator-default tests are non-hermetic and fail on any ROCm host

The three remaining failures in my full run:

FAILED tests/app/util/test_torch_cuda_allocator.py::TestRocmWindowsAllocatorDefault::test_other_installs_are_left_alone[windows-cuda]
FAILED ...[windows-cpu]
FAILED ...[no-torch-metadata]

_installed_torch_is_rocm() now has two sources of truth, and the fixture stubs only the first:

_installed_torch_version() patched to '2.7.1+cu128'
_installed_torch_is_rocm() = True   <- version.py unpatched: hip: Optional[str] = '7.2.53211'

The product logic is correct, and I confirmed it does not import torch (the before-torch-import contract holds). The problem is the test: it is red on every ROCm developer machine — on a ROCm PR — and on CUDA/CPU CI the new version.py branch never returns True, so the locally-built-wheel case it was added for has no passing coverage. Stubbing the file lookup the way TestRocmBuildDetection does fixes both halves.


Non-blocking

  • Contradictory docstrings about the same bytes. model_cache.py:2342 says a large allocation cannot use holes in an expandable segment, citing a measured hard OOM; :2557 says the allocator reuses them for the load this is making room for. One has to be wrong. The practical risk is bounded — the credit is loop-local (2563-2585) and both consumers re-measure honestly (lock() at 2109, make_room_in_vram at 3061), so a wrong credit means under-offloading that the streaming fallback still catches — but for the same reason the stated benefit only partly reaches the load.
  • Handle leak fix is one short. wddm.py:226 info = pending.pop() removes the entry before the body runs, so if D3DKMTQueryAdapterInfo raises, that handle is in neither matches nor pending and the except at 241-246 misses exactly the one that failed — despite the comment claiming otherwise. Bounded to one per device per process.
  • Anima's gate still uses a flat unmeasured 2900 (anima_latents_to_image.py:112) while sharing the VAE whose measured curve this delta just added; qwen_image_untiled_decode_peak_bytes looks like the right input there.
  • The classic-VAE MIOpen branch has no test. Nothing asserts that estimate_vae_working_memory_sd15_sdxl / _sd3 / _cogview4 take the MIOpen column on ROCm, or that a cpu_only VAE takes cuDNN. Worth covering, because the change is significant: an SD1.5 1024px fp32 decode reservation goes 9.49 GB → 15.36 GB on ROCm, which on an 8 GiB card forces full eviction on every decode.
  • sdpa_score_matrix_bytes's unconditional chunk cap is now documented rather than conditioned — still latent, no current caller prices a causal/3-D shape.
  • The Qwen encode path has no gate at all, though the ROCm encode constant (6300) is the largest in the module and the "Windows never raises OOM" premise applies identically.
  • estimate_vae_working_memory_qwen_image treats device=None as build-wide while the FLUX family resolves the session device — harmless today since every production caller passes a device, but the two neighbouring functions now mean different things by the same omission.

Checks

  • Full suite, -n 4: 3 failed, 11201 passed, 177 skipped, 8 xfailed — the 3 are R4. Previous round was 7 failed / 8959 passed.
  • tests/app/invocations/test_flux2_working_memory.py: 149 passed (the 4 pre-existing failures are fixed).
  • ruff check / format --check on the changed files: clean.
  • scripts/generate_docs_json.py and scripts/generate_openapi_schema.py both reproduce the committed files byte-identically; schema.ts matches.
  • My two regression tests from the last round now pass.

- The ceiling is the card minus foreign residency (new wddm.other_process_local_bytes): measured on an RX 9060 XT,
  the budget reads 15.09 GiB next to an 8 GiB holder with ~7.7 GiB available, and 7.62 GiB when Invoke itself holds
  12 GiB and eviction would free the card. The budget stays as the fallback
- Both backends get their measured decode curve; past the last measured area the decision uses the reservation
  constant instead of extrapolating through a calibration point that ran out of memory
- Make the allocator and pre-tile tests hermetic on a ROCm host, rename the guard call in the LTX-2 slow test, and
  cover the MIOpen branch of the SD/SD3/CogView4 estimators
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

All four are fixed. R1 took a measurement round to get right, and it ended somewhere other than either of our proposals — details below, with the numbers.

R1 — right diagnosis, and the fix is neither a credit nor a remembered budget

You are right that budget + (total - free) cancels the budget out, and your worked case is real: Invoke holding 3 GiB next to a foreign 8 GiB stopped tiling a decode that does not fit. That is the failure this PR exists to remove, and I put it back. Thank you for running it rather than reading it.

Your suggested shape (remember the un-trimmed budget) does not survive the measurement either, though. I probed the budget against adapter-wide usage on the RX 9060 XT (16 GB) while a second process held 8 GiB:

                     budget   adapter-wide   ours   others
idle, holder present  15.09           8.40   0.14     8.26
we hold 2 GiB         15.09          10.40   2.19     8.21
we hold 4 GiB         10.95          12.41   4.19     8.21
we hold 6 GiB          8.95          14.41   6.20     8.21
we hold 8 GiB          7.55          14.41   8.20     6.20

In the state your case describes — the foreign 8 GiB present, Invoke small — the budget is not trimmed at all: it reads 15.09 GiB while only ~7.7 GiB can be had. Windows trims it as our own residency grows, not when someone else takes memory. So there is no un-trimmed figure to restore: the un-trimmed value is exactly the one that is wrong there. Both readings mislead, in opposite directions:

state budget says truth
another program holds 8 GiB, Invoke idle 15.09 GiB ~7.7 GiB available
Invoke holds 12 GiB, nothing else on the card 7.62 GiB ~15.9 GiB, released by eviction

What the decision actually needs is what other processes hold, which neither figure answers. wddm.other_process_local_bytes reads the adapter-wide GPU Adapter Memory\Dedicated Usage counter and subtracts our own usage, and the ceiling is total - others. That gives the right answer in both rows above, and in your case: others 8.2 → ceiling 7.7 → line 6.9 → the 8.93 GiB Qwen decode tiles.

Two things fell out of the measurement that are now in the code:

  • Our own usage comes from torch, not from PDH's per-process counter: with two processes on the card, GPU Process Memory swapped its instance attribution between samples (rows 4-5 above report "ours 6.21" while we held 12 GiB). Adapter-wide was stable throughout.
  • Where the counters cannot answer, the fallback is the bare budget — too pessimistic mid-session, never too optimistic.

Three tests now pin it: the foreign-holder case (tiles although the budget looks fine), the trimmed-by-our-own-residency case (does not tile), and the fallback. You were right that the old test discriminated against the old code only; the first of these is the one that fails for "ignore the budget entirely".

R2 — CUDA has its own curve now, and nothing extrapolates

Both backends carry their measured per-point table. CUDA's grid ends at 1792² because 2048² ran out of memory when it was calibrated, and past the last measured point the decision falls back to the reservation constant rather than the maximum seen — "we have not measured this" has to mean the conservative figure, not a flat carry. On a 24 GiB card:

   px  gate peak   reservation   line     tiles?
 1536     12.60G        13.68G  23.19G    False
 1792     14.65G        18.63G  23.19G    False
 1920     21.38G        21.38G  23.19G    False
 2048     24.33G        24.33G  23.19G    True

2048² tiles again, and the gate figure no longer sits under the line while its own reservation exceeds the card.

R3 — guard rename reconciled

tests/backend/ltx2/test_ltx2_real_files.py calls install_rocm_sdpa_guard now. You are right about the lane: -m slow on an accelerator is exactly where this branch's own hardware tests live, so it should have been caught here.

R4 — allocator tests are hermetic

The fixture stubs both sources now, the file lookup the way TestRocmBuildDetection does. 16 pass on the CUDA venv and on the ROCm venv, where 3 failed before.

Non-blocking, also done

  • Handle leak was one short — correct, and the comment claimed otherwise. The entry being processed is held in its own list now, so a raising D3DKMTQueryAdapterInfo closes that handle too.
  • Contradictory docstrings — reconciled at the source: the credit is the same optimistic figure _get_reclaimable_allocator_bytes refuses to grant, which is why it is loop-local, and the comment now says that plus what the consumers do instead of implying the bytes are dependable.
  • Classic-VAE MIOpen branch had no test — added, nine cases across the three estimators: MIOpen on ROCm, cuDNN on CUDA, and cuDNN for a cpu_only VAE on a ROCm build. Your 9.49 → 15.36 GB figure is the reason it earns a test.

Still open, deliberately

  • Anima's gate keeps its flat 2900. Swapping in the Qwen curve would raise its 1024px figure from ~6.1 GB to ~9.6 GB and move a threshold that was calibrated end-to-end on 8 GB cards (0.65 s untiled against 1.05 s tiled, and 7 s+ once the transformer is evicted). That is a measurement I have not made, so I would rather not move it blind.
  • Qwen encode has no gate. Agreed in principle; it needs its own measurement of when an encode actually fails, and encode tiling changes the latents rather than the pixels, so it is a separate change.
  • device=None means different things in the two estimator families. Left as is on purpose: aligning them moves the figure for every caller that passes nothing, and the last attempt at exactly that made the same call answer differently on my box and on CI. Both are documented at their branch.

Also fixed from your non-blocking list, after a second pass

An internal blocker review caught that your test_the_pretile_gate_uses_the_memory_the_device_will_keep_resident now needs other_process_local_bytes stubbed too, or it reads the real card on a ROCm host and flips on any foreign reading below ~1.2 GiB -- the same host-leakage class as R4. Stubbed, and that case is explicitly the fallback's case now. other_process_local_bytes also has its own unit tests (adapter-minus-own, the clamp at zero, and no reading), which is what the fake's new parameter is for.

Verification

  • Full suite, -n 6, INVOKEAI_DEVICE=cpu (the CI regime): 11440 passed. ruff check and format --check clean.
  • ROCm venv, focused suites: 1123 passed. -m slow on the RX 9060 XT: 5 passed, 1 skipped (no fused kernel for that shape on gfx1200).
  • End-to-end on the card, Z-Image nvfp4: 1024px pixel-identical to the merge-base (mean 0.000, max 0) at 36.8 s against 132-168 s at the merge-base; 1536px still pre-tiled, decode 2.3 s; shared system memory 0.28 GiB peak. The 0.06/255 drift against my older reference is upstream: the merge-base differs from it by the same 0.060.
  • Krea-2 fp8: 60-61 s, bit-identical run-to-run within a process, shared memory 0.34 GiB. Across processes it is not reproducible at all, which I only established because a cross-commit comparison looked alarming: same seed, two server starts on the same commit differ by 10.7/255 (max 223), and the second run then agreed with the merge-base to 0.12. So Krea-2 image comparisons across commits say nothing at that granularity on this path; Z-Image nvfp4 is the reference that holds. The gate did not tile either Krea-2 decode (1.43 s here against 19.07 s at the merge-base, where the unbounded score matrix pages).

@lstein lstein left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up review — round 3

All four requests from the last round are fixed, and I verified each by running the shipped code rather than reading the commit:

  • R1 — the ceiling is now total − foreign residency with the budget as fallback. The double-count is gone, and both directions you measured now behave: foreign 8 GiB holder with the budget misreporting 15.09 → ceiling 7.92 GiB → tiles; Invoke holding 12 GiB with the budget misreporting 7.62 → ceiling 15.91 GiB → does not tile.
  • R2 — per-backend curves with no extrapolation past the last measured area. The reported 2048²/24 GiB case tiles again. Full grid (512–2560 px × 8/12/16/24/48 GiB, both backends): no decode that fails to fit is left untiled.
  • R3 — symbol gone repo-wide; the LTX-2 slow test imports one that exists.
  • R4 — the fixture stubs find_spec to a synthetic version.py, and TestRocmBuildDetection still covers the hip = '7.14.60850' → True case, so both sources of truth are exercised. Confirmed the detector still does not import torch.

The docstring contradiction, the one-handle leak (current is a good fix) and the missing classic-VAE MIOpen coverage were all picked up too. Full suite here: 11476 passed, 0 failed — 7 → 3 → 0 across the three rounds.

One issue left, and it is narrow.


The new foreign-residency figure subtracts two different accounting bases

invokeai/backend/util/wddm.py:412

return max(0, adapter_bytes - (total_bytes - free_bytes))

adapter_bytes is \GPU Adapter Memory(*)\Dedicated Usage — a driver-side committed figure. total_bytes - free_bytes is torch's live-allocation figure. This module's header says those two are not the same thing, and measures the gap:

Not minus CurrentUsage: that also counts memory HIP keeps after a free and hands back to the next allocation (1-2 GiB measured, still held after empty_cache())

The adapter counter is the same driver-side view as CurrentUsage, so it should include HIP's retained pool while torch's figure does not. If so, the subtraction yields foreign + our own retained, not foreign. max(0, …) only clamps the negative side, so the over-count passes through: the ceiling drops by 1–2 GiB and the 0.9 line with it, on every decode after the first load/unload cycle, with nothing else on the card. Decodes landing in that band get tiled when they would have fitted — the non-pixel-identical output this whole line of work is about avoiding, reaching the same place by a different route.

The fix is already in the file: query D3DKMT's CurrentUsage for this process so both sides are on the driver basis. The header's reason for avoiding CurrentUsage — "it would hide what an offload just freed" — is about the headroom question, not about how much of the adapter other processes hold; those two genuinely want different bases.

I could not confirm this on Linux. It rests on Dedicated Usage behaving like CurrentUsage with respect to HIP's retained pool, which is an inference from your own measurement rather than a measurement of its own. You have the card and have already characterised the CurrentUsage gap, so this should be quick to settle either way — and if Dedicated Usage turns out to exclude the retained pool, say so in the docstring and I will drop it.

Two hardening items (not blocking)

No plausibility bound on the new reading, and no lower clamp on the ceiling. video_memory_budget validates 0 < Budget <= total; other_process_local_bytes validates nothing, and should_pretile_vae_decode does min(total, total - other) with no floor. A bogus reading is not hypothetical, because the per-item PDH CStatus is still unchecked (wddm.py:335), so an instance with invalid data contributes whatever largeValue happens to hold:

other=  0 GiB on a 16 GiB card -> ceiling=  16.00 GiB  512px decode (2.51 GiB) tiled=False
other= 20 GiB on a 16 GiB card -> ceiling=  -4.00 GiB  512px decode (2.51 GiB) tiled=True
other= 40 GiB on a 16 GiB card -> ceiling= -24.00 GiB  512px decode (2.51 GiB) tiled=True

A negative ceiling tiles every decode on the machine, and the debug line reports a negative GiB figure. Clamping the reading to total (or validating it the way the budget is) makes this unreachable.

paged_bytes is now coupled to the new counter. Both counters share one query and any PdhAddEnglishCounterW failure returns None for the whole query, which _pdh_attempted latches for the process lifetime. On a Windows build where \GPU Adapter Memory is unavailable, the shipped "Windows kept N GiB of Invoke's GPU memory in shared system memory" warning silently stops working — an unrelated feature disabled by the absence of a counter it does not use. Separate queries, or tolerating a missing counter per-counter, would keep them independent.

Not a blocker, to save you the trip

For completeness, since it looks alarming: CUDA 1792×1792 on a 16 GiB card flips tiled → untiled in this commit, and a 0.13% smaller 1728×1856 stays tiled.

1728x1856  const=2671 (interpolated)  gate=15.96G  tiled=True
1792x1792  const=2281 (measured)      gate=13.64G  tiled=False

I do not think this is the 2048² defect relocated. There the gate extrapolated a flat zero-headroom constant past the last measured point, into a region where calibration had run out of memory. Here it is using an actual measurement at the point it was measured — and if 2281 is right, that decode really does peak at 13.64 GiB and fits the card with room to spare, so tiling it would be the false positive we have been removing. The neighbours only look inconsistent because bracket-max interpolation is deliberately conservative around a real dip, the same shape as the ROCm 3273 dip at 1536². Worth knowing that one measurement is now load-bearing there with no smoothing, but I would not change it without evidence that 2281 is noise.

Checks

  • Full suite, -n 4: 11476 passed, 177 skipped, 8 xfailed, 0 failed.
  • tests/backend/util/test_wddm.py, test_vae_pretile.py, test_torch_cuda_allocator.py, test_vae_working_memory_backend.py: 53 passed.
  • ruff check / format --check on all nine changed files: clean.
  • The new pre-tile tests do catch the old bug: reverting only the ceiling logic fails test_another_process_holding_the_card_tiles_even_while_the_budget_looks_fine.
  • New wddm entry point confirmed inert off Windows (None, nothing at import time); counter paths and instance naming check out — \GPU Adapter Memory keyed by luid_… is adapter-wide, \GPU Process Memory keyed by pid_N_luid_… is per-process, and PdhAddEnglishCounterW is the right call on a localized Windows.

… counter

- The adapter-wide PDH counter cannot be used: measured with an 8 GiB holder, the holding process read 15.66 GiB and
  another read 0.176 GiB for the same instance at the same moment. It also shares the driver's basis, so subtracting
  torch's live figure invented ~2 GiB of foreign residency after every load/unload cycle
- The ceiling is min(total, max(budget, our CurrentUsage)): a budget below our own residency asks this process to
  trim itself, which the cache does for the reservation. A foreign holder that leaves the budget alone is documented
  as undetectable
- One PDH query per counter, per-item CStatus checked, usage clamped rather than dropped with the budget
- Interpolate decode peaks in bytes so a smaller image is never priced above a larger one
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

Your inference was right, and settling it on the card also killed my fix. Both halves below, with the measurements.

The basis mismatch is real

Exactly as you predicted. RX 9060 XT, 4 GiB allocated then freed and empty_cache()-ed, under expandable segments (the default this PR ships on Windows ROCm):

                           adapter-wide   CurrentUsage   torch live   reserved
held 4 GiB                         4.22           4.16         4.16       4.01
freed + empty_cache                2.17           2.13         0.15       0.00

The adapter counter and CurrentUsage track each other to ~40 MB throughout; torch's figure does not. So adapter − torch_live invented ~2 GiB of foreign residency after every load/unload cycle, dropping the ceiling and the line with it. Your reading of the module header against the new code was correct.

It is allocation-pattern dependent, which is why my earlier characterisation was too narrow: with the default allocator, one 4 GiB block retains nothing, four 1 GiB blocks retain 1.0 GiB, sixteen 256 MiB blocks retain 1.8 GiB. Under expandable segments it is 2.0 GiB in all three shapes.

But the adapter counter cannot be used at all

Before switching the subtraction to CurrentUsage, I checked whether the adapter figure itself is sound. It is not. With a second process holding 8 GiB, sampled within the same second, same counter, same instance key luid_0x00000000_0x0002221a_phys_0:

  • from the holding process: adapter-wide 15.66 GiB
  • from another process: adapter-wide 0.176 GiB (while its own CurrentUsage was 0.14)

So total − foreign has no reliable input on this platform, whichever basis I subtract. other_process_local_bytes and the adapter counter are gone again, along with the PDH coupling that came with them.

What ships instead

Only figures I could verify: the budget and this process's own driver-side usage, both from the one D3DKMT call.

ceiling = min(total_memory, max(budget, current_usage))

A budget below what this process already holds is Windows asking it to trim itself — which is exactly what the cache does to honour the reservation — so it must not drive tiling. Above that, the budget is the ceiling, which is your round-2 finding unchanged.

What it does in the states I measured:

state budget ours ceiling 7.0 GiB decode 15.8 GiB decode
idle 15.09 0.14 15.09 untiled tiled
12 GiB held, budget trimmed 7.62 12.17 12.17 untiled tiled
after a load/unload cycle 15.09 2.13 15.09 untiled tiled
foreign 8 GiB + 8.2 GiB ours 7.55 8.20 8.20 untiled tiled

Your round-2 test passes unchanged, and the mid-session case that made me write the credit in the first place stays fixed.

The limit is now documented rather than papered over: a foreign process that takes memory without moving our budget is invisible here. Windows keeps reporting 15.09 GiB next to an 8 GiB holder, torch's free figure ignores other processes by construction, and the adapter counter contradicts itself as above. That decode pages, and the end-of-session paging warning is what surfaces it. If you know a driver-side adapter figure that holds up cross-process — D3DKMTQueryStatistics segment data is the candidate I did not attempt — that would close it properly.

Hardening

  • Per-counter PDH queries. One query per counter path, so a build without a given counter disables only what reads it. Your point stands independently of the adapter counter being gone.
  • Per-item CStatus checked. Instances whose data PDH marks invalid are skipped instead of contributing whatever largeValue holds.
  • CurrentUsage is plausibility-bounded the way the budget already was, and the ceiling can no longer go negative: max(budget, usage) is at least one real figure, and min with the total caps it.

The 1792² neighbour you flagged

You were right not to touch the 2281 measurement, and right that bracket-max interpolation was what made the neighbours look wrong. Fixed at the interpolation instead: the bytes are interpolated between measured points rather than the per-pixel constant, since the constant dips while the bytes rise monotonically. Measured points keep their measured value, and past the last point the reservation constant still takes over.

1536x1536  11.74 GiB
1728x1856  13.63 GiB   (was 15.96)
1792x1792  13.64 GiB   measured
1920x1920  19.91 GiB
2048x2048  22.66 GiB -> tiled on a 24 GiB card

Monotonic now, and a test pins that a smaller image is never priced above a larger one.

Verification

  • Full suite, -n 6, INVOKEAI_DEVICE=cpu: 11750 passed. ruff check and format --check clean.
  • ROCm venv focused suites: 857 passed. -m slow on the card: 5 passed, 1 skipped.
  • The shipped rule probed on the card in four states (table above): 1024px never tiled, 1536px always tiled.
  • End-to-end Z-Image nvfp4: 1024px pixel-identical to the merge-base (max diff 0) at 36.3 s, decode 0.86 s; 1536px still pre-tiled at 178 s, decode 2.3 s; shared system memory 0.3 GiB peak.

@lstein lstein left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up review — round 4

Everything from the last round is resolved, and removing the adapter counter outright was the right call — measuring that it contradicts itself between processes is a better answer than the same-basis fix I suggested. Confirmed fixed:

  • The accounting-basis mismatch is gone with the counter.
  • local_video_memory returns both figures from one D3DKMT query, budget validated 0 < Budget <= total.
  • Per-item CStatus is checked — mutation-verified, re-admitting PDH_CSTATUS_INVALID_DATA brings the garbage reading back.
  • One query per counter path, so paged_bytes is no longer disabled by an unrelated counter.
  • Monotonic pricing: 0 area-ordered inversions across 1089 (H,W) pairs on both backends, and the regression sweep shows no untiled→tiled flips on any card, so the pricing change costs no pixel identity.

One regime is still wrong, and it is narrow — this is not a reopened round.


The max(budget, usage) override fails open in two reachable states

invokeai/backend/util/vae_working_memory.py:51

total_bytes = min(total_bytes, max(budget_bytes, usage_bytes))

The override reads a budget below our own usage as "Windows is asking this process to trim itself, and trimming recovers that much". That is true when we are the only pressure on the card. It is not true when someone else is, and CurrentUsage cannot tell the two apart — it is commitment, and wddm.py:16-17 notes it "counts paged bytes as local".

Trigger 1 — a foreign holder that does move the budget. 16 GiB card, foreign process holding 8 GiB, Invoke committed 12 GiB with part of it already paged out:

budget 7.6, usage 12.0, Qwen 1024px   ceiling=12.00G  price=8.93G  tiles=False
  -> evicting every model Invoke owns still leaves only ~7.9G (16 - 8 foreign)
parent min(total, budget)             ceiling= 7.60G  price=8.93G  tiles=True

The decode pages. This is a regression against the parent commit and against the pre-round-3 min(total, budget) rule, and it is not the case the new docstring excuses — the "known limit" is a holder that leaves the budget alone, whereas here the budget moved correctly and the override discarded it.

Trigger 2 — the clamp maximizes the ceiling in exactly the paging state. wddm.py:303:

return int(info.Budget), max(0, min(int(info.CurrentUsage), adapter.total_bytes))
CurrentUsage=  12.00G -> usage=12.00G -> ceiling=12.00G
CurrentUsage=  17.00G -> usage=15.92G -> ceiling=15.92G   <- correction disabled
CurrentUsage= garbage -> usage=15.92G -> ceiling=15.92G   <- correction disabled

The comment at wddm.py:300-302 states this state is reachable, and it is the over-commit state — i.e. the correction switches itself off precisely when Windows is already paging. Budget gets a plausibility check; CurrentUsage gets only a clamp, and the clamp pushes the ceiling to the nameplate rather than rejecting the reading. This trigger does not depend on any contested measurement.

Direction, not a prescription: bound the usage override by something actually resident. paged_bytes() is already read once per session for the warning and is the per-process counter this commit still trusts, so CurrentUsage - paged is a residency figure on the same basis. And CurrentUsage > total looks like it should drop the override while keeping the budget for cuda_mem_get_info, rather than clamp to the most optimistic value.

While you are in there: the budget's behaviour is now documented both ways

Trigger 1's reachability depends on whether a foreign process lowers the budget, and the PR currently asserts both answers:

  • invokeai/backend/util/wddm.py:10-11 — "next to another GPU process, Windows lowers the budget within about a second"
  • PR description, item 2 — "15.09 of 15.92 GiB alone; lowered within a second when another GPU process starts"
  • invokeai/backend/util/vae_working_memory.py:41-42 (new) — "Windows keeps reporting 15.09 GiB next to an 8 GiB holder"

The round-4 design rests on the third while the module header still asserts the first. Whichever matches the card, the other two should follow — and if the budget really does track foreign pressure, that is an argument for trusting it over the usage override rather than the reverse.

Non-blocking

  • test_a_counter_this_build_does_not_have_only_disables_what_reads_it (tests/backend/util/test_wddm.py:248) passes unchanged on the parent commit — after the adapter counter is removed there is only one counter path left, so there is no second reader for it to protect. It documents an intent rather than guarding behaviour.
  • tests/backend/util/test_vae_pretile.py:85,98 still monkeypatch torch.cuda.mem_get_info, which the gate no longer calls. Dead setup that reads as coverage.
  • The pricing change is larger than "fix non-monotonicity": on ROCm/16 GiB, 197 sizes move from tiled to untiled, 134 of them landing more than 1 GiB under the line (worst 1600x1472, 20.05 → 14.37 GiB). The direction is intended and nothing became newly tiled, but the 1024²→1536² chord spans the flattest measured bracket, so that band is the least evidenced part of the curve. A running max over the bracket-max prices would have removed the inversions without lowering any price, if you would rather not carry that.
  • No test covers either blocker regime, and monotonicity is asserted for one CUDA pair with no ROCm counterpart (ROCm had its own inversion on the parent: 1600x1472 20.05 GiB vs 1024x2304 14.38 GiB).

Checks

  • Full suite, -n 4: 11759 passed, 177 skipped, 8 xfailed, 0 failed.
  • Targeted: test_wddm.py, test_vae_pretile.py, test_torch_cuda_allocator.py, test_vae_working_memory_backend.py, test_qwen_image_working_memory.py, test_anima_vae.py, test_flux2_working_memory.py, model_cache/ → 731 passed, no failures.
  • All 16 CI checks green.
  • ruff check / format --check on the five changed files: clean.
  • test_a_budget_below_our_own_residency_does_not_tile_a_decode_that_fits genuinely fails if the ceiling reverts to the bare budget, so the new coverage is load-bearing.
  • New local_video_memory confirmed inert off Windows (None, nothing at import time).

…dows has paged out

Only what this process keeps in VRAM (CurrentUsage minus its paged bytes) may lift the ceiling above the WDDM budget; usage beyond the card or an unknown paged amount leaves the budget in charge.
The measured budget behaviour next to a foreign holder now reads the same in wddm.py and the gate docstring.
Tests cover the foreign-holder, over-commit and unknown-paged states and area monotonicity of the Qwen pricing on both backends.
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

Thanks — both triggers were real, and measuring the budget next to a holder settled the other point too.

The override is bounded by residency now

should_pretile_vae_decode starts from the budget and lets only what this process keeps resident lift it: CurrentUsage − paged_bytes(), as you suggested. Where either figure is unknown, the budget decides alone:

  • CurrentUsage beyond the adapter: local_video_memory now returns (budget, None) instead of clamping. cuda_mem_get_info keeps the budget.
  • No PDH reading.
State budget / usage / paged ceiling Qwen 1024px (8.93 GiB)
Alone, trimmed (round-3 case) 7.6 / 12.0 / 0 12.0 untiled (7.0 GiB Z-Image decode stays untiled)
Trigger 1, foreign holder 7.6 / 12.0 / 4.3 7.7 tiles
Trigger 2, over-commit 7.6 / >card / – 7.6 tiles
Paged unknown 7.6 / 12.0 / – 7.6 tiles

The residency figure is checked on the card: an RX 9060 XT next to an 8 GiB holder, this process having allocated 9.5 GiB.

  • CurrentUsage 9.65 GiB, paged_bytes 2.01 GiB → 7.64 GiB resident.
  • PDH process_dedicated 7.70 GiB.
  • The gate tiles the Qwen decode.

The budget, measured, and now described the same way everywhere

An 8 GiB holder, then a second process filling in 256 MiB steps (torch 2.12+rocm7.14):

  • While the filler holds little: its budget stays at 15.09 GiB, up to 3.25 GiB of its own.
  • From 3.5 GiB of its own: it steps down, 11.51 → 11.01 → 10.51 → 10.01 → 9.26 → 8.51 → 7.76 → 7.58.
  • From 7.75 GiB: the filler starts to page, with its dedicated memory flat at 7.70 GiB while CurrentUsage keeps growing.

So both earlier statements were half true. The budget does track foreign pressure, but only once this process grows into it. By the time it could page, the budget is its share of the card, which is why trusting it over a raw usage override is right. The wddm.py header and the gate docstring now say exactly this, including the alone case, where it drops to 7.62 GiB at ~12 GiB without paging. The PR description's item 2 is corrected as well. The remaining known limit is a decode sized while this process still holds little next to a big holder; it is stated in the docstring, and the paging warning surfaces it.

Non-blocking items

  • test_a_counter_this_build_does_not_have_only_disables_what_reads_it → replaced by test_a_missing_counter_answers_unknown, which guards what is still behaviour: no counter means unknown, not an error. The stale per-counter comment is gone too.
  • Dead torch.cuda.mem_get_info / video_memory_budget patches in test_vae_pretile.py → removed.
  • Regime tests → one each for trigger 1, trigger 2 and the unknown-paged case. All three fail on the previous commit; the over-commit one raises a TypeError in the old max.
  • Monotonicity → the ROCm pair (1600×1472 vs 1024×2304) is added, plus a property test over a 33×33 grid (512–2560 px) for both backends: no size is priced below a smaller one.
  • Pricing → I kept the chord. At every measured point it returns the measured peak, and it is monotone by construction because the measured bytes rise with area. The running max over bracket maxima would keep prices the measurements show to be too high, e.g. 1600×1472 at 20.05 GiB against a measured 14.4 GiB at 1536². The flattest bracket (1024²→1536² on ROCm) is the least evidenced part, as you say. The reservation there still uses the flat 5500, so an under-priced decode still gets the headroom; only the tiling decision follows the curve. The description now states the scale: 197 sizes on ROCm/16 GiB move from tiled to untiled, and none become newly tiled.

Checks

  • origin/main merged in (conflict-free; none of the files this PR touches changed on main).
  • Full suite -n 6: 11,787 passed, 184 skipped, 9 xfailed.
  • Targeted (test_wddm, test_vae_pretile, test_qwen_image_working_memory, allocator, Anima VAE, FLUX.2 working memory, model cache): 729 passed.
  • ruff check and format --check: clean.

# Conflicts:
#	invokeai/app/invocations/vae/flux_vae_decode.py
#	invokeai/app/invocations/vae/flux_vae_encode.py
#	invokeai/app/invocations/vae/wan_latents_to_video.py
#	invokeai/app/invocations/vae/z_image_image_to_latents.py
#	invokeai/app/invocations/vae/z_image_latents_to_image.py
#	invokeai/backend/util/vae_working_memory.py
#	tests/app/invocations/test_anima_vae.py
#	tests/app/invocations/test_flux_vae_decode_tiling.py
#	tests/app/invocations/test_wan_working_memory.py
@Pfannkuchensack
Pfannkuchensack changed the base branch from main to upstream-merge October 1, 2026 20:12
…indows-memory

Carries the branch's low-VRAM docs and regenerated API contracts to their moved locations under configuration/Optimization and frontend/api.
Keeps the corrected force_tiled_decode description next to auto_tiled_decode and regenerates the API contracts.
Folds the automatic-tiling docs into the existing Tiled VAE decode and encode section and fixes the moved page's system-requirements link.
@Pfannkuchensack
Pfannkuchensack merged commit 59b48c7 into upstream-merge Oct 1, 2026
13 checks passed
@Pfannkuchensack
Pfannkuchensack deleted the feat/rocm-windows-memory branch October 1, 2026 21:18
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants