Repository navigation
ROCm on Windows: keep VRAM resident and bound attention and VAE memory - #304
Conversation
With pytorch_cuda_alloc_conf unset, a ROCm build on Windows now runs with expandable_segments:True, so fragmented VRAM no longer pushes allocations into shared system memory. An explicit allocator setting or allocator env var still wins; docs and generated config artifacts updated.
On ROCm under Windows, free VRAM is now capped by the WDDM budget (gdi32 D3DKMTQueryVideoMemoryInfo), past which Windows pages allocations into system memory instead of failing them. torch's figure there ignores other processes and runs up to the physical total; other platforms are unchanged.
After each session, a worker on a Windows ROCm device reads the process's shared GPU memory (PDH) and warns once per episode when more than 512 MiB of it sits in system memory, where every generation that touches it slows down. Adds a low-VRAM docs section on AMD GPUs under Windows.
The ROCm SDPA guard now splits any math-kernel call over 1 GiB of scores into head groups (bitwise identical) or query rows (one oversized head), and working-memory estimates price one chunk instead of the whole matrix. Z-Image 1024px peaks at 0.5 GiB instead of 8 GiB per attention on an RX 9060 XT; renamed to install_rocm_sdpa_guard.
The Krea-2 working-memory estimate now adds the score matrix where the build materializes it (no fused kernel), priced from the loaded model's heads and the attended sequence; one 1 GiB chunk on ROCm, nothing on CUDA. Before, Krea-2 fp8 on an RX 9060 XT loaded the whole transformer and paged each 12.9 GiB attention into system RAM.
On ROCm the FLUX.1 VAE (also Z-Image's) now budgets 3600 decode / 2750 encode bytes per pixel·byte instead of 2200/1100, measured 3451/2688 on an RX 9060 XT: the same convolution stack and numbers as the FLUX.2 VAE, so both share one table. The estimate follows the VAE's compute device, so a cpu_only VAE keeps the cuDNN column.
…code The FLUX.1, Z-Image, Qwen-Image (Krea-2) and Wan decodes now tile when the untiled estimate exceeds 90% of the VAE's GPU memory (Anima keeps its measured 70%), instead of relying on an OOM that Windows paging never raises. New auto_tiled_decode setting (default on) turns it off; force_tiled_decode and the tiled field still win.
- Cap free VRAM at the WDDM budget minus live allocations: HIP keeps freed memory that CurrentUsage still counts - Empty the allocator cache after each offload under expandable segments so the cache stops once enough is free - Hardware tests for the headroom after a free and for the PDH lookup
- Skip the allocator default once torch is imported; stop the test fixture leaking it - Warn only about memory paged across two sessions, with actionable advice - Leave K/V broadcasts the chunks cannot slice whole; cover masked row chunks and Krea-2 CFG/regional pricing
- Price and pre-tile FLUX.1 decodes with a diffusers-layout VAE - Log up-front tiling at debug; document what auto_tiled_decode off means - Emit POSIX path defaults in the generated docs settings
|
I can verify this on Linux ROCm, but I don't have a Windows boot on my AMD rig. |
|
Passes all the ROCm tests, including |
These tests fail on purpose. They pin behaviour the current up-front tiling rule does not have, and should go green with the fix. - A Qwen-Image/Krea-2 decode whose measured peak fits the card is tiled anyway, because the gate is fed the padded reservation figure (a flat 5500 B/pixel-byte) instead of the expected peak (3273 measured at 1536px). Tiling is not pixel-identical, so this silently changes output for images that would have decoded in a single pass. - should_pretile_vae_decode compares against the card's nameplate total, not the Windows video-memory budget this PR added TorchDevice.cuda_mem_get_info for, so a decode Windows will page into system memory is left untiled.
Adversarial reviewReviewed against I've pushed 33e16af with three intentionally-failing tests for the two findings I think are blockers. CI will be red by design — they pin behaviour the current rule does not have and should go green with the fix. Happy to drop that commit if you'd rather not carry red tests. 1. Reservation headroom is reused as a tiling trigger, so decodes that fit are silently tiled
The estimator constants are deliberately over-provisioned ("max observed + ~8% headroom"). Over-reserving used to cost only cache eviction. Comparing that padded figure against a 90% line converts the conservatism into a non-pixel-identical output change, with no error and nothing above DEBUG in the log. Qwen-Image is the sharp case, because its ROCm decode constant is a flat
On a 24 GiB ROCm card a 1536² Qwen-Image decode that previously ran untiled and succeeded now tiles. Krea-2 decodes through this node too. On CUDA FLUX.1 the crossovers are benign (shipped 2200 vs measured 2185), so this is specifically Qwen/Krea-2 on MIOpen. The reservation is right as it is — it's the tiling decision that shouldn't be carrying reservation headroom. Test: One knock-on: 2. The gate measures total VRAM on the platform this PR taught the cache not to trust
Trigger: 16 GiB RX 9060 XT, a browser holding ~3 GiB so the budget is ~12.5 GiB. FLUX.1/Z-Image at 1408² estimates 13.29 GiB — under the 14.4 GiB threshold, over the budget. No pre-tile, no OOM (Windows pages instead), so the retry at Also reachable on Linux/CUDA: Test: Other material findings (no tests pushed)3. The expandable-segments offload fix is inert on multi-GPU. 4. The "bitwise identical" claim for head-group chunking is false. 5. 6. 7. SD1/SDXL, SD3 and CogView4 left on the cuDNN constant. Smaller
Verified as claimed
I could not complete a full-suite run — this box was saturated and |
…droom - Compare against the Windows video-memory budget where there is one, not the card's total - Qwen-Image ROCm constants now follow the bounded math attention this branch ships: measured 2650-2772 decode and 1541-1552 encode across 512-2048px, against 5500/6300 fitted while a call built its whole score matrix - Calibrate with the attention guard installed, so the script measures the path the app runs
…ld up - Credit freed bytes while a peer device defers empty_cache, so multi-GPU stops over-unloading - Never let the budget or paged-bytes lookups raise, and close adapter handles if enumeration throws - Detect a ROCm build from torch's version.py too, so a locally built wheel gets the allocator default - Correct the bitwise-identical claim for head-group chunking
…kend - auto_tiled_decode no longer disables Anima's own rule, which is a speed optimization, not an OOM fallback - SD1/SDXL, SD3 and CogView4 take the MIOpen constant on ROCm, like the FLUX.1 autoencoder they share a stack with - Encode nodes place their tensors on the VAE's device, matching the decode nodes
…eiling - Keep the padded Qwen-Image reservations; the tiling decision reads the measured curve instead, so a decode whose real peak fits the card is no longer tiled by a reservation's headroom - Compare against the video-memory budget plus what this process can release: Windows halves the budget once the process passes ~12 of 16 GiB, and the bare figure tiled a 7.0 GiB decode that fits
|
Thanks — this was a genuinely useful review, and the two tests made the first finding much easier to act on. All three are green now, and everything else you raised is either fixed or answered below. One of my first attempts at finding 1 was wrong in a way worth recording, so I've left that in rather than quietly dropping it. 1. Reservation headroom as a tiling trigger — fixed the way you framed itImplemented as you described it: the reservation keeps its headroom, and the tiling decision no longer carries it. I went the wrong way first and want to record why, since it nearly shipped: I re-measured on an RX 9060 XT (gfx1200, torch 2.12+rocm7.14, fp16) and got a flat 2650-2772 decode / 1541-1552 encode across 512²..2048², attributed that to the score-matrix chunking this branch adds, and made the constants conditional on the guard. That attribution does not survive your own table. At 512² the gap between the two cards is 1.24 GB while the entire score matrix there is ~0.3 GB, and at 1536² an unbounded matrix would be ~23 GB against a measured 15.4 GB total — so the W7900 run was never dominated by an unbounded score matrix in the first place. The spread is the card, not the guard, and shipping the lower pair would have under-reserved by 1.9x (decode) to 3.8x (encode) on the hardware the shipped figures were measured on, with Qwen Image Edit's encode having no tiled retry to fall back on. Reverted; the constants are unchanged. The gfx1200 numbers stay in the comment, because that spread is the actual argument for your finding: a reservation that has to cover a 1.9x card difference is not a figure to decide non-pixel-identical output with. One thing I did keep: 2. The gate measured the nameplate total — fixedThe gate consults the budget now. My first version compared against the bare budget, and an end-to-end run on the card caught what that does mid-session: Windows holds the budget at 15.09 GiB until the process passes about 12 of 16 GiB and then halves it to 7.62 — below what we already hold. With models resident, the gate read 7.6 GiB, drew a 6.9 GiB line, and tiled a 7.0 GiB Z-Image decode that fits. Output stopped being pixel-identical (PSNR 42 dB) and the decode went from 0.89 s to 1.5 s: your finding 1, reintroduced at the other end. So the ceiling is the budget plus what this process itself holds, capped at the card: the cache evicts models to honour the reservation, so that memory is available to the decode. Measured:
Your test passes unchanged (it patches The 3. Offload fix inert on multi-GPU — fixed
4. "Bitwise identical" — corrected, thank youYou are right, and the CPU test was validating the claim on the one platform it isn't about. The comment and the docstring now say what actually holds: exact in exact arithmetic, bitwise where the kernel's arithmetic does not depend on the batch it runs over (CPU), and within a bf16 ulp on a card whose batched GEMM re-tiles with the group count — with your 0.0005 measurement named. The test keeps its exact assertion and says why it is exact there. The PR description is updated too. 5.
|
|
@lstein can you test again? |
lstein
left a comment
There was a problem hiding this comment.
Follow-up review
Thanks — all seven findings from the last round are genuinely addressed, and splitting the tiling decision from the reservation is a better shape than what I suggested. The 4 pre-existing FLUX.2 failures are fixed too, and both generated artifacts still regenerate byte-identical.
Two of the fixes overshoot, though: R1 and R2 each make a specific case worse than your own immediately preceding commit, and neither is caught by CI or by the new tests. Those plus the slow-test ImportError are what I'd like resolved before merge.
Verified on this box (torch 2.13.0+rocm7.2) by running the shipped functions, not by reading.
R1 — the budget ceiling cancels the budget out, and regresses the case it was written for
invokeai/backend/util/vae_working_memory.py:47 (feb4d8e)
total_bytes = min(total_bytes, budget_bytes + (total_bytes - free_bytes))total_bytes - free_bytes is this process's live allocations — which the budget already bounds — so adding them back double-counts. Once our usage reaches total - budget, the budget has zero effect:
ours= 0.00 GiB -> effective ceiling 7.62 GiB
ours= 4.00 GiB -> effective ceiling 11.62 GiB
ours= 8.29 GiB -> effective ceiling 15.91 GiB <-- budget no longer participates
The credit assumes releasing our memory raises the budget. That holds when we depressed it, but not when another process did — and that is the case the gate exists for. RX 9060 XT 16 GB, a second GPU program takes 8 GiB so Windows trims the budget to 7.62 GiB, Invoke holds 3.0 GiB:
budget=7.62G ours=3.00G -> ceiling=10.62G line=9.56G
Qwen 1024px peak = 8.93G
NEW code tiles? False
OLD (min(total,budget)) would tile? True
TRUE ceiling after evicting everything = 7.62G -> decode fits? False
At 40a78452df this tiled. Now it does not, Windows pages the decode, no OOM is raised, and the Qwen node has no retry — the generation just crawls, which is the exact failure mode this PR exists to remove.
The intent is right and the measurement behind it is real (15.09 GiB un-trimmed vs 7.62 trimmed). It just needs to be expressed as a restored budget rather than a credit for our own usage — something closer to min(total, max(budget, untrimmed_budget)), remembering the un-trimmed figure from when the process was small.
Note also that test_the_gate_counts_memory_this_process_can_still_release sits in the clamped regime (ceiling pinned to the full card), so "ignore the budget entirely" passes it too. It discriminates against the old code but not against this fix.
R2 — the CUDA tiling constant has zero headroom and extrapolates through an OOM
invokeai/backend/util/vae_working_memory.py:697
_QWEN_VAE_MEASURED_DECODE_PEAK_CUDA = 2690 is exactly max(2660, 2519, 2690, 2671, 2281) — the maximum observation, not a bound — and the calibration table's 2048² cell for CUDA decode reads (OOM), so the flat constant extrapolates straight through the point the calibration could not complete. On a 24 GiB 4090:
px gate peak reservation line tiled?
1792 17.28G 18.63G 22.85G new=False prev-commit=False
1920 19.83G 21.38G 22.85G new=False prev-commit=False
2048 22.57G 24.33G 22.85G new=False prev-commit=True
2176 25.47G 27.46G 22.85G new=True prev-commit=True
At 2048² the gate now leaves it untiled by 280 MB while the reservation (24.33 GB) exceeds the whole card, so the cache logs the shortfall and streams, and the decode then peaks at ≥22.6 GB alongside the context, VAE weights and latents. This is not worse than main (which never pre-tiled), but the PR added the protection and this commit removes it at the one CUDA resolution its own data says fails, on a node with no OOM retry.
Simplest fix: give CUDA the same per-point table as ROCm, or at minimum pad the constant and stop extrapolating past 1792².
R3 — slow LTX-2 tests are broken by the guard rename
tests/backend/ltx2/test_ltx2_real_files.py:197-199 still imports and calls install_rocm_sdpa_head_dim_guard, renamed here to install_rocm_sdpa_guard:
ImportError: cannot import name 'install_rocm_sdpa_head_dim_guard' from 'invokeai.backend.util.attention'
Four tests call _accelerator(). The file is pytestmark = pytest.mark.slow, so neither CI nor the default suite sees it — it only fires on an accelerator box under -m slow, which is the lane this PR's own hardware tests use. main added the caller, this branch did the rename, and the merge did not reconcile them.
R4 — the allocator-default tests are non-hermetic and fail on any ROCm host
The three remaining failures in my full run:
FAILED tests/app/util/test_torch_cuda_allocator.py::TestRocmWindowsAllocatorDefault::test_other_installs_are_left_alone[windows-cuda]
FAILED ...[windows-cpu]
FAILED ...[no-torch-metadata]
_installed_torch_is_rocm() now has two sources of truth, and the fixture stubs only the first:
_installed_torch_version() patched to '2.7.1+cu128'
_installed_torch_is_rocm() = True <- version.py unpatched: hip: Optional[str] = '7.2.53211'
The product logic is correct, and I confirmed it does not import torch (the before-torch-import contract holds). The problem is the test: it is red on every ROCm developer machine — on a ROCm PR — and on CUDA/CPU CI the new version.py branch never returns True, so the locally-built-wheel case it was added for has no passing coverage. Stubbing the file lookup the way TestRocmBuildDetection does fixes both halves.
Non-blocking
- Contradictory docstrings about the same bytes.
model_cache.py:2342says a large allocation cannot use holes in an expandable segment, citing a measured hard OOM;:2557says the allocator reuses them for the load this is making room for. One has to be wrong. The practical risk is bounded — the credit is loop-local (2563-2585) and both consumers re-measure honestly (lock()at 2109,make_room_in_vramat 3061), so a wrong credit means under-offloading that the streaming fallback still catches — but for the same reason the stated benefit only partly reaches the load. - Handle leak fix is one short.
wddm.py:226info = pending.pop()removes the entry before the body runs, so ifD3DKMTQueryAdapterInforaises, that handle is in neithermatchesnorpendingand theexceptat 241-246 misses exactly the one that failed — despite the comment claiming otherwise. Bounded to one per device per process. - Anima's gate still uses a flat unmeasured 2900 (
anima_latents_to_image.py:112) while sharing the VAE whose measured curve this delta just added;qwen_image_untiled_decode_peak_byteslooks like the right input there. - The classic-VAE MIOpen branch has no test. Nothing asserts that
estimate_vae_working_memory_sd15_sdxl/_sd3/_cogview4take the MIOpen column on ROCm, or that acpu_onlyVAE takes cuDNN. Worth covering, because the change is significant: an SD1.5 1024px fp32 decode reservation goes 9.49 GB → 15.36 GB on ROCm, which on an 8 GiB card forces full eviction on every decode. sdpa_score_matrix_bytes's unconditional chunk cap is now documented rather than conditioned — still latent, no current caller prices a causal/3-D shape.- The Qwen encode path has no gate at all, though the ROCm encode constant (6300) is the largest in the module and the "Windows never raises OOM" premise applies identically.
estimate_vae_working_memory_qwen_imagetreatsdevice=Noneas build-wide while the FLUX family resolves the session device — harmless today since every production caller passes a device, but the two neighbouring functions now mean different things by the same omission.
Checks
- Full suite,
-n 4: 3 failed, 11201 passed, 177 skipped, 8 xfailed — the 3 are R4. Previous round was 7 failed / 8959 passed. tests/app/invocations/test_flux2_working_memory.py: 149 passed (the 4 pre-existing failures are fixed).ruff check/format --checkon the changed files: clean.scripts/generate_docs_json.pyandscripts/generate_openapi_schema.pyboth reproduce the committed files byte-identically;schema.tsmatches.- My two regression tests from the last round now pass.
- The ceiling is the card minus foreign residency (new wddm.other_process_local_bytes): measured on an RX 9060 XT, the budget reads 15.09 GiB next to an 8 GiB holder with ~7.7 GiB available, and 7.62 GiB when Invoke itself holds 12 GiB and eviction would free the card. The budget stays as the fallback - Both backends get their measured decode curve; past the last measured area the decision uses the reservation constant instead of extrapolating through a calibration point that ran out of memory - Make the allocator and pre-tile tests hermetic on a ROCm host, rename the guard call in the LTX-2 slow test, and cover the MIOpen branch of the SD/SD3/CogView4 estimators
|
All four are fixed. R1 took a measurement round to get right, and it ended somewhere other than either of our proposals — details below, with the numbers. R1 — right diagnosis, and the fix is neither a credit nor a remembered budgetYou are right that Your suggested shape (remember the un-trimmed budget) does not survive the measurement either, though. I probed the budget against adapter-wide usage on the RX 9060 XT (16 GB) while a second process held 8 GiB: In the state your case describes — the foreign 8 GiB present, Invoke small — the budget is not trimmed at all: it reads 15.09 GiB while only ~7.7 GiB can be had. Windows trims it as our own residency grows, not when someone else takes memory. So there is no un-trimmed figure to restore: the un-trimmed value is exactly the one that is wrong there. Both readings mislead, in opposite directions:
What the decision actually needs is what other processes hold, which neither figure answers. Two things fell out of the measurement that are now in the code:
Three tests now pin it: the foreign-holder case (tiles although the budget looks fine), the trimmed-by-our-own-residency case (does not tile), and the fallback. You were right that the old test discriminated against the old code only; the first of these is the one that fails for "ignore the budget entirely". R2 — CUDA has its own curve now, and nothing extrapolatesBoth backends carry their measured per-point table. CUDA's grid ends at 1792² because 2048² ran out of memory when it was calibrated, and past the last measured point the decision falls back to the reservation constant rather than the maximum seen — "we have not measured this" has to mean the conservative figure, not a flat carry. On a 24 GiB card: 2048² tiles again, and the gate figure no longer sits under the line while its own reservation exceeds the card. R3 — guard rename reconciled
R4 — allocator tests are hermeticThe fixture stubs both sources now, the file lookup the way Non-blocking, also done
Still open, deliberately
Also fixed from your non-blocking list, after a second passAn internal blocker review caught that your Verification
|
lstein
left a comment
There was a problem hiding this comment.
Follow-up review — round 3
All four requests from the last round are fixed, and I verified each by running the shipped code rather than reading the commit:
- R1 — the ceiling is now
total − foreign residencywith the budget as fallback. The double-count is gone, and both directions you measured now behave: foreign 8 GiB holder with the budget misreporting 15.09 → ceiling 7.92 GiB → tiles; Invoke holding 12 GiB with the budget misreporting 7.62 → ceiling 15.91 GiB → does not tile. - R2 — per-backend curves with no extrapolation past the last measured area. The reported 2048²/24 GiB case tiles again. Full grid (512–2560 px × 8/12/16/24/48 GiB, both backends): no decode that fails to fit is left untiled.
- R3 — symbol gone repo-wide; the LTX-2 slow test imports one that exists.
- R4 — the fixture stubs
find_specto a syntheticversion.py, andTestRocmBuildDetectionstill covers thehip = '7.14.60850'→Truecase, so both sources of truth are exercised. Confirmed the detector still does not import torch.
The docstring contradiction, the one-handle leak (current is a good fix) and the missing classic-VAE MIOpen coverage were all picked up too. Full suite here: 11476 passed, 0 failed — 7 → 3 → 0 across the three rounds.
One issue left, and it is narrow.
The new foreign-residency figure subtracts two different accounting bases
invokeai/backend/util/wddm.py:412
return max(0, adapter_bytes - (total_bytes - free_bytes))adapter_bytes is \GPU Adapter Memory(*)\Dedicated Usage — a driver-side committed figure. total_bytes - free_bytes is torch's live-allocation figure. This module's header says those two are not the same thing, and measures the gap:
Not minus
CurrentUsage: that also counts memory HIP keeps after a free and hands back to the next allocation (1-2 GiB measured, still held afterempty_cache())
The adapter counter is the same driver-side view as CurrentUsage, so it should include HIP's retained pool while torch's figure does not. If so, the subtraction yields foreign + our own retained, not foreign. max(0, …) only clamps the negative side, so the over-count passes through: the ceiling drops by 1–2 GiB and the 0.9 line with it, on every decode after the first load/unload cycle, with nothing else on the card. Decodes landing in that band get tiled when they would have fitted — the non-pixel-identical output this whole line of work is about avoiding, reaching the same place by a different route.
The fix is already in the file: query D3DKMT's CurrentUsage for this process so both sides are on the driver basis. The header's reason for avoiding CurrentUsage — "it would hide what an offload just freed" — is about the headroom question, not about how much of the adapter other processes hold; those two genuinely want different bases.
I could not confirm this on Linux. It rests on Dedicated Usage behaving like CurrentUsage with respect to HIP's retained pool, which is an inference from your own measurement rather than a measurement of its own. You have the card and have already characterised the CurrentUsage gap, so this should be quick to settle either way — and if Dedicated Usage turns out to exclude the retained pool, say so in the docstring and I will drop it.
Two hardening items (not blocking)
No plausibility bound on the new reading, and no lower clamp on the ceiling. video_memory_budget validates 0 < Budget <= total; other_process_local_bytes validates nothing, and should_pretile_vae_decode does min(total, total - other) with no floor. A bogus reading is not hypothetical, because the per-item PDH CStatus is still unchecked (wddm.py:335), so an instance with invalid data contributes whatever largeValue happens to hold:
other= 0 GiB on a 16 GiB card -> ceiling= 16.00 GiB 512px decode (2.51 GiB) tiled=False
other= 20 GiB on a 16 GiB card -> ceiling= -4.00 GiB 512px decode (2.51 GiB) tiled=True
other= 40 GiB on a 16 GiB card -> ceiling= -24.00 GiB 512px decode (2.51 GiB) tiled=True
A negative ceiling tiles every decode on the machine, and the debug line reports a negative GiB figure. Clamping the reading to total (or validating it the way the budget is) makes this unreachable.
paged_bytes is now coupled to the new counter. Both counters share one query and any PdhAddEnglishCounterW failure returns None for the whole query, which _pdh_attempted latches for the process lifetime. On a Windows build where \GPU Adapter Memory is unavailable, the shipped "Windows kept N GiB of Invoke's GPU memory in shared system memory" warning silently stops working — an unrelated feature disabled by the absence of a counter it does not use. Separate queries, or tolerating a missing counter per-counter, would keep them independent.
Not a blocker, to save you the trip
For completeness, since it looks alarming: CUDA 1792×1792 on a 16 GiB card flips tiled → untiled in this commit, and a 0.13% smaller 1728×1856 stays tiled.
1728x1856 const=2671 (interpolated) gate=15.96G tiled=True
1792x1792 const=2281 (measured) gate=13.64G tiled=False
I do not think this is the 2048² defect relocated. There the gate extrapolated a flat zero-headroom constant past the last measured point, into a region where calibration had run out of memory. Here it is using an actual measurement at the point it was measured — and if 2281 is right, that decode really does peak at 13.64 GiB and fits the card with room to spare, so tiling it would be the false positive we have been removing. The neighbours only look inconsistent because bracket-max interpolation is deliberately conservative around a real dip, the same shape as the ROCm 3273 dip at 1536². Worth knowing that one measurement is now load-bearing there with no smoothing, but I would not change it without evidence that 2281 is noise.
Checks
- Full suite,
-n 4: 11476 passed, 177 skipped, 8 xfailed, 0 failed. tests/backend/util/test_wddm.py,test_vae_pretile.py,test_torch_cuda_allocator.py,test_vae_working_memory_backend.py: 53 passed.ruff check/format --checkon all nine changed files: clean.- The new pre-tile tests do catch the old bug: reverting only the ceiling logic fails
test_another_process_holding_the_card_tiles_even_while_the_budget_looks_fine. - New
wddmentry point confirmed inert off Windows (None, nothing at import time); counter paths and instance naming check out —\GPU Adapter Memorykeyed byluid_…is adapter-wide,\GPU Process Memorykeyed bypid_N_luid_…is per-process, andPdhAddEnglishCounterWis the right call on a localized Windows.
… counter - The adapter-wide PDH counter cannot be used: measured with an 8 GiB holder, the holding process read 15.66 GiB and another read 0.176 GiB for the same instance at the same moment. It also shares the driver's basis, so subtracting torch's live figure invented ~2 GiB of foreign residency after every load/unload cycle - The ceiling is min(total, max(budget, our CurrentUsage)): a budget below our own residency asks this process to trim itself, which the cache does for the reservation. A foreign holder that leaves the budget alone is documented as undetectable - One PDH query per counter, per-item CStatus checked, usage clamped rather than dropped with the budget - Interpolate decode peaks in bytes so a smaller image is never priced above a larger one
|
Your inference was right, and settling it on the card also killed my fix. Both halves below, with the measurements. The basis mismatch is realExactly as you predicted. RX 9060 XT, 4 GiB allocated then freed and The adapter counter and It is allocation-pattern dependent, which is why my earlier characterisation was too narrow: with the default allocator, one 4 GiB block retains nothing, four 1 GiB blocks retain 1.0 GiB, sixteen 256 MiB blocks retain 1.8 GiB. Under expandable segments it is 2.0 GiB in all three shapes. But the adapter counter cannot be used at allBefore switching the subtraction to
So What ships insteadOnly figures I could verify: the budget and this process's own driver-side usage, both from the one D3DKMT call. ceiling = min(total_memory, max(budget, current_usage))A budget below what this process already holds is Windows asking it to trim itself — which is exactly what the cache does to honour the reservation — so it must not drive tiling. Above that, the budget is the ceiling, which is your round-2 finding unchanged. What it does in the states I measured:
Your round-2 test passes unchanged, and the mid-session case that made me write the credit in the first place stays fixed. The limit is now documented rather than papered over: a foreign process that takes memory without moving our budget is invisible here. Windows keeps reporting 15.09 GiB next to an 8 GiB holder, torch's free figure ignores other processes by construction, and the adapter counter contradicts itself as above. That decode pages, and the end-of-session paging warning is what surfaces it. If you know a driver-side adapter figure that holds up cross-process — Hardening
The 1792² neighbour you flaggedYou were right not to touch the 2281 measurement, and right that bracket-max interpolation was what made the neighbours look wrong. Fixed at the interpolation instead: the bytes are interpolated between measured points rather than the per-pixel constant, since the constant dips while the bytes rise monotonically. Measured points keep their measured value, and past the last point the reservation constant still takes over. Monotonic now, and a test pins that a smaller image is never priced above a larger one. Verification
|
lstein
left a comment
There was a problem hiding this comment.
Follow-up review — round 4
Everything from the last round is resolved, and removing the adapter counter outright was the right call — measuring that it contradicts itself between processes is a better answer than the same-basis fix I suggested. Confirmed fixed:
- The accounting-basis mismatch is gone with the counter.
local_video_memoryreturns both figures from one D3DKMT query, budget validated0 < Budget <= total.- Per-item
CStatusis checked — mutation-verified, re-admittingPDH_CSTATUS_INVALID_DATAbrings the garbage reading back. - One query per counter path, so
paged_bytesis no longer disabled by an unrelated counter. - Monotonic pricing: 0 area-ordered inversions across 1089 (H,W) pairs on both backends, and the regression sweep shows no untiled→tiled flips on any card, so the pricing change costs no pixel identity.
One regime is still wrong, and it is narrow — this is not a reopened round.
The max(budget, usage) override fails open in two reachable states
invokeai/backend/util/vae_working_memory.py:51
total_bytes = min(total_bytes, max(budget_bytes, usage_bytes))The override reads a budget below our own usage as "Windows is asking this process to trim itself, and trimming recovers that much". That is true when we are the only pressure on the card. It is not true when someone else is, and CurrentUsage cannot tell the two apart — it is commitment, and wddm.py:16-17 notes it "counts paged bytes as local".
Trigger 1 — a foreign holder that does move the budget. 16 GiB card, foreign process holding 8 GiB, Invoke committed 12 GiB with part of it already paged out:
budget 7.6, usage 12.0, Qwen 1024px ceiling=12.00G price=8.93G tiles=False
-> evicting every model Invoke owns still leaves only ~7.9G (16 - 8 foreign)
parent min(total, budget) ceiling= 7.60G price=8.93G tiles=True
The decode pages. This is a regression against the parent commit and against the pre-round-3 min(total, budget) rule, and it is not the case the new docstring excuses — the "known limit" is a holder that leaves the budget alone, whereas here the budget moved correctly and the override discarded it.
Trigger 2 — the clamp maximizes the ceiling in exactly the paging state. wddm.py:303:
return int(info.Budget), max(0, min(int(info.CurrentUsage), adapter.total_bytes))CurrentUsage= 12.00G -> usage=12.00G -> ceiling=12.00G
CurrentUsage= 17.00G -> usage=15.92G -> ceiling=15.92G <- correction disabled
CurrentUsage= garbage -> usage=15.92G -> ceiling=15.92G <- correction disabled
The comment at wddm.py:300-302 states this state is reachable, and it is the over-commit state — i.e. the correction switches itself off precisely when Windows is already paging. Budget gets a plausibility check; CurrentUsage gets only a clamp, and the clamp pushes the ceiling to the nameplate rather than rejecting the reading. This trigger does not depend on any contested measurement.
Direction, not a prescription: bound the usage override by something actually resident. paged_bytes() is already read once per session for the warning and is the per-process counter this commit still trusts, so CurrentUsage - paged is a residency figure on the same basis. And CurrentUsage > total looks like it should drop the override while keeping the budget for cuda_mem_get_info, rather than clamp to the most optimistic value.
While you are in there: the budget's behaviour is now documented both ways
Trigger 1's reachability depends on whether a foreign process lowers the budget, and the PR currently asserts both answers:
invokeai/backend/util/wddm.py:10-11— "next to another GPU process, Windows lowers the budget within about a second"- PR description, item 2 — "15.09 of 15.92 GiB alone; lowered within a second when another GPU process starts"
invokeai/backend/util/vae_working_memory.py:41-42(new) — "Windows keeps reporting 15.09 GiB next to an 8 GiB holder"
The round-4 design rests on the third while the module header still asserts the first. Whichever matches the card, the other two should follow — and if the budget really does track foreign pressure, that is an argument for trusting it over the usage override rather than the reverse.
Non-blocking
test_a_counter_this_build_does_not_have_only_disables_what_reads_it(tests/backend/util/test_wddm.py:248) passes unchanged on the parent commit — after the adapter counter is removed there is only one counter path left, so there is no second reader for it to protect. It documents an intent rather than guarding behaviour.tests/backend/util/test_vae_pretile.py:85,98still monkeypatchtorch.cuda.mem_get_info, which the gate no longer calls. Dead setup that reads as coverage.- The pricing change is larger than "fix non-monotonicity": on ROCm/16 GiB, 197 sizes move from tiled to untiled, 134 of them landing more than 1 GiB under the line (worst
1600x1472, 20.05 → 14.37 GiB). The direction is intended and nothing became newly tiled, but the1024²→1536²chord spans the flattest measured bracket, so that band is the least evidenced part of the curve. A running max over the bracket-max prices would have removed the inversions without lowering any price, if you would rather not carry that. - No test covers either blocker regime, and monotonicity is asserted for one CUDA pair with no ROCm counterpart (ROCm had its own inversion on the parent:
1600x147220.05 GiB vs1024x230414.38 GiB).
Checks
- Full suite,
-n 4: 11759 passed, 177 skipped, 8 xfailed, 0 failed. - Targeted:
test_wddm.py,test_vae_pretile.py,test_torch_cuda_allocator.py,test_vae_working_memory_backend.py,test_qwen_image_working_memory.py,test_anima_vae.py,test_flux2_working_memory.py,model_cache/→ 731 passed, no failures. - All 16 CI checks green.
ruff check/format --checkon the five changed files: clean.test_a_budget_below_our_own_residency_does_not_tile_a_decode_that_fitsgenuinely fails if the ceiling reverts to the bare budget, so the new coverage is load-bearing.- New
local_video_memoryconfirmed inert off Windows (None, nothing at import time).
…dows has paged out Only what this process keeps in VRAM (CurrentUsage minus its paged bytes) may lift the ceiling above the WDDM budget; usage beyond the card or an unknown paged amount leaves the budget in charge. The measured budget behaviour next to a foreign holder now reads the same in wddm.py and the gate docstring. Tests cover the foreign-holder, over-commit and unknown-paged states and area monotonicity of the Qwen pricing on both backends.
|
Thanks — both triggers were real, and measuring the budget next to a holder settled the other point too. The override is bounded by residency now
The residency figure is checked on the card: an RX 9060 XT next to an 8 GiB holder, this process having allocated 9.5 GiB.
The budget, measured, and now described the same way everywhereAn 8 GiB holder, then a second process filling in 256 MiB steps (torch 2.12+rocm7.14):
So both earlier statements were half true. The budget does track foreign pressure, but only once this process grows into it. By the time it could page, the budget is its share of the card, which is why trusting it over a raw usage override is right. The Non-blocking items
Checks
|
# Conflicts: # invokeai/app/invocations/vae/flux_vae_decode.py # invokeai/app/invocations/vae/flux_vae_encode.py # invokeai/app/invocations/vae/wan_latents_to_video.py # invokeai/app/invocations/vae/z_image_image_to_latents.py # invokeai/app/invocations/vae/z_image_latents_to_image.py # invokeai/backend/util/vae_working_memory.py # tests/app/invocations/test_anima_vae.py # tests/app/invocations/test_flux_vae_decode_tiling.py # tests/app/invocations/test_wan_working_memory.py
…eat/rocm-windows-memory
…indows-memory Carries the branch's low-VRAM docs and regenerated API contracts to their moved locations under configuration/Optimization and frontend/api.
Keeps the corrected force_tiled_decode description next to auto_tiled_decode and regenerates the API contracts. Folds the automatic-tiling docs into the existing Tiled VAE decode and encode section and fixes the moved page's system-requirements link.
Summary
Generations on a ROCm build of PyTorch under Windows ran 4-5x slower than the hardware allows, and Krea-2 fp8 did not run at all. Windows never fails an allocation that does not fit: it moves it into shared system memory and keeps it there, so nothing in the log says why a run crawls. Measured on an RX 9060 XT (gfx1200, torch 2.12+rocm7.14.1): Z-Image Turbo nvfp4 at 1024px denoised in 144/139/186 s over three runs; it now takes 34 s, with pixel-identical images. Krea-2 fp8 produced no step in two minutes; it now takes 60 s per image. Z-Image at 1536px is possible for the first time.
Four causes, addressed in order:
expandable_segments:True, set before torch is imported and only when no allocator variable is configured. Any explicitpytorch_cuda_alloc_confstill wins.torch.cuda.mem_get_infois the device total minus this process's own usage there: it ignores other processes, and Windows starts paging once the process passes its WDDM budget (15.09 of 15.92 GiB alone; next to another GPU process it stays there while this process holds little and steps down to this process's share once the two together pass ~12 GiB -- 7.58 GiB next to an 8 GiB holder). The cache's free figure is now capped by that budget throughD3DKMTQueryVideoMemoryInfo, and a worker says so when memory stays in system RAM across two sessions (PDHGPU Process Memory\Shared Usage). Both are best-effort and silent everywhere else.TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1, which InvokeAI does not set. On torch 2.12+rocm7.14.1, every call that reaches them with the flag set fails withhipErrorInvalidValue(Z-Image, FLUX.1, CLIP and VAE shapes, masked and unmasked), and a full CLIP text encoder forward does not finish. torch 2.13.0+rocm10.0.0 runs them correctly with the flag, the VAE's 512-wide heads included; turning the flag on belongs to the ROCm 10 upgrade, not to this PR. The existing ROCm SDPA guard now computes such a call in chunks of at most 1 GiB of scores - head groups first, query rows otherwise - and the working-memory estimates are capped to match. Chunking is exact in exact arithmetic and lands within one bf16 ulp of the unchunked call; head groups are bitwise identical where the kernel's arithmetic does not depend on the batch it runs over, which holds on CPU but not for a rocBLAS batched GEMM that re-tiles with the group count. Krea-2's denoise reservation prices the score matrix, which it did not before.auto_tiled_decodesetting. On Windows the line is the WDDM budget, overridden only by what this process keeps resident (CurrentUsageminus its paged bytes). For the Qwen-Image VAE the decision uses the measured peak curve (the chord between measured points) instead of the reservation constant; reservations keep the constant. On ROCm/16 GiB that moves 197 sizes from tiled to untiled; none become newly tiled. The FLUX.1 autoencoder's working-memory constants follow the convolution backend, as the FLUX.2 ones already did: MIOpen needs 3600 B/pixel-byte to decode where cuDNN needs 2200.Items 3 and 4 change behaviour on every platform: math-kernel attention runs in chunks on Linux ROCm too, and decodes that would take most of the card are tiled everywhere. On CUDA the estimates and the fused path are untouched.
Documentation says only what is true: AMD remains supported on Linux only; the Windows ROCm notes are hints.
Related Issues / Discussions
None.
QA Instructions
Gates
uv run --no-sync pytest -n 6(after mergingorigin/main): 11,787 passed, 184 skipped, 9 xfailed.uv tool run ruff@0.11.2 check/format --checkon the changed paths: clean.pnpm -C docs build: 326 pages, link validation clean. Generatedopenapi.json,schema.tsanddocs/src/generated/settings.jsonwere regenerated with their generators;generate_docs_json.pynow emits Path defaults with.as_posix(), so regenerating on Windows no longer flips the separators and failscheck-docs-data.-m slow, RX 9060 XT): 5 passed, 1 skipped. The skip istest_a_fused_kernel_is_still_wrong_for_the_wide_head: no fused backend runs that shape on gfx1200 with this build, so there is nothing to compare.tests/app/invocations/test_flux2_working_memory.pyfail depending on the build (5 on gfx1200 / torch 2.12+rocm7.14 here); they fail identically on the base commit, as they assume a CUDA/Linux backend.End-to-end, from this branch's code with no runtime patches
RX 9060 XT, 16 GB (port 9091):
RTX 4090, 24 GB (port 9090), for the CUDA regression:
Measurements behind two design decisions
CurrentUsage. HIP under Windows keeps 1-2 GiB after a free plusempty_cache()and hands it to the next allocation, whileCurrentUsagestill counts it: 4 GiB allocated, then freed, leavesCurrentUsageat 1.14 GiB withmemory_reservedat 0 and torch's free figure back at its starting value; the next 2 GiB cost only 1 GiB of new usage. Capping againstCurrentUsagewould hide what an offload just freed.delfreed 0.00 GiB,empty_cache()freed 2.99 GiB), so the loop saw no progress and unloaded every unlocked model.Not verified
Linux ROCm (the query-row chunks in the VAE and the FLUX.1 MIOpen constant), multi-GPU under Windows ROCm, other Windows ROCm builds, MPS/XPU pre-tiling, and small CUDA cards on real hardware. The Qwen-Image VAE's ROCm constants are unchanged (measurable only on Windows here), and
auto_detect_slice_sizestill usestorch.cuda.mem_get_infodirectly.Review
Material findings resolved:
A
settings.jsonregenerated on Windows flipped two path defaults and would have failed the docs job; the generator now emits POSIX paths.The allocator test fixture leaked
PYTORCH_CUDA_ALLOC_CONFinto later tests, which made a model-cache test fail depending on file order.The offload loop over-unloaded under expandable segments (see the measurement above), now covered by a regression test that fails without the fix.
The allocator default silently did nothing when torch was already imported; it now warns instead.
The paging warning could fire on an overflow that returns to VRAM by itself, and advised a restart; it now requires two consecutive readings and names settings to change.
Chunking sliced 3-D and wider-batch K/V wrongly; those calls stay whole.
The FLUX.1 decode node reserved nothing for a diffusers-layout
AutoencoderKL, which is the exact failure this PR targets; it is now priced and pre-tiled like the Z-Image node.The up-front tiling log was at INFO on every Anima decode and suggested turning the setting off, which is bad advice for Wan (no OOM retry) and Anima (tiling measured faster).
Test gaps closed: masked query-row chunking, Krea-2 with CFG and a regional mask,
auto_tiled_decode=falsefor Wan and Anima, a Wancpu_onlytest that passed regardless of the code, PDH struct offsets plus a hardware test, and a test that the worker loop calls the warning at all.The pre-tiling ceiling over-trusted
CurrentUsage: next to a foreign process, or while over-committed, it counts what Windows has already paged out. The override is now bounded by residency, with tests for both states and area-monotonicity tests for the Qwen pricing on both backends.Compatibility / Rollout
auto_tiled_decode(default true); generated OpenAPI, frontend types and docs settings regenerated.wddm.pyis new and best-effort: every failure answers "unknown" and leaves the previous behaviour in place.Checklist
What's Newcopy (if doing a release after this PR)