Repository navigation
[Loom] Separate immutable AMDGPU wait classification - #1320
Merged
Merged
Conversation
Effect dependencies are produced in one contiguous construction phase, but AMDGPU wait planning rediscovered them by scanning the full mixed dependency graph. Publish that owned range in the completed schedule and consume it directly in target planning. The range is an eight-byte view into existing graph storage, so it adds no arena allocation or persistent dependency index. Tests now exercise the same range contract when inspecting and synthesizing schedules.
AMDGPU wait planning classified target hazards, effects, explicit bounds, and structural behavior inside the same translation unit that simulated counter progress and emitted actions. Move that construction and its compact node representation into a dedicated classification component with an explicit result consumed by dependency, frontier, loop, and action planning. Preserved-result flags are now established during classification and elided authored waits remain represented by the existing output bitset, keeping classification flags stable. Dependency-head allocation moves to dependency construction, removing its hidden dependence on the former broad planner allocator.
Build classification directly into its restricted result, restore the planner table allocation order, and keep tied-source preservation with dependency construction so the extracted boundary does not obscure ownership or introduce result copies. Retain caller specialization for the single hot classification call while isolating the compact hazard scan from its register pressure. This keeps the standalone component usable by ordinary builds without making the JIT pay for the source split.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
AMDGPU wait planning currently rediscovers schedule facts that their producers already know and classifies immutable target state inside the same implementation that mutates counter epochs and emits wait actions. This change retains the scheduler-owned effect range and moves AMDGPU wait classification into an explicit result consumed by dependency, frontier, loop, and action planning.
The resulting boundary keeps immutable packet and effect facts separate from mutable path simulation. It also makes the scheduler's effect subset reusable without a full dependency index or a scan over unrelated SSA, ordering, and state edges.
Design
The scheduler records the exact contiguous range produced by its effect-dependency phase as an eight-byte graph view. AMDGPU wait planning consumes that range directly, and schedule tests inspect the same range contract rather than filtering the mixed graph themselves.
AMDGPU classification now owns compact per-node flags and decoded effects, frontier and completion seeds, explicit-wait bounds, and the counts that size later planner state. The 16-byte node rows remain stable after construction. Authored wait elision stays in the plan's existing output bitset, while downstream analyses refine only the completion and frontier fields whose ownership is documented.
The extraction builds directly into its restricted result, preserves the planner's prior transient allocation order, and keeps the single hot caller specialized. Tied-source preservation remains with dependency construction because it describes a dependency relation rather than a packet classification fact.
Cost
On the maintained attention frame, the wait planner now visits 8,165 exact effect rows instead of scanning 15,933 mixed dependency rows, without another index or arena allocation.
Across the long planner matrix, the selected stack retires about 1.33% fewer aggregate instructions and branches while aggregate cycles remain neutral to slightly lower. Attention alone retires 2.33% fewer instructions and 2.42% fewer branches; repeated cycle measurements stayed within roughly one percent. Frame signatures, planner arena use, waits, and hazards remain exact. The complete benchmark executable grows by 160 bytes.
Reviewer Notes
The highest-value review surfaces are the ownership line between immutable classification and mutable producer/action state, the schedule range publication point, and the code-shape choices that preserve the hot caller.