Repository navigation
[Loom] Keep every AIE2P fixed permutation vectorized - #1318
Merged
Merged
Conversation
Derive MOVXM encoding adapters for both scalar halves of each allocatable predicate register. Add descriptor forms that start a predicate with a signed 32-bit constant and complete its high half while preserving the low-half storage continuation. This gives target callback lowerings a general compile-time lane-mask primitive without routing predicate construction through fixed selector state or source-level scalar IR.
Preserve rank-one vector.shuffle operations whenever their source and result fit an AIE2P register carrier. Selection records complete source-packet aliases in an eight-byte plan, so identity and packet routing do not build generic shuffle state. Compose remaining packets from native lane broadcasts and compile-time predicate selects. Predicate carriers widen once to byte vectors and pack back; accumulator carriers move one logical packet through vector registers, while 64-bit lanes duplicate each selector bit across the 32-bit select granularity. Keep the structured i8 4x4 transpose path ahead of the fallback and scalarize only carrierless shapes. Exact source-to-Low coverage spans every scalar type and carrier boundary, and the shared execution corpus compares repeated-lane semantics against a scalar oracle.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
AIE2P now keeps every fixed rank-one permutation in register form whenever the vector has a hardware carrier. A constant
vector.table.lookupsuch as[3, 2, 1, 0]previously became scalar extracts followed by a scalar insert chain; it now remains avector.shuffleand lowers to lane broadcasts plus predicate selects. This removes the scalar-register pressure and allocator repair that made even fully static permutations expensive to compile and difficult to schedule.The implementation covers the complete carrier family rather than the first failing element type:
i1,i8, both float8 types,i16,f16,bf16,i32,f32,i64,f64,index, andoffset. It handles ordinary vectors through the maximum 1024-bit carrier, one- and two-unit predicates, and the two- and four-unit accumulator forms used byi32,f32, andi64. Shapes without an AIE2P carrier retain the reference legalization contract.Route complete packets before composing lanes
Selection consumes the canonical shuffle's
source_lanesattribute and records the physical carrier, element width, logical packet count, and source alias for each result packet in an eight-byte target plan. That plan gives emission three bounded routes:For example, a three-lane
i16permutation[1, 0, 1]has this complete shape:A maximum-carrier packet swap is cheaper still:
The composer scans at most one 512-bit packet, so its bound is 64 lanes regardless of carrier size. It caches source-packet conversions and derives every routing decision during selection instead of rediscovering structure during emission. The existing specialized i8 4x4 transpose remains ahead of this fallback and retains its native sequence.
Preserve predicate and accumulator carriers
Predicate payloads widen once to byte vectors, use the same packet composer, and compare back to packed predicate bits. Accumulator packets move through vector512 only at the composition boundary; exact packet aliases stay in accumulator registers. Physical padding repeats the final logical packet, matching the existing carrier contract.
Arbitrary byte and predicate maps need all 64 selector bits. The prerequisite commit adds descriptor forms that initialize the low 32 bits of an allocatable predicate register and complete its high half while retaining the low-half storage continuation. Zero and arbitrary high halves share that representation. For 64-bit elements, each logical lane expands to two adjacent selector bits because the hardware select operates at 32-bit granularity.
Worst-case allocation and schedule
An i8x128 reversal deliberately defeats packet aliasing and repeats no lanes, forcing the general fallback to its maximum 128 broadcasts and 126 selects. This is the hostile end of the family; identities and packet routes introduce no ALU work, and repeated-lane maps need one broadcast per distinct lane.
The emitted leaf exposes the tradeoff directly:
The general vector fallback spends 332 additional code bytes on this maximum-entropy map, while removing every allocator repair and reducing the schedule by 91 issue cycles. This is an explicit fallback cost rather than a claim that every permutation shrinks. The ordinary table-lookup and packet-routing cases that motivated the change are much smaller.
Coverage and diagnostics
Exact Source-to-Low goldens cover every payload width, all 13 scalar and address types, the first partial packet beyond 512 bits, maximum ordinary carriers, a two-packet predicate with a full 64-bit selector, the special two-unit
f32accumulator, and four-uniti32andi64accumulators with padding. Existing lookup-table cases now show the complete scalar-to-vector transition rather than checking for isolated mnemonics.The shared execution corpus checks a repeated-lane cross-half byte permutation against an independent scalar oracle. Compile reports expose
fixed-permutation.packet-routingor the width-specific broadcast/select strategy and confirmscalarized=false, so future regressions remain visible inloom-compile-report suggest.Reviewer Notes
The highest-value review surface is the compact plan and packet-alias classification in
lower/shuffle.c, followed by predicate and accumulator conversion at the packet boundary. The generatedvector_shufflegolden shows every distinct emitted mechanism, while the lookup-table golden shows the user-visible path that originally scalarized.