Skip to content

[Loom] Keep every AIE2P fixed permutation vectorized - #1318

Merged
benvanik merged 2 commits into
mainfrom
users/benvanik/loom-xdna-fixed-permutation-family
Oct 7, 2026
Merged

benvanik merged 2 commits into
mainfrom
users/benvanik/loom-xdna-fixed-permutation-family

Conversation

@benvanik

@benvanik benvanik commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

AIE2P now keeps every fixed rank-one permutation in register form whenever the vector has a hardware carrier. A constant vector.table.lookup such as [3, 2, 1, 0] previously became scalar extracts followed by a scalar insert chain; it now remains a vector.shuffle and lowers to lane broadcasts plus predicate selects. This removes the scalar-register pressure and allocator repair that made even fully static permutations expensive to compile and difficult to schedule.

The implementation covers the complete carrier family rather than the first failing element type: i1, i8, both float8 types, i16, f16, bf16, i32, f32, i64, f64, index, and offset. It handles ordinary vectors through the maximum 1024-bit carrier, one- and two-unit predicates, and the two- and four-unit accumulator forms used by i32, f32, and i64. Shapes without an AIE2P carrier retain the reference legalization contract.

Route complete packets before composing lanes

Selection consumes the canonical shuffle's source_lanes attribute and records the physical carrier, element width, logical packet count, and source alias for each result packet in an eight-byte target plan. That plan gives emission three bounded routes:

  • An identity aliases the source value directly.
  • A result packet that exactly aliases a source packet uses a slice, with a concat only when the carrier has several packets.
  • A packet with an arbitrary lane map broadcasts each distinct source lane once and folds those broadcasts with compile-time predicate masks.

For example, a three-lane i16 permutation [1, 0, 1] has this complete shape:

%lane_1 = vextbcst.16 %source, 1
%lane_0 = vextbcst.16 %source, 0
%mask = predicate.constant.low32 2
%result = vsel.16.mask64 %lane_1, %lane_0, %mask

A maximum-carrier packet swap is cheaper still:

%high = slice %source[2]
%low = slice %source[0]
%result = concat(%high, %low)

The composer scans at most one 512-bit packet, so its bound is 64 lanes regardless of carrier size. It caches source-packet conversions and derives every routing decision during selection instead of rediscovering structure during emission. The existing specialized i8 4x4 transpose remains ahead of this fallback and retains its native sequence.

Preserve predicate and accumulator carriers

Predicate payloads widen once to byte vectors, use the same packet composer, and compare back to packed predicate bits. Accumulator packets move through vector512 only at the composition boundary; exact packet aliases stay in accumulator registers. Physical padding repeats the final logical packet, matching the existing carrier contract.

Arbitrary byte and predicate maps need all 64 selector bits. The prerequisite commit adds descriptor forms that initialize the low 32 bits of an allocatable predicate register and complete its high half while retaining the low-half storage continuation. Zero and arbitrary high halves share that representation. For 64-bit elements, each logical lane expands to two adjacent selector bits because the hardware select operates at 32-bit granularity.

Worst-case allocation and schedule

An i8x128 reversal deliberately defeats packet aliasing and repeats no lanes, forcing the general fallback to its maximum 128 broadcasts and 126 selects. This is the hostile end of the family; identities and packet routes introduce no ALU work, and repeated-lane maps need one broadcast per distinct lane.

The emitted leaf exposes the tradeoff directly:

Metric Scalar construction Vector composition
Scalar extracts 128 0
Scalar register inserts 126 0
Allocation-repair events 81 0
Issue cycles 536 445
Coissued bundles 1 63
Code bytes 1,964 2,296

The general vector fallback spends 332 additional code bytes on this maximum-entropy map, while removing every allocator repair and reducing the schedule by 91 issue cycles. This is an explicit fallback cost rather than a claim that every permutation shrinks. The ordinary table-lookup and packet-routing cases that motivated the change are much smaller.

Coverage and diagnostics

Exact Source-to-Low goldens cover every payload width, all 13 scalar and address types, the first partial packet beyond 512 bits, maximum ordinary carriers, a two-packet predicate with a full 64-bit selector, the special two-unit f32 accumulator, and four-unit i32 and i64 accumulators with padding. Existing lookup-table cases now show the complete scalar-to-vector transition rather than checking for isolated mnemonics.

The shared execution corpus checks a repeated-lane cross-half byte permutation against an independent scalar oracle. Compile reports expose fixed-permutation.packet-routing or the width-specific broadcast/select strategy and confirm scalarized=false, so future regressions remain visible in loom-compile-report suggest.

Reviewer Notes

The highest-value review surface is the compact plan and packet-alias classification in lower/shuffle.c, followed by predicate and accumulator conversion at the packet boundary. The generated vector_shuffle golden shows every distinct emitted mechanism, while the lookup-table golden shows the user-visible path that originally scalarized.

Derive MOVXM encoding adapters for both scalar halves of each
allocatable predicate register. Add descriptor forms that start a
predicate with a signed 32-bit constant and complete its high half while
preserving the low-half storage continuation.

This gives target callback lowerings a general compile-time lane-mask
primitive without routing predicate construction through fixed selector
state or source-level scalar IR.
Preserve rank-one vector.shuffle operations whenever their source and
result fit an AIE2P register carrier. Selection records complete
source-packet aliases in an eight-byte plan, so identity and packet
routing do not build generic shuffle state.

Compose remaining packets from native lane broadcasts and compile-time
predicate selects. Predicate carriers widen once to byte vectors and
pack back; accumulator carriers move one logical packet through vector
registers, while 64-bit lanes duplicate each selector bit across the
32-bit select granularity.

Keep the structured i8 4x4 transpose path ahead of the fallback and
scalarize only carrierless shapes. Exact source-to-Low coverage spans
every scalar type and carrier boundary, and the shared execution corpus
compares repeated-lane semantics against a scalar oracle.
@benvanik
benvanik marked this pull request as ready for review October 6, 2026 22:45
@benvanik
benvanik requested a review from a team as a code owner October 6, 2026 22:45
@benvanik
benvanik merged commit 087addc into main Oct 7, 2026
25 of 28 checks passed
@benvanik
benvanik deleted the users/benvanik/loom-xdna-fixed-permutation-family branch October 7, 2026 00:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant