Skip to content

[Loom] Emit HAL artifacts through core target emitters - #1322

Merged
benvanik merged 2 commits into
mainfrom
users/benvanik/loom-tooling-pipeline-next
Oct 7, 2026
Merged

benvanik merged 2 commits into
mainfrom
users/benvanik/loom-tooling-pipeline-next

Conversation

@benvanik

@benvanik benvanik commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Live AMDGPU and SPIR-V execution now uses the same core target emitters as LoomC and offline compilation. HAL tooling selects a target for the active device and invokes that emitter directly; it no longer owns a second target-specific artifact compilation layer.

The old AMDGPU wrapper repackaged the native kernel library, while the SPIR-V wrapper repeated entry selection, target compatibility analysis, bundle retention, allocation, and artifact ownership around the core module compiler. Besides living at the wrong boundary, the two implementations made live execution a separate compiler route that could drift from public emission.

Make core emitters the only artifact producers

loom_target_emit_artifact_t can now retain the exact target bundle selected after function refinement and an optional target listing. Both products are caller-requested: ordinary LoomC emission pays no bundle-copy or listing cost, while live HAL execution requests the metadata it needs to load, report, and bundle the result after compiler scratch is released.

The shared HAL candidate adapter builds one loom_target_emit_request_t from the execution session, compiler-retained function versions, manifest options, diagnostic policy, and compile report. The selected device provider contributes its target profile, HAL executable row, and core emitter. AMDGPU and SPIR-V then follow the same path from prepared target-low IR to an owned core artifact.

The emitter request also carries the diagnostic budget instead of letting the AMDGPU adapter silently substitute a fixed value. LoomC, loom-check, VM testbench emission, and HAL execution all preserve the policy established by their own public or tooling boundary.

Preserve live execution and bundle behavior

iree-run-loom one-shot execution, iree-test-loom scenarios, and iree-benchmark-loom candidates now load the primary core artifact directly. Benchmark bundles still receive the target artifact, HAL executable, target listing, manifest sidecar, compile report, and exact function-refined target identity. Provider-facing names remain stable: the Vulkan device provider is still reported as spirv-vulkan-hal even though its core emitter is spirv.

The target-specific tooling artifact providers and their wrapper tests are deleted. SPIR-V module compiler coverage moves beside the compiler implementation and constructs its inputs from target record bundles and function-version facts, so the test no longer reaches across the target/tooling boundary.

Reduce compile-time state and work

The SPIR-V live path no longer performs a second entry-selection and compatibility pass or creates a private block-pool owner around emission; the shared adapter uses session scratch and immediately returns it after the core emitter finishes. The AMDGPU live path removes its wrapper allocation and indirection while retaining the same native kernel-library producer.

A retained HAL candidate shrinks from 144 bytes to 96 bytes. The transient device target shrinks from 32 bytes to 16 bytes by deriving its key from the HAL executable row instead of storing a second string view and validating that both copies agree. Successful provider selection is compiler-owned state and is consumed directly rather than converted into additional internal failure paths.

Runtime dispatch, executable loading, and emitted bytes are unchanged. The change is confined to the cold compile/emission path, net-deletes production machinery, and adds no build visibility exceptions.

Ownership and coverage

Core AMDGPU and SPIR-V emitter tests cover default emission without retained metadata and requested bundle, listing, manifest, and report retention. Device-provider tests cover stable provider identity, emitter identity, target selection, and artifact format. The shared candidate test covers target-environment and function-version propagation, optional listing flags, diagnostic limits, report preservation, and artifact teardown. End-to-end AMDGPU and Vulkan scenario execution exercises the resulting artifacts through the production HAL loader.

Reviewer Notes

The first commit adds opt-in metadata retention to core emitter artifacts. The second switches HAL execution to those artifacts, simplifies device-target state, moves SPIR-V compiler coverage to its owner, and deletes both target-specific tooling compilers.

The highest-value review surfaces are loom/target/provider.h, the AMDGPU and SPIR-V emitter retention paths, and loom/tooling/execution/hal/candidate.c. The central invariant is that tooling selects and stitches capabilities, while core target emitters alone interpret prepared compiler state and construct artifacts.

Extend target emission artifacts with opt-in exact bundle and textual
listing retention. AMDGPU and SPIR-V now copy the post-refinement bundle
into artifact-owned storage only when requested, and AMDGPU transfers
its assembly listing through the same core contract. Default LoomC
requests allocate no storage for this metadata.

This establishes core output ownership for live HAL execution without
requiring target-specific tooling compilers.
Live HAL execution wrapped the core AMDGPU and SPIR-V emitters in
target-specific tooling artifact providers. Replace both wrappers with
one shared candidate adapter that invokes the selected device provider's
core emitter using the session target environment and compiler-retained
function versions. The core artifact now carries the exact target
bundle, listing, and sidecars through loading and benchmark bundle
production.

Device providers now own only live target selection and the association
with a core emitter. Collapse duplicated target-key state onto the HAL
executable target, trust successful provider selection instead of
revalidating compiler-owned results, and preserve the caller's
diagnostic budget in the emitter request.

Delete both duplicate artifact compilers and their wrapper tests, and
move SPIR-V module compiler coverage beside the implementation it
exercises.
@benvanik
benvanik marked this pull request as ready for review October 7, 2026 00:41
@benvanik
benvanik requested a review from a team as a code owner October 7, 2026 00:41
@benvanik
benvanik merged commit 0ce79b5 into main Oct 7, 2026
24 of 28 checks passed
@benvanik
benvanik deleted the users/benvanik/loom-tooling-pipeline-next branch October 7, 2026 00:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant