HIP port of antirez/h3.c for AMD GPUs.
One tree, three supported products — pass HIP_ARCH to match the GPU you
compile for. Timed scoreboard SKUs are Strix Halo (gfx1151), MI210
(gfx90a), and MI300X (gfx942). MI250 / MI250X also use gfx90a but
were not timed.
| Product | HIP_ARCH |
Default DiT | SDPA |
|---|---|---|---|
| Strix Halo (RDNA) | gfx1151 |
INT8 + BF16 activations | wave32 rocWMMA |
| MI210 (CDNA2) | gfx90a |
INT8 (default) | wave64 MFMA flash |
| MI300X (CDNA3) | gfx942 |
INT8 (default) | wave64 MFMA flash |
Tagged v0.15.0 adds three opt-in FastH3 paths, all off by default.
--fasth3-lora merges a dense 4-step LoRA and runs timesteps 999 / 749 /
500 / 250. --taeh3 replaces the video VAE with a tiny decoder.
--vsa turns on sparse video attention and requires the vsa-datafree
adapter. They do not combine with --token-reduction, --sol-attn,
--fbc, or --reuse above 1, and they keep all 50 layers. On MI300X
the 832×480 · 5 s INT8 process is dense 166.87 s, --fbc 53.31 s,
FastH3 30.12 s, FastH3+TAEH3 27.72 s, VSA 32.61 s, VSA+TAEH3
26.02 s. The fox-15s canvas (864×480 · 15 s; FastH3 uses 4 steps and
reuse 1, not the script's 20 steps / 45 layers / reuse 2) is dense
213.3 s, fox-15s.sh 146.21 s, --fbc 136.43 s, FastH3
108.68 s, FastH3+TAEH3 89.12 s, VSA 95.06 s, VSA+TAEH3
76.72 s. The 4-step picture follows the prompt and is a different
composition from dense; TAEH3 is softer, and the 15 s VSA ending has a
second person in the office. v0.14.0 remains the --fbc release.
v0.13.0 remains the quality-path CDNA flash SDPA and Sol-Attn-past-64k
release. INT8 DiT is default on all ISAs (H3_INT8_MLP=0 for BF16).
v0.12.0 was TR schedule + fused AdaLN + 3-GPU retune; v0.11.x had opt-in
INT8; v0.10.x is the dual-ISA history; v0.9.x is Strix Halo only. The original project is a native MiniMax-H3
inference engine (Apple Metal / macOS); this repository reimplements the
GPU backend in pure HIP so the same CLI and model stack run on ROCm.
Project ident, generated on Strix Halo (gfx1151) (864×480, 56 frames, --steps 20 --layers 50 --reuse 1). Click the poster for the MP4:
h3-hip-ident.mp4.
Project page: alexhegit.github.io/h3-hip.c Wiki: github.com/alexhegit/h3-hip.c/wiki Original project: antirez/h3.c CUDA sibling (DGX Spark): alexhegit/h3-spark.c Official weights: MiniMaxAI/MiniMax-H3
Headline T2VA (same knobs; details in docs/PERFORMANCE.md):
| Preset | Strix Halo (gfx1151) | MI210 (gfx90a) | MI300X (gfx942) |
|---|---|---|---|
| fox-s2 | ~85–90 s (I/O) | 11.2 s | 10.6 s |
| fox-fast | ~2 min (I/O) | 19.5 s | 9.8 s |
| 15 s cinematic (864×480, 362 f) | 40 min 4 s (no TR) / 26 min 56 s (fox-15s.sh) |
12 min 12 s (no TR) / 8 min 13 s (fox-15s.sh) |
199 s (no TR) / 149 s (fox-15s.sh) |
These are complete muxed MP4s (video + audio), not stubs. Headline
cells are process E2E (/usr/bin/time until the MP4 is written)
except Halo short clips (I/O-bound). MI300X short clips are the 2026-09-18
v0.12.0 retake (b0ff702): fox-s2 warm n=5 mean 10.6 s
(denoise 0.29 s); fox-fast warm n=5 mean 9.8 s (denoise 1.94 s,
DiT total 5.0 s). v0.13.0 15 s no-TR is 199 s (quality-path
SDPA A/B 215 → 199 s; v0.12.0 was 213 s).
fox-15s-fast.sh process E2E is 118 s. Do not quote DiT total as
time-to-MP4 (fox-fast DiT 5.0 s). Cold
fox-fast is ~20 s. fox-s2 and fox-fast are both 512² · 22 frames
(~0.9 s at 24 fps). The README fox showcase clip is a third preset
(--layers 50 --reuse 1). Ledger:
docs/perf-runs/MI300X_2026-09-18_full.md
and docs/perf-runs/MI300X_2026-09-21_1344x768-15s-sdpa.md.
| Name | Size | Knobs | Role |
|---|---|---|---|
| fox-s2 | 512² · 22 f | --steps 2 --layers 35 --reuse 1 |
HIP A/B + gfx1151 md5 gate |
| fox-fast | 512² · 22 f | --steps 20 --layers 45 --reuse 2 |
Complete short clip; 11 DiT evals; upstream “first fast video” |
| fox showcase | 512² · 22 f | --steps 20 --layers 50 --reuse 1 |
README / wiki gallery fox |
| 15 s cinematic | 864×480 · 362 f | --steps 20 --layers 45 --reuse 2 |
Same quality path as fox-fast, long duration |
Clips below were generated on an AMD Strix Halo iGPU (gfx1151) with this HIP
port. Click a poster for the MP4. Several clips are untitled model output
(no ffmpeg captions). The 15 s three-way reel is labelled left-to-right.
| Mode | Sample |
|---|---|
| T2VA — fox tutorial clip | mp4 |
| T2VA — Linux 35 tribute (untitled) | mp4 |
| T2VA — project ident draft (untitled) | mp4 |
| Ref2VA — AMD developer community (untitled) | mp4 |
| T2VA — 15 s cinematic office (untitled) | mp4 |
| T2VA — 15 s Halo three-way (dense / TR / Sol-Attn) | mp4 |
| T2VA — 15 s MI300X FastH3 (4-step / 4-step+TAEH3) | mp4 |
| T2VA — 10 s cinematic office (untitled) | mp4 |
Long clips (864×480, --steps 20 --layers 45 --reuse 2): 15 s E2E 40 min 4 s
without TR (main 2026-09-12); ./bench/fox-15s.sh (TR 4:30) is 26 min 56 s
on Strix Halo (gfx1151); 15 s E2E 12 min 12 s / fox-15s.sh 8 min 13 s
on MI210 (gfx90a); 15 s process E2E 199 s / fox-15s.sh 149 s
on MI300X (gfx942) / fox-15s-fast.sh 118 s.
Halo same-seed 15 s lossy A/B (2026-09-16): fox-15s.sh TR is 27 min 46 s
at 18.23 dB / 0.667 vs dense; --sol-attn is 29 min 47 s at
19.55 dB / 0.721 (audio SNR 3.11 dB vs 10.99 dB). TR is faster; Sol-Attn
is closer to dense. Triptych:
fox-15s-3way-compare-gfx1151.mp4
· docs/SOL_ATTN.md.
Timings and reproduce commands:
docs/perf-runs/LONG_VIDEO.md ·
Wiki: Long video ·
Project page § Long T2VA.
FL2VA and Ref2VA (image, silent video, embedded soundtrack, image+audio) are wired and smoke-tested on the same binary; see Status.
The titled ident at the top of this README is a later overlay of a 864×480
run (--layers 50 --reuse 1). The ident row in the table is the untitled
512×288 draft.
Build first (make HIP_ARCH=gfx1151 or HIP_ARCH=gfx90a). Then:
MODEL=/path/to/MiniMax-H3
# Fox (512², 22 frames, full layers)
./h3 -d "$MODEL" \
-p "A red fox walks through fresh snow." \
--width 512 --height 512 --frames 22 \
--steps 20 --layers 50 --reuse 1 \
-o assets/showcase/t2va-fox.mp4
# Linux 35 tribute (untitled; 864×480, 56 frames)
./h3 -d "$MODEL" \
-p "A single penguin stands on a snowy ridge at blue hour, facing a glowing amber terminal screen floating in the cold air, scrolling lines of green monospaced code reflected in its eyes. Slow dolly-in, shallow depth of field, volumetric mist, cinematic teal and amber grade. Ambient wind, soft mechanical keyboard clicks, a low warm synth pad." \
--width 864 --height 480 --seconds 2.33 \
--steps 20 --layers 50 --reuse 1 \
--seed 1991 \
-o assets/showcase/linux35.mp4
# Project ident draft (untitled; 512×288, 56 frames)
./h3 -d "$MODEL" \
-p "Cinematic technology project ident. In a dark graphite silicon landscape, streams of glowing model weights flow from an SSD into a powerful red GPU compute array. Magenta and amber wavefronts race through precise matrix tiles, code particles transform into a grid of moving video frames, and a realistic red fox briefly emerges from a snowy digital scene while a luminous audio waveform pulses beneath it. Slow controlled camera pullback reveals the complete accelerator board, elegant engineering visualization, AMD-red and HIP-magenta color palette, volumetric light, crisp reflections, premium cinematic motion graphics, no readable text, no logos. Sound: deep electronic startup pulse, rapid soft data clicks, rising synthesized tone, clean final impact." \
--width 512 --height 288 --frames 56 \
--steps 20 --layers 45 --reuse 2 \
--seed 1991 \
-o assets/showcase/h3-hip-ident-draft-raw.mp4
# AMD developer community (untitled Ref2VA; 864×480, 56 frames)
# Needs Ref2VA weights. Image encode on this HIP path is slow (~5 min per still).
./h3 -d "$MODEL" \
-p "Cinematic AMD developer community invitation. Preserve the two referenced black developer T-shirt designs: first, the gold Helios chariot artwork on the back; second, the futuristic runner and open-world artwork on the front. Begin with a close view of the golden Helios shirt worn by a developer in a modern hardware lab. The camera makes a smooth orbit to reveal another developer wearing the runner shirt. Together they walk through a luminous orange open doorway into a welcoming diverse group of software engineers. Every engineer in the group, including the people further away in the background, wears the same matching black AMD team T-shirt with the crisp geometric AMD arrow emblem clearly printed on the chest and on the back, sharp white and orange print, consistent uniform team branding across the whole room. They gather around glowing code displays and an AMD-powered workstation. Confident, inclusive, collaborative energy, black graphite with AMD orange and electric blue accents, premium technology campaign, realistic fabric, crisp apparel print detail, cinematic lighting, no generated captions. Sound: subtle server room ambience, keyboard clicks, warm rising electronic pulse, uplifting final impact." \
--ref-image assets/showcase/refs/helios-shirt.png \
--ref-image assets/showcase/refs/run-open-shirt.png \
--ref-image-size max \
--width 864 --height 480 --seconds 2.33 \
--steps 20 --layers 50 --reuse 1 \
--seed 2026 \
-o assets/showcase/amd-developer-community-raw.mp4
# Long T2VA — 15 s cinematic office (864×480, 362 frames;
# E2E 40 min 4 s Strix Halo (gfx1151) no TR / 12 min 12 s MI210 (gfx90a);
# ./bench/fox-15s.sh → 26 min 56 s Halo / 8 min 13 s MI210
./h3 --profile -d "$MODEL" \
-p "15 seconds, 16:9 landscape cinematic. A lone software engineer works late in a dim home office lit only by monitor glow and a desk lamp. Photoreal live-action feel with subtle handheld camera breathing.
[0–3 seconds] Medium shot from behind the desk. Code scrolls on dual monitors; warm red accent light reflects on glass. Ambient: quiet keyboard clicks, soft fan hum, distant city rain.
[3–6 seconds] Slow push-in over the shoulder. On screen, glowing matrix tiles and magenta wavefronts visualize a neural network training. The engineer pauses, sips coffee. Sound: gentle electronic pulse, a single soft notification chime.
[6–9 seconds] Cut to close-up of hands typing, then rack focus to a small window showing a red fox walking through digital snow inside the monitor reflection. Sound: rising synthesized tone, subtle wind.
[9–12 seconds] Smooth lateral move across the desk: terminal windows, GPU metrics, and a grid of video frames assembling on screen. Warm amber grade, volumetric dust in the lamp beam.
[12–15 seconds] Controlled pullback reveals the full workspace at rest. The engineer leans back, satisfied. Sound: clean final impact, room tone fades.
No readable text, no logos, no subtitles. Premium technology documentary aesthetic." \
--width 864 --height 480 --seconds 15 \
--steps 20 --layers 45 --reuse 2 --seed 42 \
-o assets/showcase/long-15s-cinematic.mp4Current tagged line is v0.15.0. h3 --info prints h3-hip 0.15.0.
| Capability | Status |
|---|---|
| T2VA (text → video+audio) | ✅ |
Dual HIP ISA (HIP_ARCH=gfx1151 / gfx90a / gfx942) |
✅ |
FL2VA (--first-frame / --last-frame) |
✅ |
Ref2VA (--ref-image, --ref-silent-video, --ref-video, --ref-audio) |
✅ |
| Runtime INT8 DiT (hipBLAS) | ✅ default on all ISAs (H3_INT8_MLP=0 for BF16) |
--frames-dir / --ssd-streaming |
✅ |
--token-reduction |
✅ opt-in; off by default; H3_TOKEN_REDUCTION_SCHEDULE for per-step control |
--sol-attn |
✅ opt-in lossy long SDPA on Strix Halo + MI210 + MI300X (measured) |
--fbc |
✅ opt-in lossy first-block cache on gfx1151 and gfx942 (measured). Best on 50-step reuse-1 clips; small on 15 s reuse-2. Not with --token-reduction |
--fasth3-lora |
✅ opt-in dense 4-step LoRA. MI300X only in v0.15.0. Forces 999/749/500/250, 50 layers, --reuse 1 |
--taeh3 |
✅ opt-in tiny video decoder. Softer picture. Audio VAE unchanged |
--vsa |
✅ opt-in sparse video attention. Requires the vsa-datafree --fasth3-lora. Not with token reduction, Sol-Attn, --fbc, or reuse above 1 |
--serve HTTP daemon (protocol v1alpha) |
--serve starts a loopback HTTP daemon so dsh-plugin-h3-hip
(or any v1alpha client) can submit jobs without changing the CLI generate path.
Protocol v1alpha is not stable; a breaking change increments the protocol
version. Bind stays on 127.0.0.1 unless you set H3D_BIND (do not expose
this port on a public interface).
MODEL=/path/to/MiniMax-H3
./h3 -d "$MODEL" --serve
# GET http://127.0.0.1:8571/v1/info with header X-H3-Protocol: v1alphaUseful env: H3D_BIND / H3D_PORT (default 127.0.0.1:8571),
H3D_MODEL_PATH (or H3D_MODEL_PATH_FL2VA / H3D_MODEL_PATH_REF2VA),
H3D_OUTPUT_ROOT, H3D_MEDIA_ROOT, H3D_GPUS, H3D_QUOTA_BYTES.
Full contract: docs/DESIGN_DHS.md.
- Project page (speedup ladder and clips): alexhegit.github.io/h3-hip.c
- Wiki sources (also published to the GitHub wiki): Getting started, T2VA pipeline, Long video
- Timings (Strix Halo / MI210 / MI300X):
docs/PERFORMANCE.md - Best practice (recommended settings):
docs/BEST_PRACTICE.md --sol-attn(lossy long SDPA; off by default):docs/SOL_ATTN.md--fbc(lossy first-block cache; off by default):docs/PERFORMANCE.md- Known gaps:
docs/KNOWN_ISSUES.md --servedaemon / DSH plugin:docs/DESIGN_DHS.md
The fox showcase uses --steps 20 --layers 50 --reuse 1. Tagged scoreboard
commands are in PERFORMANCE.md. --token-reduction is the same opt-in
speed flag as h3-spark.c (pair
video tokens in middle DiT blocks). It is off by default. Use it when
wall clock matters more than fox-s2 bit identity — long T2VA is DiT-SDPA
bound on all ISAs. Strix Halo fox-fast denoise with CLI TR was 34.6 s →
25.8 s (v0.9.0); on 2026-09-12 main default fox-fast denoise is 26.4 s.
Same 15 s cinematic: Strix Halo (gfx1151) no-TR 40 min 4 s;
./bench/fox-15s.sh 26 min 56 s; ./bench/fox-15s-fast.sh 20 min 44 s.
MI210 (gfx90a) no-TR 12 min 12 s; fox-15s.sh 8 min 13 s;
fox-15s-fast.sh 6 min 18 s.
MI300X (gfx942) no-TR 199 s; fox-15s.sh 149 s;
fox-15s-fast.sh 118 s.
--sol-attn is a separate opt-in. Default stays dense (quality). 15 s no-TR
versus dense: MI210 E2E −18% / SDPA −29% (18.73 dB / 6.69 dB SNR);
MI300X E2E −11% / SDPA −33% (19.19 dB / 8.69 dB SNR); Strix Halo
E2E −27% / SDPA −42% (19.55 dB / 10.99 dB SNR). MI300X
1344×768 · 15 s: Sol-Attn E2E 719 s vs dense 947 s (−24%).
Do not enable for
publication or audio-sensitive output. Details:
docs/SOL_ATTN.md.
Tagged scoreboard stays without TR. Per-step schedule is on-demand only:
./bench/fox-15s-sched.sh (not the default bench/ retune). Generate prints a
stderr warning when --token-reduction or H3_TOKEN_REDUCTION_SCHEDULE is on.
Wiki pages not mirrored under docs/wiki/ (Home, CLI, Showcase, Performance,
Known issues) live only on GitHub wiki. In-tree copies of Getting started,
T2VA pipeline, and Long video are under docs/wiki/. After this retake,
republish the GitHub wiki Performance / Home pages so they do not still
say MI300X fox-fast “5.2 s E2E” or 15 s “179.5 s E2E” (those were DiT total).
-
Linux + ROCm (
hipcc,libamdhip64). Three HIP offload ISAs; you must passHIP_ARCHto match the GPU you are building for. Other AMD targets are still experimental.- Strix Halo (gfx1151) — RDNA. Runtime: INT8 DiT + rocWMMA SDPA.
- MI210 (gfx90a) — CDNA2. Runtime: BF16 DiT + MFMA flash SDPA. MI250 / MI250X share this ISA; they were not timed.
- MI300X (gfx942) — CDNA3. Same CDNA kernel paths as MI210.
-
Official BF16 checkpoint at
MiniMax-H3/(FL2VA/*, optionalRef2VA/*). The full repo is about 464 GiB; this tree does not read the roottransformer/,transformer_ref/,text_encoder/, orvae/copies. Partial download:pip install -U "huggingface_hub[cli]" ./tools/download_weights.sh --dir ./MiniMax-H3 # T2VA / FL2VA, ~134 GiB ./tools/download_weights.sh --dir ./MiniMax-H3 --both # also Ref2VA, ~268 GiB export H3_MODEL=$PWD/MiniMax-H3 ./tools/download_fasth3_lora.sh --dir ./FastH3-4-step-LoRA # optional 4-step LoRA, ~1.4 GiB
-
FFmpeg / FFprobe on
PATH -
ICU (
libicu-dev)
Pick the ISA that matches rocminfo / hipGetDeviceProperties on this
machine. The Makefile does not probe the GPU.
git clone https://github.com/alexhegit/h3-hip.c.git
cd h3-hip.c
git checkout v0.14.0
# Strix Halo
make HIP_ARCH=gfx1151 -j$(nproc) h3
# MI210 / MI250X
make HIP_ARCH=gfx90a -j$(nproc) h3
# MI300X
make HIP_ARCH=gfx942 -j$(nproc) h3
./h3 --info -d /path/to/MiniMax-H3make clean does not need HIP_ARCH. After changing arch, run make clean
before rebuilding. Halo fox-s2 md5 gate: make HIP_ARCH=gfx1151 halo-regression
(see tools/halo_regression.sh).
Bind one card of a multi-GPU box with H3_HIP_DEVICE=N (default 0) or
HIP_VISIBLE_DEVICES=N. On the four-GPU MI210 box: GPU 0 compile/debug,
GPU 1 quality gates, GPU 2 long T2VA, GPU 3 spare. Do not run two
weight-streaming T2VA jobs at once.
Apple Metal sources remain in-tree for reference against antirez/h3.c
and are not the Linux build path.
Same fox preset as the T2VA showcase clip:
./h3 --profile \
-d /path/to/MiniMax-H3 \
-p "A red fox walks through fresh snow." \
--width 512 --height 512 \
--frames 22 --steps 20 \
--layers 50 --reuse 1 \
-o outputs/fox.mp4First run pays model load from disk (~107 GiB on this T2VA path). On Strix Halo
(gfx1151) the BIOS carveout leaves ~31 GiB host RAM, so repeat runs still miss most of
the page cache. MI210 (gfx90a) is a discrete GPU; its E2E is still dominated by weight
I/O on fox-s2. For the tagged scoreboard commands, see
docs/PERFORMANCE.md.
Ordered references (--ref-image, --ref-video, --ref-audio, …) and
first/last-frame anchors (--first-frame, --last-frame) follow the same CLI
as antirez/h3.c.
--ref-video with an embedded soundtrack needs at least 2 seconds of audio
(request --frames 56 or more). --ref-video-audio VIDEO AUDIO is the same
generation path with the soundtrack supplied as a separate file.
--frames-dir DIR— write each decoded frame as PPM (works everywhere)--show— live denoise preview in Kitty/Ghostty/iTerm2/WezTerm/Konsole. Override detection:H3_TERMINAL=kitty(useful over SSH). Default terminal zoom: 1× on Linux, 2× on macOS.
make HIP_ARCH=gfx1151 -j$(nproc) test # Strix Halo
make HIP_ARCH=gfx90a -j$(nproc) test # MI210
make HIP_ARCH=gfx1151 hip-functional
make HIP_ARCH=gfx1151 halo-regression # Halo fox-s2 md5 gate
make HIP_ARCH=gfx90a hip-test # MI210 unit + smokesSet H3_MODEL=/path/to/MiniMax-H3 if weights are not at the Makefile default.
MLX toy-block fixtures are optional; without misc/fixtures/ those parity
targets are skipped.
Host/model/CLI code follows antirez/h3.c.
The Linux build uses the HIP backend (backends/h3_gpu_hip.c,
kernels/h3_kernels.hip, kernels/h3_kernels_extra.hip). Pass HIP_ARCH
explicitly (gfx1151 or gfx90a). Apple Metal sources (h3_gpu.m,
h3_shaders.metal) remain for reference against the original project and are
not the HIP build path. User-facing extra docs are under docs/wiki/,
docs/PERFORMANCE.md, and docs/KNOWN_ISSUES.md.
See LICENSE. Model weights are subject to the MiniMax-H3 license on
Hugging Face.








