Skip to content
alexhegitPublic

About

HIP port of antirez/h3.c (MiniMax-H3) for AMD GPUs

Resources

Stars

11 stars

Watchers

1 watching

Forks

Latest commit

 

History

339 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

h3-hip.c

HIP port of antirez/h3.c for AMD GPUs. One tree, three supported products — pass HIP_ARCH to match the GPU you compile for. Timed scoreboard SKUs are Strix Halo (gfx1151), MI210 (gfx90a), and MI300X (gfx942). MI250 / MI250X also use gfx90a but were not timed.

Product HIP_ARCH Default DiT SDPA
Strix Halo (RDNA) gfx1151 INT8 + BF16 activations wave32 rocWMMA
MI210 (CDNA2) gfx90a INT8 (default) wave64 MFMA flash
MI300X (CDNA3) gfx942 INT8 (default) wave64 MFMA flash

Tagged v0.15.0 adds three opt-in FastH3 paths, all off by default. --fasth3-lora merges a dense 4-step LoRA and runs timesteps 999 / 749 / 500 / 250. --taeh3 replaces the video VAE with a tiny decoder. --vsa turns on sparse video attention and requires the vsa-datafree adapter. They do not combine with --token-reduction, --sol-attn, --fbc, or --reuse above 1, and they keep all 50 layers. On MI300X the 832×480 · 5 s INT8 process is dense 166.87 s, --fbc 53.31 s, FastH3 30.12 s, FastH3+TAEH3 27.72 s, VSA 32.61 s, VSA+TAEH3 26.02 s. The fox-15s canvas (864×480 · 15 s; FastH3 uses 4 steps and reuse 1, not the script's 20 steps / 45 layers / reuse 2) is dense 213.3 s, fox-15s.sh 146.21 s, --fbc 136.43 s, FastH3 108.68 s, FastH3+TAEH3 89.12 s, VSA 95.06 s, VSA+TAEH3 76.72 s. The 4-step picture follows the prompt and is a different composition from dense; TAEH3 is softer, and the 15 s VSA ending has a second person in the office. v0.14.0 remains the --fbc release. v0.13.0 remains the quality-path CDNA flash SDPA and Sol-Attn-past-64k release. INT8 DiT is default on all ISAs (H3_INT8_MLP=0 for BF16). v0.12.0 was TR schedule + fused AdaLN + 3-GPU retune; v0.11.x had opt-in INT8; v0.10.x is the dual-ISA history; v0.9.x is Strix Halo only. The original project is a native MiniMax-H3 inference engine (Apple Metal / macOS); this repository reimplements the GPU backend in pure HIP so the same CLI and model stack run on ROCm.

h3-hip.c ident

Project ident, generated on Strix Halo (gfx1151) (864×480, 56 frames, --steps 20 --layers 50 --reuse 1). Click the poster for the MP4: h3-hip-ident.mp4.

Project page: alexhegit.github.io/h3-hip.c Wiki: github.com/alexhegit/h3-hip.c/wiki Original project: antirez/h3.c CUDA sibling (DGX Spark): alexhegit/h3-spark.c Official weights: MiniMaxAI/MiniMax-H3

Headline T2VA (same knobs; details in docs/PERFORMANCE.md):

Preset Strix Halo (gfx1151) MI210 (gfx90a) MI300X (gfx942)
fox-s2 ~85–90 s (I/O) 11.2 s 10.6 s
fox-fast ~2 min (I/O) 19.5 s 9.8 s
15 s cinematic (864×480, 362 f) 40 min 4 s (no TR) / 26 min 56 s (fox-15s.sh) 12 min 12 s (no TR) / 8 min 13 s (fox-15s.sh) 199 s (no TR) / 149 s (fox-15s.sh)

These are complete muxed MP4s (video + audio), not stubs. Headline cells are process E2E (/usr/bin/time until the MP4 is written) except Halo short clips (I/O-bound). MI300X short clips are the 2026-09-18 v0.12.0 retake (b0ff702): fox-s2 warm n=5 mean 10.6 s (denoise 0.29 s); fox-fast warm n=5 mean 9.8 s (denoise 1.94 s, DiT total 5.0 s). v0.13.0 15 s no-TR is 199 s (quality-path SDPA A/B 215 → 199 s; v0.12.0 was 213 s). fox-15s-fast.sh process E2E is 118 s. Do not quote DiT total as time-to-MP4 (fox-fast DiT 5.0 s). Cold fox-fast is ~20 s. fox-s2 and fox-fast are both 512² · 22 frames (~0.9 s at 24 fps). The README fox showcase clip is a third preset (--layers 50 --reuse 1). Ledger: docs/perf-runs/MI300X_2026-09-18_full.md and docs/perf-runs/MI300X_2026-09-21_1344x768-15s-sdpa.md.

Name Size Knobs Role
fox-s2 512² · 22 f --steps 2 --layers 35 --reuse 1 HIP A/B + gfx1151 md5 gate
fox-fast 512² · 22 f --steps 20 --layers 45 --reuse 2 Complete short clip; 11 DiT evals; upstream “first fast video”
fox showcase 512² · 22 f --steps 20 --layers 50 --reuse 1 README / wiki gallery fox
15 s cinematic 864×480 · 362 f --steps 20 --layers 45 --reuse 2 Same quality path as fox-fast, long duration

Showcase (Strix Halo)

Clips below were generated on an AMD Strix Halo iGPU (gfx1151) with this HIP port. Click a poster for the MP4. Several clips are untitled model output (no ffmpeg captions). The 15 s three-way reel is labelled left-to-right.

Mode Sample
T2VA — fox tutorial clip T2VA fox mp4
T2VA — Linux 35 tribute (untitled) Linux 35 mp4
T2VA — project ident draft (untitled) ident draft mp4
Ref2VA — AMD developer community (untitled) AMD community mp4
T2VA — 15 s cinematic office (untitled) 15 s long mp4
T2VA — 15 s Halo three-way (dense / TR / Sol-Attn) 15 s 3-way mp4
T2VA — 15 s MI300X FastH3 (4-step / 4-step+TAEH3) FastH3 2-way mp4
T2VA — 10 s cinematic office (untitled) 10 s long mp4

Long clips (864×480, --steps 20 --layers 45 --reuse 2): 15 s E2E 40 min 4 s without TR (main 2026-09-12); ./bench/fox-15s.sh (TR 4:30) is 26 min 56 s on Strix Halo (gfx1151); 15 s E2E 12 min 12 s / fox-15s.sh 8 min 13 s on MI210 (gfx90a); 15 s process E2E 199 s / fox-15s.sh 149 s on MI300X (gfx942) / fox-15s-fast.sh 118 s. Halo same-seed 15 s lossy A/B (2026-09-16): fox-15s.sh TR is 27 min 46 s at 18.23 dB / 0.667 vs dense; --sol-attn is 29 min 47 s at 19.55 dB / 0.721 (audio SNR 3.11 dB vs 10.99 dB). TR is faster; Sol-Attn is closer to dense. Triptych: fox-15s-3way-compare-gfx1151.mp4 · docs/SOL_ATTN.md. Timings and reproduce commands: docs/perf-runs/LONG_VIDEO.md · Wiki: Long video · Project page § Long T2VA.

FL2VA and Ref2VA (image, silent video, embedded soundtrack, image+audio) are wired and smoke-tested on the same binary; see Status.

The titled ident at the top of this README is a later overlay of a 864×480 run (--layers 50 --reuse 1). The ident row in the table is the untitled 512×288 draft.

Reproduce the showcase clips

Build first (make HIP_ARCH=gfx1151 or HIP_ARCH=gfx90a). Then:

MODEL=/path/to/MiniMax-H3

# Fox (512², 22 frames, full layers)
./h3 -d "$MODEL" \
  -p "A red fox walks through fresh snow." \
  --width 512 --height 512 --frames 22 \
  --steps 20 --layers 50 --reuse 1 \
  -o assets/showcase/t2va-fox.mp4

# Linux 35 tribute (untitled; 864×480, 56 frames)
./h3 -d "$MODEL" \
  -p "A single penguin stands on a snowy ridge at blue hour, facing a glowing amber terminal screen floating in the cold air, scrolling lines of green monospaced code reflected in its eyes. Slow dolly-in, shallow depth of field, volumetric mist, cinematic teal and amber grade. Ambient wind, soft mechanical keyboard clicks, a low warm synth pad." \
  --width 864 --height 480 --seconds 2.33 \
  --steps 20 --layers 50 --reuse 1 \
  --seed 1991 \
  -o assets/showcase/linux35.mp4

# Project ident draft (untitled; 512×288, 56 frames)
./h3 -d "$MODEL" \
  -p "Cinematic technology project ident. In a dark graphite silicon landscape, streams of glowing model weights flow from an SSD into a powerful red GPU compute array. Magenta and amber wavefronts race through precise matrix tiles, code particles transform into a grid of moving video frames, and a realistic red fox briefly emerges from a snowy digital scene while a luminous audio waveform pulses beneath it. Slow controlled camera pullback reveals the complete accelerator board, elegant engineering visualization, AMD-red and HIP-magenta color palette, volumetric light, crisp reflections, premium cinematic motion graphics, no readable text, no logos. Sound: deep electronic startup pulse, rapid soft data clicks, rising synthesized tone, clean final impact." \
  --width 512 --height 288 --frames 56 \
  --steps 20 --layers 45 --reuse 2 \
  --seed 1991 \
  -o assets/showcase/h3-hip-ident-draft-raw.mp4

# AMD developer community (untitled Ref2VA; 864×480, 56 frames)
# Needs Ref2VA weights. Image encode on this HIP path is slow (~5 min per still).
./h3 -d "$MODEL" \
  -p "Cinematic AMD developer community invitation. Preserve the two referenced black developer T-shirt designs: first, the gold Helios chariot artwork on the back; second, the futuristic runner and open-world artwork on the front. Begin with a close view of the golden Helios shirt worn by a developer in a modern hardware lab. The camera makes a smooth orbit to reveal another developer wearing the runner shirt. Together they walk through a luminous orange open doorway into a welcoming diverse group of software engineers. Every engineer in the group, including the people further away in the background, wears the same matching black AMD team T-shirt with the crisp geometric AMD arrow emblem clearly printed on the chest and on the back, sharp white and orange print, consistent uniform team branding across the whole room. They gather around glowing code displays and an AMD-powered workstation. Confident, inclusive, collaborative energy, black graphite with AMD orange and electric blue accents, premium technology campaign, realistic fabric, crisp apparel print detail, cinematic lighting, no generated captions. Sound: subtle server room ambience, keyboard clicks, warm rising electronic pulse, uplifting final impact." \
  --ref-image assets/showcase/refs/helios-shirt.png \
  --ref-image assets/showcase/refs/run-open-shirt.png \
  --ref-image-size max \
  --width 864 --height 480 --seconds 2.33 \
  --steps 20 --layers 50 --reuse 1 \
  --seed 2026 \
  -o assets/showcase/amd-developer-community-raw.mp4

# Long T2VA — 15 s cinematic office (864×480, 362 frames;
# E2E 40 min 4 s Strix Halo (gfx1151) no TR / 12 min 12 s MI210 (gfx90a);
# ./bench/fox-15s.sh → 26 min 56 s Halo / 8 min 13 s MI210
./h3 --profile -d "$MODEL" \
  -p "15 seconds, 16:9 landscape cinematic. A lone software engineer works late in a dim home office lit only by monitor glow and a desk lamp. Photoreal live-action feel with subtle handheld camera breathing.

[0–3 seconds] Medium shot from behind the desk. Code scrolls on dual monitors; warm red accent light reflects on glass. Ambient: quiet keyboard clicks, soft fan hum, distant city rain.

[3–6 seconds] Slow push-in over the shoulder. On screen, glowing matrix tiles and magenta wavefronts visualize a neural network training. The engineer pauses, sips coffee. Sound: gentle electronic pulse, a single soft notification chime.

[6–9 seconds] Cut to close-up of hands typing, then rack focus to a small window showing a red fox walking through digital snow inside the monitor reflection. Sound: rising synthesized tone, subtle wind.

[9–12 seconds] Smooth lateral move across the desk: terminal windows, GPU metrics, and a grid of video frames assembling on screen. Warm amber grade, volumetric dust in the lamp beam.

[12–15 seconds] Controlled pullback reveals the full workspace at rest. The engineer leans back, satisfied. Sound: clean final impact, room tone fades.

No readable text, no logos, no subtitles. Premium technology documentary aesthetic." \
  --width 864 --height 480 --seconds 15 \
  --steps 20 --layers 45 --reuse 2 --seed 42 \
  -o assets/showcase/long-15s-cinematic.mp4

Status

Current tagged line is v0.15.0. h3 --info prints h3-hip 0.15.0.

Capability Status
T2VA (text → video+audio) ✅
Dual HIP ISA (HIP_ARCH=gfx1151 / gfx90a / gfx942) ✅
FL2VA (--first-frame / --last-frame) ✅
Ref2VA (--ref-image, --ref-silent-video, --ref-video, --ref-audio) ✅
Runtime INT8 DiT (hipBLAS) ✅ default on all ISAs (H3_INT8_MLP=0 for BF16)
--frames-dir / --ssd-streaming ✅
--token-reduction ✅ opt-in; off by default; H3_TOKEN_REDUCTION_SCHEDULE for per-step control
--sol-attn ✅ opt-in lossy long SDPA on Strix Halo + MI210 + MI300X (measured)
--fbc ✅ opt-in lossy first-block cache on gfx1151 and gfx942 (measured). Best on 50-step reuse-1 clips; small on 15 s reuse-2. Not with --token-reduction
--fasth3-lora ✅ opt-in dense 4-step LoRA. MI300X only in v0.15.0. Forces 999/749/500/250, 50 layers, --reuse 1
--taeh3 ✅ opt-in tiny video decoder. Softer picture. Audio VAE unchanged
--vsa ✅ opt-in sparse video attention. Requires the vsa-datafree --fasth3-lora. Not with token reduction, Sol-Attn, --fbc, or reuse above 1
--serve HTTP daemon (protocol v1alpha) ⚠️ experimental; loopback only; see Daemon

Daemon (experimental)

--serve starts a loopback HTTP daemon so dsh-plugin-h3-hip (or any v1alpha client) can submit jobs without changing the CLI generate path. Protocol v1alpha is not stable; a breaking change increments the protocol version. Bind stays on 127.0.0.1 unless you set H3D_BIND (do not expose this port on a public interface).

MODEL=/path/to/MiniMax-H3
./h3 -d "$MODEL" --serve
# GET http://127.0.0.1:8571/v1/info  with header X-H3-Protocol: v1alpha

Useful env: H3D_BIND / H3D_PORT (default 127.0.0.1:8571), H3D_MODEL_PATH (or H3D_MODEL_PATH_FL2VA / H3D_MODEL_PATH_REF2VA), H3D_OUTPUT_ROOT, H3D_MEDIA_ROOT, H3D_GPUS, H3D_QUOTA_BYTES. Full contract: docs/DESIGN_DHS.md.

Documentation

The fox showcase uses --steps 20 --layers 50 --reuse 1. Tagged scoreboard commands are in PERFORMANCE.md. --token-reduction is the same opt-in speed flag as h3-spark.c (pair video tokens in middle DiT blocks). It is off by default. Use it when wall clock matters more than fox-s2 bit identity — long T2VA is DiT-SDPA bound on all ISAs. Strix Halo fox-fast denoise with CLI TR was 34.6 s → 25.8 s (v0.9.0); on 2026-09-12 main default fox-fast denoise is 26.4 s. Same 15 s cinematic: Strix Halo (gfx1151) no-TR 40 min 4 s; ./bench/fox-15s.sh 26 min 56 s; ./bench/fox-15s-fast.sh 20 min 44 s. MI210 (gfx90a) no-TR 12 min 12 s; fox-15s.sh 8 min 13 s; fox-15s-fast.sh 6 min 18 s. MI300X (gfx942) no-TR 199 s; fox-15s.sh 149 s; fox-15s-fast.sh 118 s. --sol-attn is a separate opt-in. Default stays dense (quality). 15 s no-TR versus dense: MI210 E2E −18% / SDPA −29% (18.73 dB / 6.69 dB SNR); MI300X E2E −11% / SDPA −33% (19.19 dB / 8.69 dB SNR); Strix Halo E2E −27% / SDPA −42% (19.55 dB / 10.99 dB SNR). MI300X 1344×768 · 15 s: Sol-Attn E2E 719 s vs dense 947 s (−24%). Do not enable for publication or audio-sensitive output. Details: docs/SOL_ATTN.md. Tagged scoreboard stays without TR. Per-step schedule is on-demand only: ./bench/fox-15s-sched.sh (not the default bench/ retune). Generate prints a stderr warning when --token-reduction or H3_TOKEN_REDUCTION_SCHEDULE is on.

Wiki pages not mirrored under docs/wiki/ (Home, CLI, Showcase, Performance, Known issues) live only on GitHub wiki. In-tree copies of Getting started, T2VA pipeline, and Long video are under docs/wiki/. After this retake, republish the GitHub wiki Performance / Home pages so they do not still say MI300X fox-fast “5.2 s E2E” or 15 s “179.5 s E2E” (those were DiT total).

Requirements

  • Linux + ROCm (hipcc, libamdhip64). Three HIP offload ISAs; you must pass HIP_ARCH to match the GPU you are building for. Other AMD targets are still experimental.

    • Strix Halo (gfx1151) — RDNA. Runtime: INT8 DiT + rocWMMA SDPA.
    • MI210 (gfx90a) — CDNA2. Runtime: BF16 DiT + MFMA flash SDPA. MI250 / MI250X share this ISA; they were not timed.
    • MI300X (gfx942) — CDNA3. Same CDNA kernel paths as MI210.
  • Official BF16 checkpoint at MiniMax-H3/ (FL2VA/*, optional Ref2VA/*). The full repo is about 464 GiB; this tree does not read the root transformer/, transformer_ref/, text_encoder/, or vae/ copies. Partial download:

    pip install -U "huggingface_hub[cli]"
    ./tools/download_weights.sh --dir ./MiniMax-H3          # T2VA / FL2VA, ~134 GiB
    ./tools/download_weights.sh --dir ./MiniMax-H3 --both   # also Ref2VA, ~268 GiB
    export H3_MODEL=$PWD/MiniMax-H3
    ./tools/download_fasth3_lora.sh --dir ./FastH3-4-step-LoRA   # optional 4-step LoRA, ~1.4 GiB
  • FFmpeg / FFprobe on PATH

  • ICU (libicu-dev)

Build

Pick the ISA that matches rocminfo / hipGetDeviceProperties on this machine. The Makefile does not probe the GPU.

git clone https://github.com/alexhegit/h3-hip.c.git
cd h3-hip.c
git checkout v0.14.0

# Strix Halo
make HIP_ARCH=gfx1151 -j$(nproc) h3

# MI210 / MI250X
make HIP_ARCH=gfx90a -j$(nproc) h3

# MI300X
make HIP_ARCH=gfx942 -j$(nproc) h3

./h3 --info -d /path/to/MiniMax-H3

make clean does not need HIP_ARCH. After changing arch, run make clean before rebuilding. Halo fox-s2 md5 gate: make HIP_ARCH=gfx1151 halo-regression (see tools/halo_regression.sh).

Bind one card of a multi-GPU box with H3_HIP_DEVICE=N (default 0) or HIP_VISIBLE_DEVICES=N. On the four-GPU MI210 box: GPU 0 compile/debug, GPU 1 quality gates, GPU 2 long T2VA, GPU 3 spare. Do not run two weight-streaming T2VA jobs at once. Apple Metal sources remain in-tree for reference against antirez/h3.c and are not the Linux build path.

Quick generate (T2VA)

Same fox preset as the T2VA showcase clip:

./h3 --profile \
  -d /path/to/MiniMax-H3 \
  -p "A red fox walks through fresh snow." \
  --width 512 --height 512 \
  --frames 22 --steps 20 \
  --layers 50 --reuse 1 \
  -o outputs/fox.mp4

First run pays model load from disk (~107 GiB on this T2VA path). On Strix Halo (gfx1151) the BIOS carveout leaves ~31 GiB host RAM, so repeat runs still miss most of the page cache. MI210 (gfx90a) is a discrete GPU; its E2E is still dominated by weight I/O on fox-s2. For the tagged scoreboard commands, see docs/PERFORMANCE.md.

Conditional paths

Ordered references (--ref-image, --ref-video, --ref-audio, …) and first/last-frame anchors (--first-frame, --last-frame) follow the same CLI as antirez/h3.c.

--ref-video with an embedded soundtrack needs at least 2 seconds of audio (request --frames 56 or more). --ref-video-audio VIDEO AUDIO is the same generation path with the soundtrack supplied as a separate file.

Preview while generating

  • --frames-dir DIR — write each decoded frame as PPM (works everywhere)
  • --show — live denoise preview in Kitty/Ghostty/iTerm2/WezTerm/Konsole. Override detection: H3_TERMINAL=kitty (useful over SSH). Default terminal zoom: 1× on Linux, 2× on macOS.

Tests

make HIP_ARCH=gfx1151 -j$(nproc) test          # Strix Halo
make HIP_ARCH=gfx90a  -j$(nproc) test          # MI210
make HIP_ARCH=gfx1151 hip-functional
make HIP_ARCH=gfx1151 halo-regression          # Halo fox-s2 md5 gate
make HIP_ARCH=gfx90a  hip-test                 # MI210 unit + smokes

Set H3_MODEL=/path/to/MiniMax-H3 if weights are not at the Makefile default. MLX toy-block fixtures are optional; without misc/fixtures/ those parity targets are skipped.

Repository layout

Host/model/CLI code follows antirez/h3.c. The Linux build uses the HIP backend (backends/h3_gpu_hip.c, kernels/h3_kernels.hip, kernels/h3_kernels_extra.hip). Pass HIP_ARCH explicitly (gfx1151 or gfx90a). Apple Metal sources (h3_gpu.m, h3_shaders.metal) remain for reference against the original project and are not the HIP build path. User-facing extra docs are under docs/wiki/, docs/PERFORMANCE.md, and docs/KNOWN_ISSUES.md.

License

See LICENSE. Model weights are subject to the MiniMax-H3 license on Hugging Face.

About

HIP port of antirez/h3.c (MiniMax-H3) for AMD GPUs

Resources

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages