Skip to content

Add GraniteHybrid routed-MoE GGUF support - #615

Merged
justinchuby merged 1 commit into
mainfrom
justinchuby-granitehybrid-moe-gguf
Aug 25, 2026
Merged

Add GraniteHybrid routed-MoE GGUF support#615
justinchuby merged 1 commit into
mainfrom
justinchuby-granitehybrid-moe-gguf

Conversation

@justinchuby

Copy link
Copy Markdown
Member

Summary

  • add exact GraniteMoeHybrid GGUF config extraction for mixed Mamba2/attention schedules, routed experts, optional shared experts, and Granite scaling/norm metadata
  • map and fuse separate routed/shared gate+up tensors with value-tested gate-then-up ordering while keeping quantized expert preservation fail-closed
  • enforce strict tensor closure and logical shapes, including optional all-or-none attention biases matching official Granite 4 H GGUFs
  • cover ORT prefill/decode/reorder/replay, package save/reload, CLI import, Transformers synthetic parity, and malformed sources
  • update the generated GGUF support census from dense-only to routed-MoE support

Stack

Validation

  • python -m pytest tests/build_graph_test.py tests/cli_test.py src/ -q -k 'not phi4mm and not apply_weights_unknown and not glm_moe_dsa' --tb=short -n auto — 7141 passed, 54 skipped
  • python -m pytest tests/synthetic_parity_test.py -k granitemoehybrid -q --tb=short — 1 passed
  • python -m pytest tests/weight_alignment_test.py -k granitemoehybrid -q --tb=short — 1 passed
  • focused GraniteHybrid GGUF closure/build tests — 19 passed
  • changed-file lintrunner — clean

The unfiltered suite additionally reached 7164 passed and 54 skipped; its only failures were two reproducible pre-existing glm_moe_dsa shape-inference/checker failures unrelated to this diff.

@github-actions

github-actions Bot commented Aug 25, 2026

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing 356fcdb8690dcd

Model Sub-model Changes Status

No architecture changes detected.


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

@github-actions

github-actions Bot commented Aug 25, 2026

Copy link
Copy Markdown

Performance Comparison

Comparing 356fcdb8690dcd

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0%
bert (feature-extraction) num_nodes 68 68 +0.0%
falcon model_size_bytes 364 KB 364 KB +0.0%
falcon num_nodes 66 66 +0.0%
gemma2 model_size_bytes 428 KB 428 KB +0.0%
gemma2 num_nodes 105 105 +0.0%
gpt2 model_size_bytes 388 KB 388 KB +0.0%
gpt2 num_nodes 54 54 +0.0%
llama model_size_bytes 425 KB 425 KB +0.0%
llama num_nodes 60 60 +0.0%
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0%
llama (static-cache) num_nodes 56 56 +0.0%
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0%
mamba (ssm-text-generation) num_nodes 94 94 +0.0%
phi3 model_size_bytes 421 KB 421 KB +0.0%
phi3 num_nodes 58 58 +0.0%
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0%
phi3 (static-cache) num_nodes 54 54 +0.0%
qwen2 model_size_bytes 425 KB 425 KB +0.0%
qwen2 num_nodes 60 60 +0.0%
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0%
qwen2 (static-cache) num_nodes 56 56 +0.0%
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0%
qwen3_5_moe (hybrid-text-generation) num_nodes 265 265 +0.0%
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0%
qwen3_5_text (hybrid-text-generation) num_nodes 127 127 +0.0%
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0%
qwen3_5_vl (hybrid-qwen-vl) num_nodes 429 429 +0.0%
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0%
t5 (seq2seq) num_nodes 176 176 +0.0%
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0%
whisper (speech-to-text) num_nodes 128 128 +0.0%

No performance regressions.

@justinchuby
justinchuby force-pushed the justinchuby-complete-nemotron-h-moe-gguf branch from 14d9be9 to 41ed446 Compare August 25, 2026 18:20
Base automatically changed from justinchuby-complete-nemotron-h-moe-gguf to main August 25, 2026 18:21
@justinchuby
justinchuby requested a review from a team August 25, 2026 18:21
@justinchuby
justinchuby force-pushed the justinchuby-granitehybrid-moe-gguf branch from 836595e to 9cbc655 Compare August 25, 2026 18:31
Copilot AI lite review requested due to automatic review settings August 25, 2026 18:31
Import mixed Mamba2/attention GraniteHybrid GGUFs with exact routed and shared expert geometry, strict tensor closure, and fail-closed quantized preservation. Add end-to-end state, parity, persistence, CLI, and malformed-source coverage, and update the GGUF support census.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the justinchuby-granitehybrid-moe-gguf branch from 9cbc655 to 8690dcd Compare August 25, 2026 18:33
@justinchuby
justinchuby merged commit cc80c77 into main Aug 25, 2026
7 checks passed
@justinchuby
justinchuby deleted the justinchuby-granitehybrid-moe-gguf branch August 25, 2026 18:34

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends Mobius GGUF import support for GraniteHybrid (GraniteMoeHybrid) to cover routed-MoE variants, including mixed Mamba2/attention schedules, optional shared experts, and stricter closure/shape validation consistent with official GraniteHybrid GGUF contracts.

Changes:

  • Add exact GraniteHybrid GGUF config extraction for routed experts + optional shared expert geometry and scaling/bias metadata.
  • Add GGUF tensor mapping + processing to fuse routed expert gate/up tensors (and shared gate/up) into the fused weights expected by the Mobius model.
  • Add extensive tests covering tensor fusion order, strict closure/shape validation, ORT prefill/decode/reorder/replay, CLI build behavior, and malformed-source rejection.

Reviewed changes

Copilot reviewed 12 out of 12 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
tests/synthetic_parity_test.py Updates HF config adaptation for GraniteMoeHybrid layer type vocabulary used in synthetic parity tests.
src/mobius/models/granitemoehybrid.py Wires routed scaling into the gate and supports optional shared MLP for routed-MoE GraniteHybrid variants.
src/mobius/integrations/gguf/_tensor_processors.py Adds GraniteHybrid expert gate/up fusion into the fused expert tensor layout.
src/mobius/integrations/gguf/_tensor_processors_test.py Adds a value test ensuring expert gate/up fusion preserves expert ordering and gate-then-up layout.
src/mobius/integrations/gguf/_tensor_mapping.py Adds GraniteHybrid routed-MoE tensor aliases (router + expert projections + optional shared expert projections).
src/mobius/integrations/gguf/_tensor_mapping_test.py Tests the new GraniteHybrid routed-MoE mapping aliases.
src/mobius/integrations/gguf/_config_mapping.py Implements GraniteHybrid routed-MoE config extraction/validation and captures bias/scaling metadata.
src/mobius/integrations/gguf/_config_mapping_test.py Adds GraniteHybrid routed-MoE config extraction and invalid-metadata rejection tests.
src/mobius/integrations/gguf/_builder.py Extends hybrid GGUF tensor-closure + logical-shape validation for GraniteHybrid routed-MoE and optional shared expert tensors.
src/mobius/integrations/gguf/_builder_test.py Adds end-to-end GraniteHybrid routed-MoE GGUF build/ORT/state/CLI tests plus malformed-source rejection cases.
src/mobius/integrations/gguf/_arch_registry.py Updates GraniteHybrid architecture support reason string to reflect routed-MoE support and dequantization constraints.
docs/api/build_from_gguf.md Updates the published GGUF support census and second-hybrid-cohort documentation to reflect routed-MoE GraniteHybrid support.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +1901 to +1905
if architecture == "granitehybrid" and not int(
metadata.get("granitehybrid.expert_count", 0)
):
require_all_or_none(
"attention output projection",
[f"blk.{index}.attn_output.bias" for index in attention_layers],
"dense shared-MLP",
justinchuby added a commit that referenced this pull request Aug 25, 2026
## Summary

- add a dedicated MiniMax configuration and exact hybrid
Lightning/full-attention graph matching pinned llama.cpp
`8d9af256337d1a501250f9bbf4c0859a654bddd6`
- add fail-closed GGUF metadata, schedule, tensor-family, shape, and
closure validation with exact float and quantized tensor mapping
- preserve quantized projection/expert and embedding roles, tied packed
embedding/output storage, FP32 router semantics, and heterogeneous
recurrent/KV graph state
- add synthetic Transformers parity plus ORT prefill/decode, replay,
batch reorder, float/quantized import, CLI, and save/reload coverage
- update the GGUF support census and pin the public MiniMax reference
revision

## Runtime scope

Released ORT GenAI packaging remains deferred to #605 because MiniMax
requires heterogeneous KV/recurrent state slots plus bounded rollback
snapshots that the current package schema cannot faithfully represent.
Graph-call state threading, deterministic replay, and batch reorder are
covered here.

No small public exact checkpoint exists; the approximately 456B public
checkpoint is distributed across 413 shards, so real-checkpoint parity
and OGA E2E are not locally feasible.

## Validation

- `51 passed, 4 skipped` — focused MiniMax config, graph, GGUF
closure/import, synthetic parity, weight alignment, Lightning,
CLI/roundtrip, and ORT state tests
- `7626 passed, 56 skipped` — broad non-integration suite
- initialized `lintrunner -a` — clean
- generated GGUF documentation and `git diff --check` — clean
- GPT-5.6 Sol medium architecture and GGUF reviews completed; all
findings resolved, including quantized/tied embedding execution and
stale runtime-deferral text

Rebased linearly onto merged #615
(`cc80c77db0e8e176a04e66611423399270bcfa63`).

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: aedc6984-fda0-45fc-bddb-a6317953da96
justinchuby added a commit that referenced this pull request Aug 25, 2026
## Summary

- clean up still-valid typing, documentation, test robustness,
error-message,
  maintainability, and performance findings left by the GGUF PR stack
- preserve the behavioral fixes merged in #625-#632
- hash reused multi-GB GGUF sources once, at the final pre-publication
integrity
  gate, while retaining cheap identity checks around staging

Exact base: `2db9d33debdc254d879a51b14434c9a81c230f4f`

Exact head: `4541ed2bc9d2ab4227484510b4a85b2d9113eb25`

## Reconstructed original 17-item low-priority tranche

The persisted audit retained only the totals, so this list was
reconstructed
from the live threads and current source. All but the already-fixed #596
comment are addressed in this PR.

| PR | Comment | Disposition |
|---|---:|---|
| #550 | 3837144656 | Implemented: correct tuple return annotation |
| #552 | 3837389183 | Implemented: multichannel waveform shape docs |
| #559 | 3837716261 | Implemented: name-based cache assertions |
| #573 | 3854308350 | Implemented: one final GGUF hash, with integrity
regression coverage |
| #574 | 3854345521 | Implemented: `TensorRole \| None` typing |
| #574 | 3854345597 | Implemented: removed obsolete verdict filtering |
| #577 | 3854472859 | Implemented: documented SSM sequence length |
| #578 | 3843112613 | Implemented: documented F64 passthrough |
| #579 | 3854518332 | Implemented: generalized fused-projection error |
| #580 | 3854576300 | Implemented: removed brittle node counts |
| #583 | 3854681386 | Implemented: documented conditional draft outputs
|
| #587 | 3854816872 | Implemented: corrected MTP output contract docs |
| #596 | 3846110059 | Already fixed on base: unambiguous GQA bias
comment |
| #600 | 3855174281 | Implemented: metadata-count-only MTP error |
| #607 | 3855776048 | Implemented: stable route-field assertions |
| #607 | 3855776086 | Implemented: public tensor iterator |
| #609 | 3856486136 | Implemented: fail-closed LM-head comment |

## Current unresolved-thread disposition

This covers all 48 Copilot threads returned by the reproducible
#600-#630
query. The one human #623 thread is excluded.

| PR | Comment | Current-main disposition and evidence |
|---|---:|---|
| #600 | 3855174177 | Already fixed by #629: package cycle and
reserved-sidecar validation |
| #600 | 3855174238 | Already fixed by #629: explicit MTP sidecar
naming/loading |
| #600 | 3855174281 | Implemented here: error no longer invents an
observed block count |
| #602 | 3848155961 | Outside exact stack; already fixed: top-level
`expert_dtype` is classified before early return |
| #602 | 3848155983 | Outside exact stack; still-valid behavioral
block-quant validation, unchanged |
| #602 | 3848156000 | Outside exact stack; still-valid truncated-read
behavioral finding, unchanged |
| #602 | 3848156022 | Outside exact stack; still-valid descriptor
byte/dtype validation, unchanged |
| #602 | 3848156040 | Outside exact stack; still-valid expert-bank
payload validation, unchanged |
| #603 | 3855249765 | Already fixed by #630: runtime preflight preserves
shard sets |
| #603 | 3855249840 | Already fixed by #630: success output follows
durable runtime publication |
| #604 | 3855343082 | Implemented here: graph-only MTP persistence
distinguished from runtime rejection |
| #604 | 3855343131 | Already fixed by #630: runtime success messages
are atomic |
| #607 | 3855776001 | Implemented here: missing generation golden skips
before provenance read |
| #607 | 3855776048 | Implemented here: only stable route fields are
asserted |
| #607 | 3855776086 | Implemented here: tensor count uses
`tensor_items_raw()` |
| #608 | 3856079371 | Still-valid behavioral cache-symlink containment
finding; unchanged |
| #608 | 3856079415 | Still-valid behavioral lowercase-digest validation
finding; unchanged |
| #609 | 3856486136 | Implemented here: comment matches value-preserving
policy |
| #610 | 3855541683 | Already fixed by #628: Falcon bias precedence is
explicit |
| #610 | 3855541761 | Already fixed by #628: CTRL tiny config exercises
projection biases |
| #611 | 3856595840 | Implemented here: runtime test resolves the
distribution providing the module |
| #612 | 3855677383 | Already fixed by #625: supported-version
endianness detection |
| #612 | 3855677427 | Implemented here: shared `INT64_MAX` sentinel |
| #612 | 3855677460 | Implemented here: shared PLaMo2 width inference |
| #612 | 3855677486 | Implemented here: accepted PLaMo2 activation
spellings are explicit |
| #613 | 3855845678 | Implemented here: canonical issue URL |
| #613 | 3855845757 | Implemented here: Mamba-1 function-registration
docs |
| #613 | 3855845806 | Implemented here: test expects the canonical issue
URL |
| #614 | 3855988545 | Implemented here: removed stale Nemotron-H
divergence comments |
| #614 | 3855988597 | Already fixed by #628: zero-head geometry raises
actionable `ValueError` |
| #615 | 3856082290 | Already fixed by #626: dense GraniteHybrid bias
closure |
| #618 | 3856729330 | Implemented here: required routes filter ORT GenAI
evidence |
| #618 | 3856729409 | Implemented here: env-selected runtime version is
authoritative |
| #618 | 3856729490 | Stale/N/A: PR-description-only matrix claim;
repository workflow claims one pinned version |
| #618 | 3856729563 | Implemented here: schema tail restored to normal
indentation |
| #619 | 3856342484 | Already fixed on base: Kimi Linear uses
`/issues/605` |
| #619 | 3856342532 | Already fixed by #628: config rejects convolution
kernels below 2 |
| #619 | 3856342580 | Already fixed by #628: GGUF contract rejects
convolution kernels below 2 |
| #620 | 3855717041 | Already fixed by #627: tied LM-head-only
checkpoints are retained |
| #621 | 3856722324 | Already fixed by #628: Kimi-K3 required metadata
is complete |
| #623 | 3855931210 | N/A to current main: comment belongs to open,
unmerged #623 |
| #623 | 3855939302 | N/A to current main: comment belongs to open,
unmerged #623 |
| #623 | 3855939358 | N/A to current main: comment belongs to open,
unmerged #623 |
| #623 | 3855939394 | N/A to current main: comment belongs to open,
unmerged #623 |
| #624 | 3856777152 | Implemented here: runtime compatibility reuses the
emitted model type |
| #625 | 3857049733 | Implemented here: unsupported header reports both
endian candidates |
| #629 | 3857313182 | Newer post-audit behavioral sidecar-symlink
cleanup finding; unchanged |
| #629 | 3857313251 | Newer post-audit cross-platform path-safety
finding; unchanged |

## Validation

- affected GGUF/package/ORT GenAI/model/schema tests: 1,040 passed
- broad non-integration suite: 7,851 passed, 56 skipped, 1 subtest
passed
- generated GGUF docs checks: 7 passed
- initialized `lintrunner`; full lint/format passed
- GPT-5.6 Sol medium review: one integrity finding fixed; re-review
found no significant issues

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants