Add GraniteHybrid routed-MoE GGUF support - #615
Merged
Conversation
Performance Comparison
|
justinchuby
force-pushed
the
justinchuby-complete-nemotron-h-moe-gguf
branch
from
August 25, 2026 18:20
14d9be9 to
41ed446
Compare
Base automatically changed from
justinchuby-complete-nemotron-h-moe-gguf
to
main
August 25, 2026 18:21
justinchuby
force-pushed
the
justinchuby-granitehybrid-moe-gguf
branch
from
August 25, 2026 18:31
836595e to
9cbc655
Compare
Import mixed Mamba2/attention GraniteHybrid GGUFs with exact routed and shared expert geometry, strict tensor closure, and fail-closed quantized preservation. Add end-to-end state, parity, persistence, CLI, and malformed-source coverage, and update the GGUF support census. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby
force-pushed
the
justinchuby-granitehybrid-moe-gguf
branch
from
August 25, 2026 18:33
9cbc655 to
8690dcd
Compare
Contributor
There was a problem hiding this comment.
Pull request overview
This PR extends Mobius GGUF import support for GraniteHybrid (GraniteMoeHybrid) to cover routed-MoE variants, including mixed Mamba2/attention schedules, optional shared experts, and stricter closure/shape validation consistent with official GraniteHybrid GGUF contracts.
Changes:
- Add exact GraniteHybrid GGUF config extraction for routed experts + optional shared expert geometry and scaling/bias metadata.
- Add GGUF tensor mapping + processing to fuse routed expert gate/up tensors (and shared gate/up) into the fused weights expected by the Mobius model.
- Add extensive tests covering tensor fusion order, strict closure/shape validation, ORT prefill/decode/reorder/replay, CLI build behavior, and malformed-source rejection.
Reviewed changes
Copilot reviewed 12 out of 12 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| tests/synthetic_parity_test.py | Updates HF config adaptation for GraniteMoeHybrid layer type vocabulary used in synthetic parity tests. |
| src/mobius/models/granitemoehybrid.py | Wires routed scaling into the gate and supports optional shared MLP for routed-MoE GraniteHybrid variants. |
| src/mobius/integrations/gguf/_tensor_processors.py | Adds GraniteHybrid expert gate/up fusion into the fused expert tensor layout. |
| src/mobius/integrations/gguf/_tensor_processors_test.py | Adds a value test ensuring expert gate/up fusion preserves expert ordering and gate-then-up layout. |
| src/mobius/integrations/gguf/_tensor_mapping.py | Adds GraniteHybrid routed-MoE tensor aliases (router + expert projections + optional shared expert projections). |
| src/mobius/integrations/gguf/_tensor_mapping_test.py | Tests the new GraniteHybrid routed-MoE mapping aliases. |
| src/mobius/integrations/gguf/_config_mapping.py | Implements GraniteHybrid routed-MoE config extraction/validation and captures bias/scaling metadata. |
| src/mobius/integrations/gguf/_config_mapping_test.py | Adds GraniteHybrid routed-MoE config extraction and invalid-metadata rejection tests. |
| src/mobius/integrations/gguf/_builder.py | Extends hybrid GGUF tensor-closure + logical-shape validation for GraniteHybrid routed-MoE and optional shared expert tensors. |
| src/mobius/integrations/gguf/_builder_test.py | Adds end-to-end GraniteHybrid routed-MoE GGUF build/ORT/state/CLI tests plus malformed-source rejection cases. |
| src/mobius/integrations/gguf/_arch_registry.py | Updates GraniteHybrid architecture support reason string to reflect routed-MoE support and dequantization constraints. |
| docs/api/build_from_gguf.md | Updates the published GGUF support census and second-hybrid-cohort documentation to reflect routed-MoE GraniteHybrid support. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+1901
to
+1905
| if architecture == "granitehybrid" and not int( | ||
| metadata.get("granitehybrid.expert_count", 0) | ||
| ): | ||
| require_all_or_none( | ||
| "attention output projection", | ||
| [f"blk.{index}.attn_output.bias" for index in attention_layers], | ||
| "dense shared-MLP", |
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - add a dedicated MiniMax configuration and exact hybrid Lightning/full-attention graph matching pinned llama.cpp `8d9af256337d1a501250f9bbf4c0859a654bddd6` - add fail-closed GGUF metadata, schedule, tensor-family, shape, and closure validation with exact float and quantized tensor mapping - preserve quantized projection/expert and embedding roles, tied packed embedding/output storage, FP32 router semantics, and heterogeneous recurrent/KV graph state - add synthetic Transformers parity plus ORT prefill/decode, replay, batch reorder, float/quantized import, CLI, and save/reload coverage - update the GGUF support census and pin the public MiniMax reference revision ## Runtime scope Released ORT GenAI packaging remains deferred to #605 because MiniMax requires heterogeneous KV/recurrent state slots plus bounded rollback snapshots that the current package schema cannot faithfully represent. Graph-call state threading, deterministic replay, and batch reorder are covered here. No small public exact checkpoint exists; the approximately 456B public checkpoint is distributed across 413 shards, so real-checkpoint parity and OGA E2E are not locally feasible. ## Validation - `51 passed, 4 skipped` — focused MiniMax config, graph, GGUF closure/import, synthetic parity, weight alignment, Lightning, CLI/roundtrip, and ORT state tests - `7626 passed, 56 skipped` — broad non-integration suite - initialized `lintrunner -a` — clean - generated GGUF documentation and `git diff --check` — clean - GPT-5.6 Sol medium architecture and GGUF reviews completed; all findings resolved, including quantized/tied embedding execution and stale runtime-deferral text Rebased linearly onto merged #615 (`cc80c77db0e8e176a04e66611423399270bcfa63`). --------- Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: aedc6984-fda0-45fc-bddb-a6317953da96
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - clean up still-valid typing, documentation, test robustness, error-message, maintainability, and performance findings left by the GGUF PR stack - preserve the behavioral fixes merged in #625-#632 - hash reused multi-GB GGUF sources once, at the final pre-publication integrity gate, while retaining cheap identity checks around staging Exact base: `2db9d33debdc254d879a51b14434c9a81c230f4f` Exact head: `4541ed2bc9d2ab4227484510b4a85b2d9113eb25` ## Reconstructed original 17-item low-priority tranche The persisted audit retained only the totals, so this list was reconstructed from the live threads and current source. All but the already-fixed #596 comment are addressed in this PR. | PR | Comment | Disposition | |---|---:|---| | #550 | 3837144656 | Implemented: correct tuple return annotation | | #552 | 3837389183 | Implemented: multichannel waveform shape docs | | #559 | 3837716261 | Implemented: name-based cache assertions | | #573 | 3854308350 | Implemented: one final GGUF hash, with integrity regression coverage | | #574 | 3854345521 | Implemented: `TensorRole \| None` typing | | #574 | 3854345597 | Implemented: removed obsolete verdict filtering | | #577 | 3854472859 | Implemented: documented SSM sequence length | | #578 | 3843112613 | Implemented: documented F64 passthrough | | #579 | 3854518332 | Implemented: generalized fused-projection error | | #580 | 3854576300 | Implemented: removed brittle node counts | | #583 | 3854681386 | Implemented: documented conditional draft outputs | | #587 | 3854816872 | Implemented: corrected MTP output contract docs | | #596 | 3846110059 | Already fixed on base: unambiguous GQA bias comment | | #600 | 3855174281 | Implemented: metadata-count-only MTP error | | #607 | 3855776048 | Implemented: stable route-field assertions | | #607 | 3855776086 | Implemented: public tensor iterator | | #609 | 3856486136 | Implemented: fail-closed LM-head comment | ## Current unresolved-thread disposition This covers all 48 Copilot threads returned by the reproducible #600-#630 query. The one human #623 thread is excluded. | PR | Comment | Current-main disposition and evidence | |---|---:|---| | #600 | 3855174177 | Already fixed by #629: package cycle and reserved-sidecar validation | | #600 | 3855174238 | Already fixed by #629: explicit MTP sidecar naming/loading | | #600 | 3855174281 | Implemented here: error no longer invents an observed block count | | #602 | 3848155961 | Outside exact stack; already fixed: top-level `expert_dtype` is classified before early return | | #602 | 3848155983 | Outside exact stack; still-valid behavioral block-quant validation, unchanged | | #602 | 3848156000 | Outside exact stack; still-valid truncated-read behavioral finding, unchanged | | #602 | 3848156022 | Outside exact stack; still-valid descriptor byte/dtype validation, unchanged | | #602 | 3848156040 | Outside exact stack; still-valid expert-bank payload validation, unchanged | | #603 | 3855249765 | Already fixed by #630: runtime preflight preserves shard sets | | #603 | 3855249840 | Already fixed by #630: success output follows durable runtime publication | | #604 | 3855343082 | Implemented here: graph-only MTP persistence distinguished from runtime rejection | | #604 | 3855343131 | Already fixed by #630: runtime success messages are atomic | | #607 | 3855776001 | Implemented here: missing generation golden skips before provenance read | | #607 | 3855776048 | Implemented here: only stable route fields are asserted | | #607 | 3855776086 | Implemented here: tensor count uses `tensor_items_raw()` | | #608 | 3856079371 | Still-valid behavioral cache-symlink containment finding; unchanged | | #608 | 3856079415 | Still-valid behavioral lowercase-digest validation finding; unchanged | | #609 | 3856486136 | Implemented here: comment matches value-preserving policy | | #610 | 3855541683 | Already fixed by #628: Falcon bias precedence is explicit | | #610 | 3855541761 | Already fixed by #628: CTRL tiny config exercises projection biases | | #611 | 3856595840 | Implemented here: runtime test resolves the distribution providing the module | | #612 | 3855677383 | Already fixed by #625: supported-version endianness detection | | #612 | 3855677427 | Implemented here: shared `INT64_MAX` sentinel | | #612 | 3855677460 | Implemented here: shared PLaMo2 width inference | | #612 | 3855677486 | Implemented here: accepted PLaMo2 activation spellings are explicit | | #613 | 3855845678 | Implemented here: canonical issue URL | | #613 | 3855845757 | Implemented here: Mamba-1 function-registration docs | | #613 | 3855845806 | Implemented here: test expects the canonical issue URL | | #614 | 3855988545 | Implemented here: removed stale Nemotron-H divergence comments | | #614 | 3855988597 | Already fixed by #628: zero-head geometry raises actionable `ValueError` | | #615 | 3856082290 | Already fixed by #626: dense GraniteHybrid bias closure | | #618 | 3856729330 | Implemented here: required routes filter ORT GenAI evidence | | #618 | 3856729409 | Implemented here: env-selected runtime version is authoritative | | #618 | 3856729490 | Stale/N/A: PR-description-only matrix claim; repository workflow claims one pinned version | | #618 | 3856729563 | Implemented here: schema tail restored to normal indentation | | #619 | 3856342484 | Already fixed on base: Kimi Linear uses `/issues/605` | | #619 | 3856342532 | Already fixed by #628: config rejects convolution kernels below 2 | | #619 | 3856342580 | Already fixed by #628: GGUF contract rejects convolution kernels below 2 | | #620 | 3855717041 | Already fixed by #627: tied LM-head-only checkpoints are retained | | #621 | 3856722324 | Already fixed by #628: Kimi-K3 required metadata is complete | | #623 | 3855931210 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939302 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939358 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939394 | N/A to current main: comment belongs to open, unmerged #623 | | #624 | 3856777152 | Implemented here: runtime compatibility reuses the emitted model type | | #625 | 3857049733 | Implemented here: unsupported header reports both endian candidates | | #629 | 3857313182 | Newer post-audit behavioral sidecar-symlink cleanup finding; unchanged | | #629 | 3857313251 | Newer post-audit cross-platform path-safety finding; unchanged | ## Validation - affected GGUF/package/ORT GenAI/model/schema tests: 1,040 passed - broad non-integration suite: 7,851 passed, 56 skipped, 1 subtest passed - generated GGUF docs checks: 7 passed - initialized `lintrunner`; full lint/format passed - GPT-5.6 Sol medium review: one integrity finding fixed; re-review found no significant issues Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stack
Validation
python -m pytest tests/build_graph_test.py tests/cli_test.py src/ -q -k 'not phi4mm and not apply_weights_unknown and not glm_moe_dsa' --tb=short -n auto— 7141 passed, 54 skippedpython -m pytest tests/synthetic_parity_test.py -k granitemoehybrid -q --tb=short— 1 passedpython -m pytest tests/weight_alignment_test.py -k granitemoehybrid -q --tb=short— 1 passedlintrunner— cleanThe unfiltered suite additionally reached 7164 passed and 54 skipped; its only failures were two reproducible pre-existing
glm_moe_dsashape-inference/checker failures unrelated to this diff.