Add full Jamba GGUF MoE support - #613
Merged
Merged
Conversation
Performance Comparison
|
🏗️ Architecture Diff
No architecture changes detected. ✅ Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed) |
justinchuby
force-pushed
the
justinchuby-plamo2-gguf
branch
from
August 25, 2026 17:31
64aa8fb to
7743ca4
Compare
Implement exact mixed Mamba and attention execution with routed MoE schedules, strict GGUF tensor validation, compatible quantization preservation, and heterogeneous state handling. Add numerical parity, state-threading, expert-ordering, package round-trip, and failure-path coverage while keeping generic runtime packaging deferred. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Route Jamba Mamba projections through explicit float dequantization under the new fail-closed quantized loader while retaining MatMulNBits for attention, dense FFN, and routed expert projections. Update the pinned architecture verdict to match the reviewed policy. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Validate each Jamba Mamba ssm_a tensor before graph construction so non-finite or non-negative values cannot produce invalid A_log initializers. Cover the failure path in the synthetic GGUF importer test. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby
force-pushed
the
justinchuby-complete-jamba-gguf
branch
from
August 25, 2026 18:02
d34bcd4 to
f188ad8
Compare
Contributor
There was a problem hiding this comment.
Pull request overview
Adds end-to-end Jamba GGUF support spanning hybrid Mamba-1 + attention execution, routed-MoE routing/weight handling, and strict GGUF import validation, while explicitly deferring generic ORT GenAI packaging until heterogeneous-state metadata/schema work (#605).
Changes:
- Implement Jamba hybrid graph behavior (multi-token prefill, heterogeneous conv/SSM + KV state threading, padding masking) and Jamba-specific MoE routing behavior.
- Extend GGUF config/tensor mapping + builder to support routed experts (strict closure + shape validation) and selectively preserve/dequantize quantized MatMul roles.
- Expand unit/integration-style tests and update generated GGUF support documentation to reflect the routed-MoE subset and new validation guarantees.
Reviewed changes
Copilot reviewed 15 out of 15 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/synthetic_parity_test.py | Exercises multi-token prefill in synthetic parity runs. |
| tests/build_graph_test.py | Adds Jamba schedule/routing/state ABI assertions and fused-expert preprocessing tests. |
| src/mobius/tasks/_causal_lm.py | Annotates Jamba packages with runtime deferral metadata for heterogeneous state. |
| src/mobius/tasks/_cache_utils.py | Registers additional com.microsoft function bodies needed for Mamba-1 hybrid graphs. |
| src/mobius/models/jamba.py | Implements Jamba hybrid layers, router semantics, state threading, and tied/quantized head handling. |
| src/mobius/integrations/gguf/_tensor_mapping.py | Adds GGUF→HF tensor stems for routed expert tensors. |
| src/mobius/integrations/gguf/_config_mapping.py | Derives exact routed-MoE schedule/fields from GGUF metadata + tensor presence. |
| src/mobius/integrations/gguf/_config_mapping_test.py | Validates routed schedule derivation and updates Jamba dense config fixtures. |
| src/mobius/integrations/gguf/_builder.py | Enforces Jamba tensor closure + shape checks and supports mixed quantized import policies. |
| src/mobius/integrations/gguf/_builder_test.py | Adds tiny Jamba GGUF writer + tests for reorder/replay, expert ordering, and quantized role preservation. |
| src/mobius/integrations/gguf/_arch_registry.py | Updates Jamba architecture spec reason text and quantized-import verdict. |
| src/mobius/integrations/gguf/_arch_registry_test.py | Includes Jamba in importable architectures and adjusts quantized-verdict coverage expectations. |
| src/mobius/components/_ssm.py | Extends Jamba selective scan path to multi-token execution via LinearAttention recurrence. |
| src/mobius/_configs/_base.py | Updates JambaConfig schedule derivation, adds expert layer indices, and normalizes HF-derived defaults. |
| docs/api/build_from_gguf.md | Updates generated support table + narrative to reflect routed-MoE Jamba support and deferred runtime packaging. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+407
to
+411
| if config.model_type == "jamba": | ||
| model.metadata_props["mobius.runtime_support"] = ( | ||
| "Deferred: heterogeneous attention KV and Mamba recurrent state " | ||
| "discovery is tracked by https://github.com/onnxruntime/mobius#605" | ||
| ) |
Comment on lines
414
to
418
| layer_types = getattr(config, "layer_types", None) or [] | ||
| has_deltanet = "linear_attention" in layer_types | ||
| has_lightning = "lightning_attention" in layer_types | ||
| has_mamba = "mamba" in layer_types | ||
| has_mamba2 = "mamba2" in layer_types or isinstance(config, FalconH1Config) |
Comment on lines
+4855
to
+4857
| assert model.metadata_props["mobius.runtime_support"].endswith( | ||
| "onnxruntime/mobius#605" | ||
| ) |
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - add exact `nemotron_h_moe` GGUF backbone import with mixed Mamba2/attention/dense/MoE schedule reconstruction - preserve sigmoid correction-bias routing, ReLU-squared stacked experts, shared experts, and optional latent projections with strict tensor closure - require explicit dequantization for quantized sources and fail closed on unsupported sidecars, MTP blocks, static cache, and generic task overrides - add Transformers parity, expert-order, float/Q4 import, ORT mixed-state decode, package round-trip, malformed-input, registry, and documentation coverage ## Validation - `python -m pytest src/mobius/integrations/gguf -q --tb=short` — 2623 passed - `python -m pytest tests/build_graph_test.py tests/synthetic_parity_test.py -k nemotron_h -q --tb=short` — 11 passed, 6 skipped - changed-file Ruff check and format check — passed - two independent reviews completed; all findings fixed The broad non-integration suite reached 7135 passing tests and exposed two unrelated existing `glm_moe_dsa` shape-inference/checker failures. Stacked on #613 (`justinchuby-complete-jamba-gguf`). ORT GenAI packaging remains deferred to #605 because the released runtime schema does not represent heterogeneous KV/convolution/recurrent state slots. Released 30B-A3B GGUFs also remain rejected before graph construction because they append an unsupported combined attention+MoE MTP block. --------- Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - clean up still-valid typing, documentation, test robustness, error-message, maintainability, and performance findings left by the GGUF PR stack - preserve the behavioral fixes merged in #625-#632 - hash reused multi-GB GGUF sources once, at the final pre-publication integrity gate, while retaining cheap identity checks around staging Exact base: `2db9d33debdc254d879a51b14434c9a81c230f4f` Exact head: `4541ed2bc9d2ab4227484510b4a85b2d9113eb25` ## Reconstructed original 17-item low-priority tranche The persisted audit retained only the totals, so this list was reconstructed from the live threads and current source. All but the already-fixed #596 comment are addressed in this PR. | PR | Comment | Disposition | |---|---:|---| | #550 | 3837144656 | Implemented: correct tuple return annotation | | #552 | 3837389183 | Implemented: multichannel waveform shape docs | | #559 | 3837716261 | Implemented: name-based cache assertions | | #573 | 3854308350 | Implemented: one final GGUF hash, with integrity regression coverage | | #574 | 3854345521 | Implemented: `TensorRole \| None` typing | | #574 | 3854345597 | Implemented: removed obsolete verdict filtering | | #577 | 3854472859 | Implemented: documented SSM sequence length | | #578 | 3843112613 | Implemented: documented F64 passthrough | | #579 | 3854518332 | Implemented: generalized fused-projection error | | #580 | 3854576300 | Implemented: removed brittle node counts | | #583 | 3854681386 | Implemented: documented conditional draft outputs | | #587 | 3854816872 | Implemented: corrected MTP output contract docs | | #596 | 3846110059 | Already fixed on base: unambiguous GQA bias comment | | #600 | 3855174281 | Implemented: metadata-count-only MTP error | | #607 | 3855776048 | Implemented: stable route-field assertions | | #607 | 3855776086 | Implemented: public tensor iterator | | #609 | 3856486136 | Implemented: fail-closed LM-head comment | ## Current unresolved-thread disposition This covers all 48 Copilot threads returned by the reproducible #600-#630 query. The one human #623 thread is excluded. | PR | Comment | Current-main disposition and evidence | |---|---:|---| | #600 | 3855174177 | Already fixed by #629: package cycle and reserved-sidecar validation | | #600 | 3855174238 | Already fixed by #629: explicit MTP sidecar naming/loading | | #600 | 3855174281 | Implemented here: error no longer invents an observed block count | | #602 | 3848155961 | Outside exact stack; already fixed: top-level `expert_dtype` is classified before early return | | #602 | 3848155983 | Outside exact stack; still-valid behavioral block-quant validation, unchanged | | #602 | 3848156000 | Outside exact stack; still-valid truncated-read behavioral finding, unchanged | | #602 | 3848156022 | Outside exact stack; still-valid descriptor byte/dtype validation, unchanged | | #602 | 3848156040 | Outside exact stack; still-valid expert-bank payload validation, unchanged | | #603 | 3855249765 | Already fixed by #630: runtime preflight preserves shard sets | | #603 | 3855249840 | Already fixed by #630: success output follows durable runtime publication | | #604 | 3855343082 | Implemented here: graph-only MTP persistence distinguished from runtime rejection | | #604 | 3855343131 | Already fixed by #630: runtime success messages are atomic | | #607 | 3855776001 | Implemented here: missing generation golden skips before provenance read | | #607 | 3855776048 | Implemented here: only stable route fields are asserted | | #607 | 3855776086 | Implemented here: tensor count uses `tensor_items_raw()` | | #608 | 3856079371 | Still-valid behavioral cache-symlink containment finding; unchanged | | #608 | 3856079415 | Still-valid behavioral lowercase-digest validation finding; unchanged | | #609 | 3856486136 | Implemented here: comment matches value-preserving policy | | #610 | 3855541683 | Already fixed by #628: Falcon bias precedence is explicit | | #610 | 3855541761 | Already fixed by #628: CTRL tiny config exercises projection biases | | #611 | 3856595840 | Implemented here: runtime test resolves the distribution providing the module | | #612 | 3855677383 | Already fixed by #625: supported-version endianness detection | | #612 | 3855677427 | Implemented here: shared `INT64_MAX` sentinel | | #612 | 3855677460 | Implemented here: shared PLaMo2 width inference | | #612 | 3855677486 | Implemented here: accepted PLaMo2 activation spellings are explicit | | #613 | 3855845678 | Implemented here: canonical issue URL | | #613 | 3855845757 | Implemented here: Mamba-1 function-registration docs | | #613 | 3855845806 | Implemented here: test expects the canonical issue URL | | #614 | 3855988545 | Implemented here: removed stale Nemotron-H divergence comments | | #614 | 3855988597 | Already fixed by #628: zero-head geometry raises actionable `ValueError` | | #615 | 3856082290 | Already fixed by #626: dense GraniteHybrid bias closure | | #618 | 3856729330 | Implemented here: required routes filter ORT GenAI evidence | | #618 | 3856729409 | Implemented here: env-selected runtime version is authoritative | | #618 | 3856729490 | Stale/N/A: PR-description-only matrix claim; repository workflow claims one pinned version | | #618 | 3856729563 | Implemented here: schema tail restored to normal indentation | | #619 | 3856342484 | Already fixed on base: Kimi Linear uses `/issues/605` | | #619 | 3856342532 | Already fixed by #628: config rejects convolution kernels below 2 | | #619 | 3856342580 | Already fixed by #628: GGUF contract rejects convolution kernels below 2 | | #620 | 3855717041 | Already fixed by #627: tied LM-head-only checkpoints are retained | | #621 | 3856722324 | Already fixed by #628: Kimi-K3 required metadata is complete | | #623 | 3855931210 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939302 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939358 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939394 | N/A to current main: comment belongs to open, unmerged #623 | | #624 | 3856777152 | Implemented here: runtime compatibility reuses the emitted model type | | #625 | 3857049733 | Implemented here: unsupported header reports both endian candidates | | #629 | 3857313182 | Newer post-audit behavioral sidecar-symlink cleanup finding; unchanged | | #629 | 3857313251 | Newer post-audit cross-platform path-safety finding; unchanged | ## Validation - affected GGUF/package/ORT GenAI/model/schema tests: 1,040 passed - broad non-integration suite: 7,851 passed, 56 skipped, 1 subtest passed - generated GGUF docs checks: 7 passed - initialized `lintrunner`; full lint/format passed - GPT-5.6 Sol medium review: one integrity finding fixed; re-review found no significant issues Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Validation
python -m pytest src/mobius/integrations/gguf/ -q --tb=short— 2608 passedpython -m pytest tests/build_graph_test.py tests/synthetic_parity_test.py src/mobius/integrations/gguf/_config_mapping_test.py src/mobius/integrations/gguf/_builder_test.py -k 'jamba' -q --tb=short— 34 passedpython -m pytest tests/build_graph_test.py tests/cli_test.py src/ -q -k 'not phi4mm and not apply_weights_unknown and not glm_moe_dsa' --tb=short -n auto— 7096 passed, 52 skippedlintrunnerand formatted/linted all changed Python filesThe unfiltered broad suite still has the existing
glm_moe_dsasymbolic-shape inference/checker failure (2 tests), reproduced independently and unrelated to these Jamba paths.Upstream evidence
Audited llama.cpp at
8d9af256337d1a501250f9bbf4c0859a654bddd6and current Transformers Jamba.ai21labs/Jamba-tiny-devated303361004ac875426a61675edecf8e9d976882is the smallest suitable public real-weight follow-up candidate (~637 MB), but this PR makes no real-checkpoint or runtime-generation support claim without that separate evidence.Stacked on #612 (
justinchuby-plamo2-gguf).