Add ORT GenAI end-to-end CI - #618
Merged
Merged
Conversation
Performance Comparison
|
justinchuby
force-pushed
the
justinchuby-generic-genai-configs
branch
from
August 25, 2026 19:32
c684c1e to
f43a31f
Compare
Add a pinned, network-free released-runtime matrix and a scheduled real SmolLM F16 generation lane. Drive real-model enrollment from schema-validated evidence metadata and make relevant graph, tokenizer, runtime, and workflow changes select the gate automatically. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Remove the legacy dual-version matrix and validate only the latest stable 0.15.2 release. Refresh the SmolLM route and runtime package evidence after the #611 squash merge, and keep the real artifact lane limited to scheduled, manual, or explicitly requested runs. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby
force-pushed
the
justinchuby-ort-genai-e2e-ci
branch
from
August 25, 2026 19:46
7e30dcf to
11759e9
Compare
Contributor
There was a problem hiding this comment.
Pull request overview
This PR introduces a CPU-only ORT GenAI end-to-end CI lane for Mobius, adding both a fast synthetic “network-free” decoder package test and a pinned real-artifact SmolLM GGUF route test, plus schema/evidence plumbing and “affected models” detection to ensure these lanes run when relevant runtime surfaces change.
Changes:
- Add ORT GenAI E2E test coverage (synthetic decoder + pinned SmolLM real-artifact generation) and a dedicated GitHub Actions workflow to run it.
- Extend golden-case YAML schema and GGUF runtime evidence/registry to record ORT GenAI enrollment + released capability expectations.
- Expand
detect_affected_models.py“shared_infra” classification to conservatively select the new runtime lanes when core runtime/packaging/dependency surfaces change.
Reviewed changes
Copilot reviewed 18 out of 18 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
tests/yaml_schema_test.py |
Adds a coverage gate ensuring runtime-supported routes have ORT GenAI E2E enrollment markers. |
tests/ort_genai_e2e_test.py |
New synthetic, network-free ORT GenAI decoder E2E tests (load/tokenize/prefill/decode/reload + malformed-config rejection). |
tests/gguf_small_model_runtime_integration_test.py |
Adds isolated HF cache fixture and upgrades SmolLM ORT GenAI path to the production --runtime ort-genai route with artifact verification. |
testdata/cases/schema.json |
Adds ort_genai enrollment schema plus $defs capability constraints for released runtime expectations. |
testdata/cases/causal-lm/smollm-135m-gguf-f16.yaml |
Enrolls SmolLM F16 GGUF route into the ORT GenAI real lane with evidence ID, byte budget, and capabilities. |
src/mobius/integrations/gguf/_runtime_evidence.py |
Adds ORT GenAI-specific runtime evidence record for the pinned SmolLM route. |
src/mobius/integrations/gguf/_docs_test.py |
Updates docs tests to expect the new ORT GenAI evidence ID alongside existing evidence. |
src/mobius/integrations/gguf/_arch_registry.py |
Extends runtime evidence IDs for llama/SmolLM to include ORT GenAI evidence. |
src/mobius/_testing/golden.py |
Plumbs ort_genai YAML field into loaded golden test cases. |
scripts/detect_affected_models.py |
Introduces explicit shared-infra patterns/prefixes to trigger runtime matrix runs for ORT GenAI and related surfaces. |
scripts/detect_affected_models_test.py |
Updates classification expectations and adds coverage for new shared-infra triggers. |
pyproject.toml |
Registers new pytest markers ort_genai_fast and ort_genai_real. |
docs/design/multi-tier-testing-strategy.md |
Documents the ORT GenAI fast/real lanes and local commands, plus the CUDA waiver rationale. |
docs/api/build_from_gguf.md |
Updates llama runtime support narrative to reference evidenced ONNX Runtime/ORT GenAI versions. |
CONTRIBUTING.md |
Updates contributor checklist to point to the new ORT GenAI fast/real lanes. |
.github/workflows/ort_genai_e2e.yml |
Adds a dedicated ORT GenAI E2E workflow (fast lane + scheduled/manual real lane). |
.github/workflows/main.yml |
Wires the ORT GenAI E2E workflow into main CI when detect-affected indicates relevant changes. |
.agents/skills/quality-checklist/SKILL.md |
Updates quality checklist references to the new ORT GenAI lane commands. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+130
to
+142
| required_routes = { | ||
| ( | ||
| evidence.architecture, | ||
| evidence.repository, | ||
| evidence.revision, | ||
| evidence.filename, | ||
| evidence.import_route, | ||
| ) | ||
| for spec in iter_arch_specs() | ||
| if spec.runtime is Support.SUPPORTED | ||
| for evidence_id in spec.runtime_evidence_ids | ||
| if (evidence := runtime_evidence(evidence_id)) is not None | ||
| } |
Comment on lines
+207
to
+211
| expected_version = os.environ.get("MOBIUS_EXPECTED_ORT_GENAI_VERSION") | ||
| installed_version = version("onnxruntime-genai") | ||
| if expected_version: | ||
| assert installed_version == expected_version | ||
| assert installed_version == _EXPECTED_ORT_GENAI_VERSION |
Comment on lines
+23
to
+47
| fast: | ||
| name: Fast CPU / OGA 0.15.2 | ||
| runs-on: ubuntu-latest | ||
| timeout-minutes: 15 | ||
| steps: | ||
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7 | ||
| with: | ||
| persist-credentials: false | ||
| - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7 | ||
| with: | ||
| python-version: "3.12" | ||
| cache: pip | ||
| cache-dependency-path: pyproject.toml | ||
| - name: Install pinned runtime and test dependencies | ||
| run: | | ||
| python -m pip install --disable-pip-version-check \ | ||
| "onnxruntime-genai==0.15.2" \ | ||
| "pytest==8.4.2" \ | ||
| "pytest-timeout==2.4.0" | ||
| python -m pip install --disable-pip-version-check -e . | ||
| - name: Run network-free ORT GenAI E2E | ||
| env: | ||
| HF_HUB_OFFLINE: "1" | ||
| TRANSFORMERS_OFFLINE: "1" | ||
| MOBIUS_EXPECTED_ORT_GENAI_VERSION: "0.15.2" |
Comment on lines
+513
to
536
| "architecture": { | ||
| "type": "string", | ||
| "description": "Optional registry architecture key (e.g. Qwen35MtpModel) forcing a specific module class + task at build time. Needed for auxiliary heads (DFlash, MTP) that share a base checkpoint whose architectures field would otherwise auto-route to the base model." | ||
| } | ||
| }, | ||
| "$defs": { | ||
| "ortGenaiCapabilities": { | ||
| "type": "object", | ||
| "additionalProperties": false, | ||
| "required": [ | ||
| "generic_decoder", | ||
| "state_groups" | ||
| ], | ||
| "properties": { | ||
| "generic_decoder": { | ||
| "const": true | ||
| }, | ||
| "state_groups": { | ||
| "const": false | ||
| } | ||
| } | ||
| } | ||
| } | ||
| } |
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - normalize every graph-representable decoder-only ORT GenAI config emitted by Mobius to `model.type: "decoder"` - preserve specialized types only for GPT-2's unrepresentable cache ABI, LFM2's convolution cache, Phi-3-family LongRoPE, multimodal/encoder-decoder pipelines, and auxiliary sidecar topologies - update the Gemma4 decoder-only example, CLI docs, and ORT GenAI config skill; keep #618's OGA 0.15.2 CI and evidence unchanged ## Validation - `235 passed` focused ORT GenAI config tests - `7717 passed, 56 skipped, 1 subtest passed` broad non-integration suite - `37 passed` GGUF runtime tests - `2 passed` network-free OGA 0.15.2 fast E2E - `1 passed` pinned real SmolLM F16 CPU OGA 0.15.2 E2E - Sphinx documentation build - full lintrunner formatting/lint - two GPT-5.6 Sol medium reviews; final review reported no findings Base: `e591ecfaae00d312cb64c4334b61aff0a9e87f05` Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - add a dedicated Kimi-K3 config, model, and heterogeneous-state task with exact KDA/NoPE gated-MLA scheduling, AttnRes mixing, SiTU latent MoE, shared experts, and an untied output head - import pinned llama.cpp `kimi-k3` GGUF metadata and tensors with strict metadata/tensor/shape/storage closure and malformed-input rejection - preserve lossless quantization routes under #609: separate rank-3 MLA projections fail closed unless explicitly dequantized, while fused Q4_0 KV-B is split by exact packed-row reordering across weights, scales, and zero-points - preserve valid KDA convolution history across padding, validate kernel/state/task contracts, and cover replay, reorder, float/quantized import, roundtrip, CLI, and negative paths - keep generic OGA runtime packaging truthfully deferred under #605 because released cache schemas cannot represent the heterogeneous state ABI ## Reconstruction Reconstructed after #619 was admin squash-merged. The PR contains only the four-commit Kimi-K3 delta and review follow-ups on live-main base `04b7e3f6f2d9de5d5741aafb8fa6375a18eee693`. This preserves #611 graph-driven ORT GenAI decoder configs, #618 ORT GenAI end-to-end CI, #624 generic decoder config migration, and the Kimi Linear, MiniMax, runtime, and fail-closed quantization changes already on main. `git range-diff` reports all four replayed commits as patch-identical (`=`). ## Validation - Kimi-K3 model and GGUF tests: 30 passed - focused builder/build-graph Kimi-K3 checks: 4 passed, 1900 deselected - affected generic ORT config/runtime/E2E tests: 250 passed - affected OGA metadata tests: 94 passed, 1 skipped - broad non-integration suite: 7765 passed, 56 skipped, 1 subtest passed - repository lintrunner passed - final GPT-5.6 Sol medium review reported no findings ## Real-checkpoint note The pinned `yujiepan/kimi-k3-tiny-random@a5c86ee03f07f7b141508b0108304a1447fbb345` checkpoint was assessed, but macOS cannot run its required `fla-core` kernels and its selective compressed-tensors MXFP4 expert representation is unsupported by the generic HF loader. That format fails explicitly rather than producing an incorrectly quantized graph. --------- Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 25, 2026
## Summary - clean up still-valid typing, documentation, test robustness, error-message, maintainability, and performance findings left by the GGUF PR stack - preserve the behavioral fixes merged in #625-#632 - hash reused multi-GB GGUF sources once, at the final pre-publication integrity gate, while retaining cheap identity checks around staging Exact base: `2db9d33debdc254d879a51b14434c9a81c230f4f` Exact head: `4541ed2bc9d2ab4227484510b4a85b2d9113eb25` ## Reconstructed original 17-item low-priority tranche The persisted audit retained only the totals, so this list was reconstructed from the live threads and current source. All but the already-fixed #596 comment are addressed in this PR. | PR | Comment | Disposition | |---|---:|---| | #550 | 3837144656 | Implemented: correct tuple return annotation | | #552 | 3837389183 | Implemented: multichannel waveform shape docs | | #559 | 3837716261 | Implemented: name-based cache assertions | | #573 | 3854308350 | Implemented: one final GGUF hash, with integrity regression coverage | | #574 | 3854345521 | Implemented: `TensorRole \| None` typing | | #574 | 3854345597 | Implemented: removed obsolete verdict filtering | | #577 | 3854472859 | Implemented: documented SSM sequence length | | #578 | 3843112613 | Implemented: documented F64 passthrough | | #579 | 3854518332 | Implemented: generalized fused-projection error | | #580 | 3854576300 | Implemented: removed brittle node counts | | #583 | 3854681386 | Implemented: documented conditional draft outputs | | #587 | 3854816872 | Implemented: corrected MTP output contract docs | | #596 | 3846110059 | Already fixed on base: unambiguous GQA bias comment | | #600 | 3855174281 | Implemented: metadata-count-only MTP error | | #607 | 3855776048 | Implemented: stable route-field assertions | | #607 | 3855776086 | Implemented: public tensor iterator | | #609 | 3856486136 | Implemented: fail-closed LM-head comment | ## Current unresolved-thread disposition This covers all 48 Copilot threads returned by the reproducible #600-#630 query. The one human #623 thread is excluded. | PR | Comment | Current-main disposition and evidence | |---|---:|---| | #600 | 3855174177 | Already fixed by #629: package cycle and reserved-sidecar validation | | #600 | 3855174238 | Already fixed by #629: explicit MTP sidecar naming/loading | | #600 | 3855174281 | Implemented here: error no longer invents an observed block count | | #602 | 3848155961 | Outside exact stack; already fixed: top-level `expert_dtype` is classified before early return | | #602 | 3848155983 | Outside exact stack; still-valid behavioral block-quant validation, unchanged | | #602 | 3848156000 | Outside exact stack; still-valid truncated-read behavioral finding, unchanged | | #602 | 3848156022 | Outside exact stack; still-valid descriptor byte/dtype validation, unchanged | | #602 | 3848156040 | Outside exact stack; still-valid expert-bank payload validation, unchanged | | #603 | 3855249765 | Already fixed by #630: runtime preflight preserves shard sets | | #603 | 3855249840 | Already fixed by #630: success output follows durable runtime publication | | #604 | 3855343082 | Implemented here: graph-only MTP persistence distinguished from runtime rejection | | #604 | 3855343131 | Already fixed by #630: runtime success messages are atomic | | #607 | 3855776001 | Implemented here: missing generation golden skips before provenance read | | #607 | 3855776048 | Implemented here: only stable route fields are asserted | | #607 | 3855776086 | Implemented here: tensor count uses `tensor_items_raw()` | | #608 | 3856079371 | Still-valid behavioral cache-symlink containment finding; unchanged | | #608 | 3856079415 | Still-valid behavioral lowercase-digest validation finding; unchanged | | #609 | 3856486136 | Implemented here: comment matches value-preserving policy | | #610 | 3855541683 | Already fixed by #628: Falcon bias precedence is explicit | | #610 | 3855541761 | Already fixed by #628: CTRL tiny config exercises projection biases | | #611 | 3856595840 | Implemented here: runtime test resolves the distribution providing the module | | #612 | 3855677383 | Already fixed by #625: supported-version endianness detection | | #612 | 3855677427 | Implemented here: shared `INT64_MAX` sentinel | | #612 | 3855677460 | Implemented here: shared PLaMo2 width inference | | #612 | 3855677486 | Implemented here: accepted PLaMo2 activation spellings are explicit | | #613 | 3855845678 | Implemented here: canonical issue URL | | #613 | 3855845757 | Implemented here: Mamba-1 function-registration docs | | #613 | 3855845806 | Implemented here: test expects the canonical issue URL | | #614 | 3855988545 | Implemented here: removed stale Nemotron-H divergence comments | | #614 | 3855988597 | Already fixed by #628: zero-head geometry raises actionable `ValueError` | | #615 | 3856082290 | Already fixed by #626: dense GraniteHybrid bias closure | | #618 | 3856729330 | Implemented here: required routes filter ORT GenAI evidence | | #618 | 3856729409 | Implemented here: env-selected runtime version is authoritative | | #618 | 3856729490 | Stale/N/A: PR-description-only matrix claim; repository workflow claims one pinned version | | #618 | 3856729563 | Implemented here: schema tail restored to normal indentation | | #619 | 3856342484 | Already fixed on base: Kimi Linear uses `/issues/605` | | #619 | 3856342532 | Already fixed by #628: config rejects convolution kernels below 2 | | #619 | 3856342580 | Already fixed by #628: GGUF contract rejects convolution kernels below 2 | | #620 | 3855717041 | Already fixed by #627: tied LM-head-only checkpoints are retained | | #621 | 3856722324 | Already fixed by #628: Kimi-K3 required metadata is complete | | #623 | 3855931210 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939302 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939358 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939394 | N/A to current main: comment belongs to open, unmerged #623 | | #624 | 3856777152 | Implemented here: runtime compatibility reuses the emitted model type | | #625 | 3857049733 | Implemented here: unsupported header reports both endian candidates | | #629 | 3857313182 | Newer post-audit behavioral sidecar-symlink cleanup finding; unchanged | | #629 | 3857313251 | Newer post-audit cross-platform path-safety finding; unchanged | ## Validation - affected GGUF/package/ORT GenAI/model/schema tests: 1,040 passed - broad non-integration suite: 7,851 passed, 56 skipped, 1 subtest passed - generated GGUF docs checks: 7 passed - initialized `lintrunner`; full lint/format passed - GPT-5.6 Sol medium review: one integrity finding fixed; re-review found no significant issues Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
onnxruntime-genai0.14.1 and 0.15.2 releases--runtime ort-genaipackage pathLocal validation
lintrunner f --output oneline --all-files && lintrunner -a: cleanRuntime scope
Required CI is CPU-only. CUDA remains scheduled/manual-only because the existing GPU infrastructure uses a CUDA-12-specific prerelease ORT feed and does not provide a stable pinned released OGA lane; this PR makes no CUDA EP claim.
Stack
Based exactly on
justinchuby-generic-genai-configs/ #611.