Skip to content

Add ORT GenAI end-to-end CI - #618

Merged
justinchuby merged 2 commits into
mainfrom
justinchuby-ort-genai-e2e-ci
Aug 25, 2026
Merged

Add ORT GenAI end-to-end CI#618
justinchuby merged 2 commits into
mainfrom
justinchuby-ort-genai-e2e-ci

Conversation

@justinchuby

Copy link
Copy Markdown
Member

Summary

  • add a network-free CPU E2E matrix against pinned onnxruntime-genai 0.14.1 and 0.15.2 releases
  • exercise tokenizer loading, generic decoder model creation, prefill, cache-dependent multi-step decode, deterministic output, reload, and malformed-config rejection
  • add a scheduled/manual real SmolLM-135M F16 CPU lane using immutable GGUF/tokenizer revisions, verified sizes/hashes, isolated HF/Xet caches, and the production --runtime ort-genai package path
  • enroll runtime routes through schema-validated YAML evidence and expand affected-model selection across graph, task, tokenizer, qtype, runtime helper, workflow, and dependency surfaces
  • document tiering, exact commands, CPU scope, released capability differences, and the CUDA waiver

Local validation

  • OGA 0.14.1 fast lane: 2 passed
  • OGA 0.15.2 fast lane: 2 passed
  • OGA 0.15.2 pinned SmolLM F16 real lane: 1 passed
  • schema/evidence/affected detection: 340 passed
  • broad non-integration suite: 7,106 passed, 52 skipped, 1 subtest passed
  • lintrunner f --output oneline --all-files && lintrunner -a: clean
  • two independent security/reliability reviews plus a final post-fix review completed; all actionable findings resolved

Runtime scope

Required CI is CPU-only. CUDA remains scheduled/manual-only because the existing GPU infrastructure uses a CUDA-12-specific prerelease ORT feed and does not provide a stable pinned released OGA lane; this PR makes no CUDA EP claim.

Stack

Based exactly on justinchuby-generic-genai-configs / #611.

Comment thread tests/ort_genai_e2e_test.py Fixed
Comment thread tests/ort_genai_e2e_test.py Fixed
Comment thread tests/ort_genai_e2e_test.py Fixed
@github-actions

github-actions Bot commented Aug 25, 2026

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing 063eb8c11759e9

Model Sub-model Changes Status

No architecture changes detected.


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

@github-actions

github-actions Bot commented Aug 25, 2026

Copy link
Copy Markdown

Performance Comparison

Comparing 063eb8c11759e9

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0%
bert (feature-extraction) num_nodes 68 68 +0.0%
falcon model_size_bytes 364 KB 364 KB +0.0%
falcon num_nodes 66 66 +0.0%
gemma2 model_size_bytes 428 KB 428 KB +0.0%
gemma2 num_nodes 105 105 +0.0%
gpt2 model_size_bytes 388 KB 388 KB +0.0%
gpt2 num_nodes 54 54 +0.0%
llama model_size_bytes 425 KB 425 KB +0.0%
llama num_nodes 60 60 +0.0%
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0%
llama (static-cache) num_nodes 56 56 +0.0%
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0%
mamba (ssm-text-generation) num_nodes 94 94 +0.0%
phi3 model_size_bytes 421 KB 421 KB +0.0%
phi3 num_nodes 58 58 +0.0%
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0%
phi3 (static-cache) num_nodes 54 54 +0.0%
qwen2 model_size_bytes 425 KB 425 KB +0.0%
qwen2 num_nodes 60 60 +0.0%
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0%
qwen2 (static-cache) num_nodes 56 56 +0.0%
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0%
qwen3_5_moe (hybrid-text-generation) num_nodes 265 265 +0.0%
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0%
qwen3_5_text (hybrid-text-generation) num_nodes 127 127 +0.0%
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0%
qwen3_5_vl (hybrid-qwen-vl) num_nodes 429 429 +0.0%
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0%
t5 (seq2seq) num_nodes 176 176 +0.0%
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0%
whisper (speech-to-text) num_nodes 128 128 +0.0%

No performance regressions.

@justinchuby
justinchuby force-pushed the justinchuby-generic-genai-configs branch from c684c1e to f43a31f Compare August 25, 2026 19:32
Base automatically changed from justinchuby-generic-genai-configs to main August 25, 2026 19:33
@justinchuby
justinchuby requested a review from a team August 25, 2026 19:33
justinchuby and others added 2 commits August 25, 2026 12:34
Add a pinned, network-free released-runtime matrix and a scheduled real SmolLM F16 generation lane. Drive real-model enrollment from schema-validated evidence metadata and make relevant graph, tokenizer, runtime, and workflow changes select the gate automatically.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Remove the legacy dual-version matrix and validate only the latest stable 0.15.2 release. Refresh the SmolLM route and runtime package evidence after the #611 squash merge, and keep the real artifact lane limited to scheduled, manual, or explicitly requested runs.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the justinchuby-ort-genai-e2e-ci branch from 7e30dcf to 11759e9 Compare August 25, 2026 19:46
Copilot AI lite review requested due to automatic review settings August 25, 2026 19:46
@justinchuby
justinchuby merged commit e591ecf into main Aug 25, 2026
8 checks passed
@justinchuby
justinchuby deleted the justinchuby-ort-genai-e2e-ci branch August 25, 2026 19:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR introduces a CPU-only ORT GenAI end-to-end CI lane for Mobius, adding both a fast synthetic “network-free” decoder package test and a pinned real-artifact SmolLM GGUF route test, plus schema/evidence plumbing and “affected models” detection to ensure these lanes run when relevant runtime surfaces change.

Changes:

  • Add ORT GenAI E2E test coverage (synthetic decoder + pinned SmolLM real-artifact generation) and a dedicated GitHub Actions workflow to run it.
  • Extend golden-case YAML schema and GGUF runtime evidence/registry to record ORT GenAI enrollment + released capability expectations.
  • Expand detect_affected_models.py “shared_infra” classification to conservatively select the new runtime lanes when core runtime/packaging/dependency surfaces change.

Reviewed changes

Copilot reviewed 18 out of 18 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
tests/yaml_schema_test.py Adds a coverage gate ensuring runtime-supported routes have ORT GenAI E2E enrollment markers.
tests/ort_genai_e2e_test.py New synthetic, network-free ORT GenAI decoder E2E tests (load/tokenize/prefill/decode/reload + malformed-config rejection).
tests/gguf_small_model_runtime_integration_test.py Adds isolated HF cache fixture and upgrades SmolLM ORT GenAI path to the production --runtime ort-genai route with artifact verification.
testdata/cases/schema.json Adds ort_genai enrollment schema plus $defs capability constraints for released runtime expectations.
testdata/cases/causal-lm/smollm-135m-gguf-f16.yaml Enrolls SmolLM F16 GGUF route into the ORT GenAI real lane with evidence ID, byte budget, and capabilities.
src/mobius/integrations/gguf/_runtime_evidence.py Adds ORT GenAI-specific runtime evidence record for the pinned SmolLM route.
src/mobius/integrations/gguf/_docs_test.py Updates docs tests to expect the new ORT GenAI evidence ID alongside existing evidence.
src/mobius/integrations/gguf/_arch_registry.py Extends runtime evidence IDs for llama/SmolLM to include ORT GenAI evidence.
src/mobius/_testing/golden.py Plumbs ort_genai YAML field into loaded golden test cases.
scripts/detect_affected_models.py Introduces explicit shared-infra patterns/prefixes to trigger runtime matrix runs for ORT GenAI and related surfaces.
scripts/detect_affected_models_test.py Updates classification expectations and adds coverage for new shared-infra triggers.
pyproject.toml Registers new pytest markers ort_genai_fast and ort_genai_real.
docs/design/multi-tier-testing-strategy.md Documents the ORT GenAI fast/real lanes and local commands, plus the CUDA waiver rationale.
docs/api/build_from_gguf.md Updates llama runtime support narrative to reference evidenced ONNX Runtime/ORT GenAI versions.
CONTRIBUTING.md Updates contributor checklist to point to the new ORT GenAI fast/real lanes.
.github/workflows/ort_genai_e2e.yml Adds a dedicated ORT GenAI E2E workflow (fast lane + scheduled/manual real lane).
.github/workflows/main.yml Wires the ORT GenAI E2E workflow into main CI when detect-affected indicates relevant changes.
.agents/skills/quality-checklist/SKILL.md Updates quality checklist references to the new ORT GenAI lane commands.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/yaml_schema_test.py
Comment on lines +130 to +142
required_routes = {
(
evidence.architecture,
evidence.repository,
evidence.revision,
evidence.filename,
evidence.import_route,
)
for spec in iter_arch_specs()
if spec.runtime is Support.SUPPORTED
for evidence_id in spec.runtime_evidence_ids
if (evidence := runtime_evidence(evidence_id)) is not None
}
Comment on lines +207 to +211
expected_version = os.environ.get("MOBIUS_EXPECTED_ORT_GENAI_VERSION")
installed_version = version("onnxruntime-genai")
if expected_version:
assert installed_version == expected_version
assert installed_version == _EXPECTED_ORT_GENAI_VERSION
Comment on lines +23 to +47
fast:
name: Fast CPU / OGA 0.15.2
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
persist-credentials: false
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: "3.12"
cache: pip
cache-dependency-path: pyproject.toml
- name: Install pinned runtime and test dependencies
run: |
python -m pip install --disable-pip-version-check \
"onnxruntime-genai==0.15.2" \
"pytest==8.4.2" \
"pytest-timeout==2.4.0"
python -m pip install --disable-pip-version-check -e .
- name: Run network-free ORT GenAI E2E
env:
HF_HUB_OFFLINE: "1"
TRANSFORMERS_OFFLINE: "1"
MOBIUS_EXPECTED_ORT_GENAI_VERSION: "0.15.2"
Comment on lines +513 to 536
"architecture": {
"type": "string",
"description": "Optional registry architecture key (e.g. Qwen35MtpModel) forcing a specific module class + task at build time. Needed for auxiliary heads (DFlash, MTP) that share a base checkpoint whose architectures field would otherwise auto-route to the base model."
}
},
"$defs": {
"ortGenaiCapabilities": {
"type": "object",
"additionalProperties": false,
"required": [
"generic_decoder",
"state_groups"
],
"properties": {
"generic_decoder": {
"const": true
},
"state_groups": {
"const": false
}
}
}
}
}
justinchuby added a commit that referenced this pull request Aug 25, 2026
## Summary

- normalize every graph-representable decoder-only ORT GenAI config
emitted by Mobius to `model.type: "decoder"`
- preserve specialized types only for GPT-2's unrepresentable cache ABI,
LFM2's convolution cache, Phi-3-family LongRoPE,
multimodal/encoder-decoder pipelines, and auxiliary sidecar topologies
- update the Gemma4 decoder-only example, CLI docs, and ORT GenAI config
skill; keep #618's OGA 0.15.2 CI and evidence unchanged

## Validation

- `235 passed` focused ORT GenAI config tests
- `7717 passed, 56 skipped, 1 subtest passed` broad non-integration
suite
- `37 passed` GGUF runtime tests
- `2 passed` network-free OGA 0.15.2 fast E2E
- `1 passed` pinned real SmolLM F16 CPU OGA 0.15.2 E2E
- Sphinx documentation build
- full lintrunner formatting/lint
- two GPT-5.6 Sol medium reviews; final review reported no findings

Base: `e591ecfaae00d312cb64c4334b61aff0a9e87f05`

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
## Summary

- add a dedicated Kimi-K3 config, model, and heterogeneous-state task
with exact KDA/NoPE gated-MLA scheduling, AttnRes mixing, SiTU latent
MoE, shared experts, and an untied output head
- import pinned llama.cpp `kimi-k3` GGUF metadata and tensors with
strict metadata/tensor/shape/storage closure and malformed-input
rejection
- preserve lossless quantization routes under #609: separate rank-3 MLA
projections fail closed unless explicitly dequantized, while fused Q4_0
KV-B is split by exact packed-row reordering across weights, scales, and
zero-points
- preserve valid KDA convolution history across padding, validate
kernel/state/task contracts, and cover replay, reorder, float/quantized
import, roundtrip, CLI, and negative paths
- keep generic OGA runtime packaging truthfully deferred under #605
because released cache schemas cannot represent the heterogeneous state
ABI

## Reconstruction

Reconstructed after #619 was admin squash-merged. The PR contains only
the four-commit Kimi-K3 delta and review follow-ups on live-main base
`04b7e3f6f2d9de5d5741aafb8fa6375a18eee693`. This preserves #611
graph-driven ORT GenAI decoder configs, #618 ORT GenAI end-to-end CI,
#624 generic decoder config migration, and the Kimi Linear, MiniMax,
runtime, and fail-closed quantization changes already on main. `git
range-diff` reports all four replayed commits as patch-identical (`=`).

## Validation

- Kimi-K3 model and GGUF tests: 30 passed
- focused builder/build-graph Kimi-K3 checks: 4 passed, 1900 deselected
- affected generic ORT config/runtime/E2E tests: 250 passed
- affected OGA metadata tests: 94 passed, 1 skipped
- broad non-integration suite: 7765 passed, 56 skipped, 1 subtest passed
- repository lintrunner passed
- final GPT-5.6 Sol medium review reported no findings

## Real-checkpoint note

The pinned
`yujiepan/kimi-k3-tiny-random@a5c86ee03f07f7b141508b0108304a1447fbb345`
checkpoint was assessed, but macOS cannot run its required `fla-core`
kernels and its selective compressed-tensors MXFP4 expert representation
is unsupported by the generic HF loader. That format fails explicitly
rather than producing an incorrectly quantized graph.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
## Summary

- clean up still-valid typing, documentation, test robustness,
error-message,
  maintainability, and performance findings left by the GGUF PR stack
- preserve the behavioral fixes merged in #625-#632
- hash reused multi-GB GGUF sources once, at the final pre-publication
integrity
  gate, while retaining cheap identity checks around staging

Exact base: `2db9d33debdc254d879a51b14434c9a81c230f4f`

Exact head: `4541ed2bc9d2ab4227484510b4a85b2d9113eb25`

## Reconstructed original 17-item low-priority tranche

The persisted audit retained only the totals, so this list was
reconstructed
from the live threads and current source. All but the already-fixed #596
comment are addressed in this PR.

| PR | Comment | Disposition |
|---|---:|---|
| #550 | 3837144656 | Implemented: correct tuple return annotation |
| #552 | 3837389183 | Implemented: multichannel waveform shape docs |
| #559 | 3837716261 | Implemented: name-based cache assertions |
| #573 | 3854308350 | Implemented: one final GGUF hash, with integrity
regression coverage |
| #574 | 3854345521 | Implemented: `TensorRole \| None` typing |
| #574 | 3854345597 | Implemented: removed obsolete verdict filtering |
| #577 | 3854472859 | Implemented: documented SSM sequence length |
| #578 | 3843112613 | Implemented: documented F64 passthrough |
| #579 | 3854518332 | Implemented: generalized fused-projection error |
| #580 | 3854576300 | Implemented: removed brittle node counts |
| #583 | 3854681386 | Implemented: documented conditional draft outputs
|
| #587 | 3854816872 | Implemented: corrected MTP output contract docs |
| #596 | 3846110059 | Already fixed on base: unambiguous GQA bias
comment |
| #600 | 3855174281 | Implemented: metadata-count-only MTP error |
| #607 | 3855776048 | Implemented: stable route-field assertions |
| #607 | 3855776086 | Implemented: public tensor iterator |
| #609 | 3856486136 | Implemented: fail-closed LM-head comment |

## Current unresolved-thread disposition

This covers all 48 Copilot threads returned by the reproducible
#600-#630
query. The one human #623 thread is excluded.

| PR | Comment | Current-main disposition and evidence |
|---|---:|---|
| #600 | 3855174177 | Already fixed by #629: package cycle and
reserved-sidecar validation |
| #600 | 3855174238 | Already fixed by #629: explicit MTP sidecar
naming/loading |
| #600 | 3855174281 | Implemented here: error no longer invents an
observed block count |
| #602 | 3848155961 | Outside exact stack; already fixed: top-level
`expert_dtype` is classified before early return |
| #602 | 3848155983 | Outside exact stack; still-valid behavioral
block-quant validation, unchanged |
| #602 | 3848156000 | Outside exact stack; still-valid truncated-read
behavioral finding, unchanged |
| #602 | 3848156022 | Outside exact stack; still-valid descriptor
byte/dtype validation, unchanged |
| #602 | 3848156040 | Outside exact stack; still-valid expert-bank
payload validation, unchanged |
| #603 | 3855249765 | Already fixed by #630: runtime preflight preserves
shard sets |
| #603 | 3855249840 | Already fixed by #630: success output follows
durable runtime publication |
| #604 | 3855343082 | Implemented here: graph-only MTP persistence
distinguished from runtime rejection |
| #604 | 3855343131 | Already fixed by #630: runtime success messages
are atomic |
| #607 | 3855776001 | Implemented here: missing generation golden skips
before provenance read |
| #607 | 3855776048 | Implemented here: only stable route fields are
asserted |
| #607 | 3855776086 | Implemented here: tensor count uses
`tensor_items_raw()` |
| #608 | 3856079371 | Still-valid behavioral cache-symlink containment
finding; unchanged |
| #608 | 3856079415 | Still-valid behavioral lowercase-digest validation
finding; unchanged |
| #609 | 3856486136 | Implemented here: comment matches value-preserving
policy |
| #610 | 3855541683 | Already fixed by #628: Falcon bias precedence is
explicit |
| #610 | 3855541761 | Already fixed by #628: CTRL tiny config exercises
projection biases |
| #611 | 3856595840 | Implemented here: runtime test resolves the
distribution providing the module |
| #612 | 3855677383 | Already fixed by #625: supported-version
endianness detection |
| #612 | 3855677427 | Implemented here: shared `INT64_MAX` sentinel |
| #612 | 3855677460 | Implemented here: shared PLaMo2 width inference |
| #612 | 3855677486 | Implemented here: accepted PLaMo2 activation
spellings are explicit |
| #613 | 3855845678 | Implemented here: canonical issue URL |
| #613 | 3855845757 | Implemented here: Mamba-1 function-registration
docs |
| #613 | 3855845806 | Implemented here: test expects the canonical issue
URL |
| #614 | 3855988545 | Implemented here: removed stale Nemotron-H
divergence comments |
| #614 | 3855988597 | Already fixed by #628: zero-head geometry raises
actionable `ValueError` |
| #615 | 3856082290 | Already fixed by #626: dense GraniteHybrid bias
closure |
| #618 | 3856729330 | Implemented here: required routes filter ORT GenAI
evidence |
| #618 | 3856729409 | Implemented here: env-selected runtime version is
authoritative |
| #618 | 3856729490 | Stale/N/A: PR-description-only matrix claim;
repository workflow claims one pinned version |
| #618 | 3856729563 | Implemented here: schema tail restored to normal
indentation |
| #619 | 3856342484 | Already fixed on base: Kimi Linear uses
`/issues/605` |
| #619 | 3856342532 | Already fixed by #628: config rejects convolution
kernels below 2 |
| #619 | 3856342580 | Already fixed by #628: GGUF contract rejects
convolution kernels below 2 |
| #620 | 3855717041 | Already fixed by #627: tied LM-head-only
checkpoints are retained |
| #621 | 3856722324 | Already fixed by #628: Kimi-K3 required metadata
is complete |
| #623 | 3855931210 | N/A to current main: comment belongs to open,
unmerged #623 |
| #623 | 3855939302 | N/A to current main: comment belongs to open,
unmerged #623 |
| #623 | 3855939358 | N/A to current main: comment belongs to open,
unmerged #623 |
| #623 | 3855939394 | N/A to current main: comment belongs to open,
unmerged #623 |
| #624 | 3856777152 | Implemented here: runtime compatibility reuses the
emitted model type |
| #625 | 3857049733 | Implemented here: unsupported header reports both
endian candidates |
| #629 | 3857313182 | Newer post-audit behavioral sidecar-symlink
cleanup finding; unchanged |
| #629 | 3857313251 | Newer post-audit cross-platform path-safety
finding; unchanged |

## Validation

- affected GGUF/package/ORT GenAI/model/schema tests: 1,040 passed
- broad non-integration suite: 7,851 passed, 56 skipped, 1 subtest
passed
- generated GGUF docs checks: 7 passed
- initialized `lintrunner`; full lint/format passed
- GPT-5.6 Sol medium review: one integrity finding fixed; re-review
found no significant issues

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants