Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
152 commits
Select commit Hold shift + click to select a range
83c2938
Simplify composite inference metadata
justinchuby Aug 12, 2026
7f58445
Add reusable ONNX generation policy components
justinchuby Aug 12, 2026
ffc1b04
Align policy components with workflow contract
justinchuby Aug 12, 2026
d2e0810
Migrate masked diffusion metadata to SSA workflow
justinchuby Aug 12, 2026
6d265cc
Migrate codec metadata to typed workflow SSA
justinchuby Aug 12, 2026
66ab136
Resolve combined workflow migration tests
justinchuby Aug 12, 2026
5b6a3ff
Fix workflow sampling and loop correctness
justinchuby Aug 12, 2026
144dcfa
Make generation workflows request executable
justinchuby Aug 12, 2026
99820aa
Migrate diffusion metadata to workflow SSA
justinchuby Aug 12, 2026
ba813b9
Migrate multimodal metadata to workflow SSA
justinchuby Aug 12, 2026
62114b1
Add nested TTS workflow generation
justinchuby Aug 12, 2026
3a40b0b
Add speculative workflow branch joins
justinchuby Aug 12, 2026
56e76f5
Validate migrated workflows and Muse metadata
justinchuby Aug 12, 2026
69153d5
Fix workflow policy artifact and solver execution
justinchuby Aug 12, 2026
b69bbfa
Carry normalized decoder logits through workflows
justinchuby Aug 12, 2026
8a77aee
Add speculative accepted-state rollback policies
justinchuby Aug 12, 2026
13b9954
Migrate Qwen3 TTS to real nested workflow
justinchuby Aug 12, 2026
3de8157
Emit typed SSA image preprocessing programs
justinchuby Aug 12, 2026
c6fa8fd
Complete speculative prefix rollback workflow
justinchuby Aug 12, 2026
7fdfe86
Harden workflow export contracts
justinchuby Aug 12, 2026
e464ed9
Add grammar and adaptive proposal policy components
justinchuby Aug 12, 2026
91bc6c2
Migrate workflows to versioned guidance contracts
justinchuby Aug 13, 2026
7279dff
Simplify public workflow metadata
justinchuby Aug 13, 2026
ceabcae
Parameterize workflow token sampling
justinchuby Aug 13, 2026
61e3fa7
Add capture-friendly min-p sampling
justinchuby Aug 13, 2026
50a2144
Gate workflow performance against native
justinchuby Aug 13, 2026
045f20a
Make policy contracts data driven
justinchuby Aug 13, 2026
a201acd
Use lexical workflow loop iteration
justinchuby Aug 13, 2026
d880482
Align workflows with ONNX GenAI serving contracts
justinchuby Aug 13, 2026
7c6a201
Add KV service contracts to nested TTS
justinchuby Aug 13, 2026
5169a15
Check in ONNX GenAI workflow fixtures
justinchuby Aug 13, 2026
1b66dcc
Bound speculative grammar forced-token emits
justinchuby Aug 13, 2026
f20885b
Execute every ONNX GenAI workflow fixture
justinchuby Aug 13, 2026
7420d76
Stabilize TTS fixtures across serializer versions
justinchuby Aug 13, 2026
b7f9936
Fix component symbolic shape isolation
justinchuby Aug 13, 2026
532b83d
Support reduced-precision decoder logits policies
justinchuby Aug 13, 2026
37bba25
Add paired real Muse H200 benchmark harness
justinchuby Aug 13, 2026
b4ab3bc
Guard metadata against tiny config leakage
justinchuby Aug 13, 2026
632ab91
Derive KV storage mode from model ABI
justinchuby Aug 13, 2026
7f017bc
Emit capture-stable decoder workflow state
justinchuby Aug 13, 2026
dc3f362
Add executable text-only VLM branch
justinchuby Aug 13, 2026
afa6541
Refresh capture-stable workflow fixtures
justinchuby Aug 13, 2026
9d781d8
Fix shared-buffer cache recurrence metadata
justinchuby Aug 13, 2026
e93d68e
Make VLM media presence explicit
justinchuby Aug 13, 2026
3934de7
Derive generation constants from package metadata
justinchuby Aug 13, 2026
32d1c46
Mark autoregressive EOS loop termination
justinchuby Aug 13, 2026
f22c9df
Pin exact Muse benchmark prompt tokens
justinchuby Aug 13, 2026
b30e4db
Verify native Muse prompt identity
justinchuby Aug 13, 2026
8c936d6
Pin paired Muse attention kernel
justinchuby Aug 13, 2026
30cc0fe
Validate current workflow termination contract
justinchuby Aug 13, 2026
6d51fa4
Derive dynamic KV storage from decoder ABI
justinchuby Aug 13, 2026
b6ba874
Compare workflow fixtures semantically
justinchuby Aug 13, 2026
abe1797
Normalize ONNX SSA names in fixture checks
justinchuby Aug 13, 2026
53cce3b
Ignore serializer value wiring in fixture checks
justinchuby Aug 13, 2026
6dd733c
Limit fixture checks to stable ONNX semantics
justinchuby Aug 13, 2026
09071ed
Compare canonical ONNX tensor contents
justinchuby Aug 13, 2026
30454a3
Pin executable Muse workflow profiler
justinchuby Aug 13, 2026
92099e8
Restore logical shared-buffer KV contracts
justinchuby Aug 13, 2026
87cf626
Clarify runtime shared KV binding contract
justinchuby Aug 13, 2026
8932134
Separate Muse prefill and decode mask shapes
justinchuby Aug 13, 2026
1436b49
Use one persistent Muse attention mask
justinchuby Aug 13, 2026
448952f
Enforce fixed-count Muse benchmark semantics
justinchuby Aug 13, 2026
16dff43
Verify paired benchmark runtime identity
justinchuby Aug 13, 2026
b4108e5
Record paired benchmark runner identity
justinchuby Aug 13, 2026
49dd6c6
Pin exact prompt-ID runtime head
justinchuby Aug 13, 2026
9f58c2b
Format persistent-mask workflow changes
justinchuby Aug 13, 2026
cf14bd8
Run workflow conformance on ORT 1.28
justinchuby Aug 13, 2026
1b81572
Pin image-enabled workflow runtime
justinchuby Aug 13, 2026
2c0580a
Require CUDA Graph in paired workflow runs
justinchuby Aug 13, 2026
8a9aa99
Refresh persistent-mask workflow fixtures
justinchuby Aug 14, 2026
daedfed
Manifest dirty benchmark runtimes
justinchuby Aug 14, 2026
a6ac2d7
Record clean Muse release diagnostics
justinchuby Aug 14, 2026
50129f2
Pin measured Muse producer artifacts
justinchuby Aug 14, 2026
80a5ae8
Record SHA-named Muse release rerun
justinchuby Aug 14, 2026
1e103e2
Record exact prompt-ID release evidence
justinchuby Aug 14, 2026
6c22eb4
Pin landed CUDA bridge runtime
justinchuby Aug 14, 2026
7f66a6c
Pin dirty bridge evidence manifest
justinchuby Aug 14, 2026
3fdbb18
Enforce 0.99x Muse release gate
justinchuby Aug 14, 2026
130e393
Add heterogeneous batched generation policies
justinchuby Aug 14, 2026
c3028e0
Upgrade batched generation policy contracts to v2
justinchuby Aug 14, 2026
c34cde4
Bind row-wise emits to semantic request identities
justinchuby Aug 14, 2026
8784ad4
Validate workflows against semantic row runtime
justinchuby Aug 14, 2026
509c049
Expose row-output conformance errors
justinchuby Aug 14, 2026
bdc10f0
Exercise explicit batched policy controls
justinchuby Aug 14, 2026
f597878
Bind row emits to carried serving slots
justinchuby Aug 14, 2026
6b9bb7b
Match the exact batched policy v2 ABI
justinchuby Aug 14, 2026
96580bb
Lock exact batched policy bindings
justinchuby Aug 14, 2026
28f8129
Require complete diffusion workflow packages
justinchuby Aug 14, 2026
faf8686
Assert batched policy super-island admission
justinchuby Aug 14, 2026
00202cd
Declare coordinated row compaction semantics
justinchuby Aug 14, 2026
8175780
Enable nested TTS row compaction
justinchuby Aug 14, 2026
3dc1f96
Exercise nested TTS batch compaction
justinchuby Aug 14, 2026
7abe416
Add generic runtime adapter artifact model
justinchuby Aug 15, 2026
33c9929
Fix adapter framework lint diagnostics
justinchuby Aug 15, 2026
8f1efca
Align adapter groundwork with native LoRA contracts
justinchuby Aug 15, 2026
20f8df1
Validate adapter target parameter consumption
justinchuby Aug 15, 2026
7fdd437
Emit generic parameter adapter workflow artifacts
justinchuby Aug 15, 2026
d25779a
Fix adapter metadata lint findings
justinchuby Aug 15, 2026
f728255
Resolve remaining adapter lint checks
justinchuby Aug 15, 2026
d6e48f1
Pin finalized adapter runtime contract
justinchuby Aug 15, 2026
7809f21
Enforce final adapter selection contract
justinchuby Aug 15, 2026
5c9841a
Freeze generic adapter wire ABI
justinchuby Aug 15, 2026
08ce7ae
Integrate canonical LoRA adapter metadata ABI
justinchuby Aug 15, 2026
53e1271
Freeze final LoRA adapter producer ABI
justinchuby Aug 15, 2026
32eaca6
Align adapter targets with published ABI
justinchuby Aug 15, 2026
fa868ce
Test heterogeneous PEFT binding overrides
justinchuby Aug 15, 2026
219f77e
Pin executable heterogeneous adapter regression
justinchuby Aug 15, 2026
2a46e92
Express encoder-conditioned decode loops and align serving with the r…
justinchuby Aug 20, 2026
16c93ed
Fix Gemma 4 text-only export and prefill-prefix pruning on the Attent…
justinchuby Aug 20, 2026
e83a1df
Declare Gemma 4's heterogeneous KV cache geometry in workflow metadata
justinchuby Aug 20, 2026
9e1e5b3
Derive KV storage, ship chat templates, and drop vestigial row identity
JustinChu Aug 20, 2026
87accf7
Declare Whisper audio preprocessing in speech-to-text workflow metadata
justinchuby Aug 20, 2026
7616c08
Fix Wav2Vec2 CTC correctness and represent it end-to-end in inference…
justinchuby Aug 20, 2026
d93abcd
Make diffusion component exports faithful to their schedulers
justinchuby Aug 20, 2026
904495f
Produce a diffusion workflow that a runtime can actually execute
justinchuby Aug 20, 2026
6743366
Make the Qwen Image VAE exportable and loadable in bfloat16
justinchuby Aug 20, 2026
c6acf9f
Add flow-matching image-edit policy components
justinchuby Aug 20, 2026
26ba902
Emit executable onnx-genai metadata for image-edit pipelines
justinchuby Aug 20, 2026
228a2d7
Free the CogVideoX denoiser from its baked frame count and resolution
Copilot Aug 20, 2026
d23245e
Add a faithful CogVideoX causal 3D VAE decoder with conv-cache chunking
Copilot Aug 20, 2026
e58d35b
Produce video diffusion workflow metadata with temporal state
Copilot Aug 20, 2026
2dc969c
Shape generation positions for decoders with multi-axis rotary embedd…
justinchuby Aug 20, 2026
6b4511e
Add full-duplex speech-to-speech workflow metadata producer
justinchuby Aug 20, 2026
df4b6f5
Reconcile the E2E metadata producers after the rebase onto main
justinchuby Aug 20, 2026
d5353d3
Refresh the artifacts the integrated producers now emit
justinchuby Aug 20, 2026
8c38757
Execute the integrated workflows against their coordinated runtime
justinchuby Aug 20, 2026
5a37fe4
Describe fixed-capacity and FP8 KV caches instead of refusing to
justinchuby Aug 20, 2026
ac540b5
Format the static cache metadata tests
justinchuby Aug 20, 2026
ed73edb
Pin the runtime that accepts a port ABI beside a workflow
justinchuby Aug 20, 2026
e2f736f
Make the workflow the only place a package describes itself
justinchuby Aug 20, 2026
d8da245
Declare a role for every component a task builds
justinchuby Aug 20, 2026
c7e453a
Pin GenAI validation to the branch head, not an ancestor of it
justinchuby Aug 20, 2026
16fd082
Cover the cache layer annotation where it can actually be wrong
justinchuby Aug 20, 2026
c3dfeca
Pin GenAI validation to the workflow-derived scatter proof
justinchuby Aug 20, 2026
564af03
Re-pin GenAI validation after the upstream branch was rebased
justinchuby Aug 20, 2026
a02bf94
Re-pin GenAI validation after a second upstream rebase
justinchuby Aug 21, 2026
9694953
Stop transcribing an exported graph into its own workflow component
justinchuby Aug 21, 2026
e9bd70f
Pin the ONNX GenAI head rather than an ancestor of it
justinchuby Aug 21, 2026
10cd48f
Re-verify the canonical packages against a stricter package loader
justinchuby Aug 21, 2026
2a7bace
Require a sequence role wherever a fixed-capacity cache is bound
justinchuby Aug 21, 2026
76f158b
Stop committing generated fixture graphs; drop the last model.io prod…
justinchuby Aug 21, 2026
43ef64a
Restore revision threading lost to the rebase onto GLM-ASR and LFM2.5-VL
justinchuby Aug 21, 2026
fe1829d
Repin ONNX GenAI: the pinned commit was orphaned by a rebase
justinchuby Aug 21, 2026
9390138
Repin ONNX GenAI again and stop the regenerated tree from being commi…
justinchuby Aug 21, 2026
894527d
Anchor the ONNX GenAI pin against rebases instead of chasing it
justinchuby Aug 21, 2026
dd9eaaa
Repin ONNX GenAI onto the rebased head, and move the anchor with it
justinchuby Aug 21, 2026
00c748d
Give encoder embeddings their own workflow metadata, and add ESM-2
justinchuby Aug 21, 2026
ab343c3
Align metadata producer with final workflow schema
justinchuby Aug 21, 2026
6007c04
Merge origin/main into metadata producer
justinchuby Aug 21, 2026
6429957
Stop duplicating workflow version and ONNX opsets
justinchuby Aug 21, 2026
2c8ae52
Update canonical workflow metadata design
justinchuby Aug 21, 2026
1219f0b
Pin final metadata design revision
justinchuby Aug 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
62 changes: 62 additions & 0 deletions .github/workflows/main.yml
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,68 @@ concurrency:
cancel-in-progress: true

jobs:
onnx-genai-metadata:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove or made separate

name: ONNX GenAI metadata
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@v7
with:
repository: justinchuby/onnx-genai
# Pinned by SHA so the producer/runtime contract under test is
# reproducible. PR #828 is maintained by merging main rather than
# rebasing, so this commit remains reachable from its branch history.
# When bumping, verify reachability with `git merge-base --is-ancestor
# <sha> <branch-tip>`.
ref: 509cd4e9c4471f4cbc59fe44b47168f0ae128fe3
path: validation/onnx-genai
- uses: actions/setup-python@v7
with:
python-version: "3.12"
- name: Install dependencies
run: |
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements/ci/requirements.txt
pip install onnxruntime==1.28.0
pip install -e '.[testing]'
- name: Generate representative packages
run: |
PYTHONPATH=src python tests/generate_onnx_genai_validation_packages.py \
validation/generated
PYTHONPATH=src python tests/compare_onnx_genai_validation_packages.py \
tests/fixtures/onnx_genai_workflows validation/generated
- name: Validate package semantics
run: |
# Report every invalid package, not just the first: the job runs under
# `bash -e`, so an unguarded failure inside the loop would abort it and
# hide the remaining results.
failed=""
# The generated tree, not the committed one: validation resolves each
# component's artifact from disk, and the graphs are generated rather
# than committed. The step above has already proven the two trees
# carry the same metadata.
for package in validation/generated/*; do
[ -f "$package/inference_metadata.yaml" ] || continue
cargo run --quiet \
--manifest-path validation/onnx-genai/Cargo.toml \
-p onnx-genai-metadata --bin validate_metadata -- "$package" \
|| failed="$failed $package"
done
if [ -n "$failed" ]; then
echo "Invalid packages:$failed"
exit 1
fi
- name: Execute all workflow packages
run: |
cp tests/onnx_genai_workflow_conformance.rs \
validation/onnx-genai/crates/onnx-genai-engine/tests/mobius_workflow_conformance.rs
ORT_LIB="$(python -c \
'import onnxruntime, pathlib; print(next((pathlib.Path(onnxruntime.__file__).parent / "capi").glob("libonnxruntime.so*")))')"
MOBIUS_WORKFLOW_CONFORMANCE_DIR="$PWD/validation/generated" \
ONNX_GENAI_ORT_LIB="$ORT_LIB" \
cargo test --manifest-path validation/onnx-genai/Cargo.toml \
-p onnx-genai-engine --test mobius_workflow_conformance -- --nocapture

lint:
name: Lint
runs-on: ubuntu-latest
Expand Down
13 changes: 13 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -213,6 +213,19 @@ __marimo__/
*.onnx.data
*.gguf

# Generated conformance-package artifacts. The metadata beside them is the
# contract under review and is committed; the graphs and adapter weights are a
# deterministic function of
# tests/generate_onnx_genai_validation_packages.py, so committing them would
# store megabytes of bytes no reviewer reads. CI regenerates and validates
# them; tests get them from the materialized_workflow_packages fixture.
tests/fixtures/onnx_genai_workflows/**/*.safetensors

# Where CI and the test fixture materialize those packages in full. Nothing
# here is reviewable input: it is the regenerated output the validate and
# conformance steps consume.
validation/

# Common test dirs
output/**
cache_dir/**
Expand Down
138 changes: 138 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,144 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Fixed

- An exported graph is no longer transcribed into its own workflow component.
A component backed by a shipped `.onnx` file now declares only `ports.roles`;
the artifact answers every question about which ports exist and what shape
they have, and the runtime resolves them against the live session. The
removed block was a second statement of the same ABI with nothing keeping the
two in agreement — the same defect as `model.io`, one level down. Package
validation and runtime conformance both stay at 11/11 against the pinned ONNX
GenAI branch, and the contract tests now resolve every role, invocation
binding and state pair against the graph itself rather than against the
metadata's agreement with a copy of itself, which is a strictly stronger
check.

Policy graphs keep their contracts, and that boundary was measured rather
than assumed: a workflow value inherits its dtype, rank and request axis from
the port that produced it, so dropping them left row-wise emits untyped and
made 4 of the 11 packages invalid. Those contracts type the workflow's own
dataflow; they do not describe an external interface.

- Deep decoders are now covered where the cache layer annotation actually
matters. Every metadata package built in the test suite had two layers, and
below ten a cell label sorts identically whether ordered lexicographically,
numerically or by insertion — so a `layer` derived from a cell's position
rather than parsed from its port name would have passed every assertion while
transposing caches on any real model. `canonical_workflow_contract_test` now
builds twelve-layer dynamic, static-cache and hybrid decoders and pins that
the declared layer restates the port name, that ordering by it recovers the
buffer lists, and that a hybrid's alternating groups own layers their cells'
positions never equal. A producer that dropped the parse leaves the 101
pre-existing assertions green and fails seven of these.

- Every component a task builds now declares an optimization role. Qwen3-TTS's
four loop-wiring graphs (`code_predictor_prefill`,
`code_predictor_step_embedder`, `code_predictor_indices`, `talker_text_step`)
were absent from `TTSTask.model_roles`, so `build_from_module` fell back to
the `"decoder"` role and offered them the GQA / QKV-packing passes meant for
attention stacks, and `inspect_components` under-reported the package by four
components. They now declare a `"glue"` role: a parameter-free graph that
reads every tensor it uses from a graph input. `arch_validation_test` fails
any task that builds a component it does not declare, and a network-free unit
test pins the same invariant for Qwen3-TTS.

### One canonical serialized representation

#### Changed

- **`pipeline.workflow` is now the only place a package describes its graph
ABI.** No export emits `model.io`, including a bare single-file decoder: that
case is a one-component workflow, not a different kind of document. `model`
keeps package-wide geometry and capabilities and nothing else. Two writable
statements of one fact are a defect whatever they contain — nothing forces
them to agree, and a reader of either never learns the other exists — so a
runtime that wants an optimized single-graph path derives it by lowering the
workflow instead. Verified end to end: the ONNX GenAI runtime executes the
fixed-capacity decode path from the workflow alone, with no `model.io` in the
package.

#### Added

- An ONNX component that ships an artifact declares no port contracts. The
`.onnx` file travels inside the package and is authoritative for which ports
exist and what each one's dtype, rank and shape is, so transcribing that into
YAML would be a second writable statement of one fact — the very thing this
section removes — sitting one level below `model.io` rather than beside it.
The runtime resolves ports against the live session, which catches a name the
graph does not expose instead of agreeing with a stale echo of it. A
producer-synthesized policy graph is the exception and states its contracts,
because a workflow value takes its dtype, rank and request axis from the port
that produced it: those contracts are the dataflow's type annotations, not a
description of an external interface.
- Every ONNX component declares `ports.roles`: what it *does* with a value bound
to a port. An invocation records which SSA value reaches a port, not whether
that port is tokens, a mask or logits. Mobius mints these port names in its own
task builders, so it states the mapping (`input_ids`→`token_ids`,
`inputs_embeds`, `attention_mask`, `position_ids`, `logits`,
`last_hidden_state`→`hidden_states`, `encoder_hidden_states`,
`audio_features`) rather than inferring it. A port outside that vocabulary
carries no role.
- State port aliases declare `role` (`key`/`value`) and `layer`. A layer's key
and value buffers are the same dtype and shape, and a cell's label sorts
lexicographically so `cache_10` precedes `cache_2` — pairing per-layer buffers
positionally would silently transpose two layers' caches. Both fields are
emitted together or not at all, so a recurrent or convolution cache is never
given a fabricated index.
- `IndexedScatter.kv_length_ports` names the port carrying the graph-visible
valid length, beside the existing `write_indices_ports`. The two control
vectors are both rank-1 integers and are therefore indistinguishable by shape;
with both named, the whole fixed-capacity ABI is recoverable from the workflow.
- `tests/canonical_workflow_contract_test.py` pins the invariant. It asks one
set of shape-agnostic questions of dynamic, static-cache, FP8, heterogeneous
and composite packages — and of all 11 checked-in fixtures — so a future
feature cannot grow its own top-level block while every feature-specific test
keeps passing.

### Fixed-capacity (static) KV cache and FP8 KV cache metadata

#### Added

- `--features static-cache` now produces onnx-genai metadata instead of being
refused. The producer publishes the write cursor (`write_indices`), the valid
length (`nonpad_kv_seqlen`), the fixed-capacity buffer contracts, the
per-layer input/output pairs, and an `indexed_scatter` state-service update
discipline naming the cursor, the capacity and the per-component port that
carries it. The buffers are declared as `recurrence: {kind: invariant}` loop
cells and the capacity as a `package.cache_capacity` literal workflow input.
Nothing dispatches on model name; the ports are read from the graph.
- Heterogeneous caches keep their own disciplines. Gemma 4's sliding layers stay
on a growing rank-4 BNSH cache while its full-attention layers use rank-3
fixed-capacity buffers, and only the layers that own a buffer bind ports in a
state group — its KV-shared suffix owns none.
- A `static_cache` package joined the checked-in onnx-genai conformance
fixtures, so the engine exercises the fixed-capacity carry and the write
cursor rather than only the growing-tensor path.

#### Fixed

- Gemma 4's shared-KV fallback pinned a 4-D BNSH shape onto *any* borrowed KV
tensor whose rank was not 4. A static-cache source hands over a fully known
rank-3 `[batch, capacity, kv_hidden]` buffer, so this overwrote a correct
shape with a wrong one — corrupting the declared shape of
`updated_key_cache.N` and defeating the rank-3 static-source test further
down, which would then have transposed a rank-3 tensor as BNSH. The fallback
now only supplies a shape when there is none.
- Gemma 4's vision-language decoder dropped `attention_mask` whenever the export
was static, but a Gemma 4 decoder is only *partly* static: its sliding layers
keep a dynamic cache and build their bias from that mask. The hybrid decoder
therefore lost all padding information. Both builders now apply one rule — a
mask exists exactly when some layer still has a dynamic cache — so a fully
static decoder carries no unused port and a hybrid one keeps its mask.
- `--features fp8-kv-cache` no longer silently produces a float16 cache. The
gate only tested whether GQA fusion was *expected*; the pass now reports how
many caches it converted and the build fails when the answer is zero, naming
the reason: FP8 KV storage needs an attention operator with `k_scale`/
`v_scale` inputs, which a `TensorScatter` + `ai.onnx` `Attention` static-cache
graph does not have.


### Qwen3.5/3.6-MoE mixed float/quantized decoder (Olive checkpoints)

#### Fixed
Expand Down
36 changes: 36 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# Muse workflow benchmark

`muse_workflow_h200.json` is the shared workload for a paired native and
metadata-workflow benchmark of the published Muse Glimmer INT4 package. Both
paths use the exact 68 token IDs in `muse_prompt_ids.json`, greedy parameters,
token budget, warmups, and steady-decode window. Passing token IDs avoids a
tokenizer adding another beginning-of-text token to the already rendered prompt.
Both runners set `ORT_ENABLE_CUDNN_FLASH_ATTENTION=0`; this matches the native
baseline and avoids comparing different GQA attention kernels.
`request_max_length` is prompt tokens plus new tokens (68 + 128 = 196);
`model_max_context` is the independent 131072-token artifact admission ceiling.
For the text-only VLM path, omit `request.image`. The runtime initializes
`request.image_present=false`, and the false branch supplies empty image features.

Run the native ORT GenAI path:

```bash
python scripts/benchmark_muse_native.py \
--model artifacts/muse-int4-package \
--output artifacts/muse-int4-package/native-benchmark.json
```

Run the ONNX GenAI workflow path with the `profile_native` binary built from
the schema/runtime revision named in the JSON config:

```bash
python scripts/benchmark_muse_workflow.py \
--model artifacts/muse-int4-package \
--runner path/to/profile_native \
--output artifacts/muse-int4-package/workflow-benchmark.json
```

The workflow runner must support `--pipeline --backend ort --ep cuda`,
`--prompt-ids`, and an optional `--image` request binding. The text-only workload
intentionally leaves the image unset so results remain comparable to the
published 61.76 tok/s baseline.
58 changes: 58 additions & 0 deletions benchmarks/muse_568457_release_evidence.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
{
"kind": "workflow_release_diagnostic",
"producer_commit": "80acfc069f0959f9e6580785fad3172bcc4cc0aa",
"runtime_source_head": "568457045ea003e787816995a380bcd9d7443169",
"runner_sha256": "fad10c7c8735c56e42acdfc092c4c6c327090c50de823efc45d4f3309b2d4aa9",
"overlay_manifest_sha256": "fd77589ffb9333d69efe78909b7f4a489da1071a35affb49b243e87a76330c6a",
"metadata_sha256": "8727cb6ab44ff65a4e74b6281c5af9eadba336483f5fa1e6033996c3b99fcced",
"decoder_state_initializer_sha256": "77e4eae0e35cc618214476e0776af3c1b7930862924f431edb62724cadb97af7",
"workload": {
"prompt_tokens": 68,
"generated_tokens": 128,
"warmups": 1,
"runs": 3,
"decode_skip": 8,
"stop_on_eos": false,
"image": null
},
"metrics": {
"ttft_ms": 67.512,
"decode_ms_per_token": 29.168,
"throughput_tok_s": 34.28,
"native_ttft_ms": 49.02172100264579,
"native_throughput_tok_s": 63.36527373646786,
"throughput_ratio": 0.5409903244885887
},
"diagnostics": {
"captures": 1,
"replays": 510,
"synchronizations": 2580,
"h2d_calls": 1556,
"h2d_bytes": 413802656,
"d2h_calls": 512,
"d2h_bytes": 4096,
"d2d_calls": 3068,
"d2d_bytes": 17376,
"island_elapsed_ms": 2974.922
},
"token_parity": {
"exact": false,
"first_difference": {
"index": 38,
"runtime": 4243,
"native": 33386
},
"runtime_count": 128,
"native_count": 128
},
"gate": {
"required_throughput_ratio": 0.99,
"correctness": false,
"paired_performance": false,
"upload_allowed": false
},
"notes": [
"The decoder remained outside the captured policy island.",
"This exact prompt-ID runtime still materialized decoder logits through host memory."
]
}
Loading
Loading