Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,14 @@ Eight curated matrices ship with this bundle, plus one explicit-name alias

> `openai-knob-consistent` was **removed on 2026-09-07**. Once the `preset:` block became `openai`'s default on 2026-09-02, the two files were the same matrix under two names. Select **`openai`** instead -- it is byte-for-byte what `openai-knob-consistent` used to give you.

### Opt-in canary matrices

Not "curated" in the same permanent sense as the eight above, and not selected by anyone automatically -- these exist so a rollout can be tried, measured and dropped without touching a default:

| Matrix | When to use |
|--------|-------------|
| **openai-gpt6-canary** | **Opt-in only** -- select explicitly (`amplifier routing use openai-gpt6-canary` or `routing.matrix: openai-gpt6-canary` in settings.yaml). GPT-6 (exact `gpt-6-luna` / `gpt-6-sol` / `gpt-6-astra` IDs, no globs) on `fast`/`coding`/`reasoning` only; every other role is byte-for-byte identical to `openai.yaml`. Strict knob-consistent delegation on a four-rung ladder (`luna` < `terra` < `sol` < `astra`). **Exact IDs are not preflighted against any model catalog or provider capability declaration** -- `fast`/`coding` have no runtime fallback and depend entirely on the companion `microsoft/amplifier-module-provider-openai` PR (GPT-6 Sol/Luna support) landing FIRST; do not enable this canary for those two roles before that PR merges. Tracks `microsoft-amplifier/amplifier-support#524` -- see the file header in [`routing/openai-gpt6-canary.yaml`](routing/openai-gpt6-canary.yaml) for the full rationale, the dependency note, and its documented limitations. |

Browse the matrix files directly in the [`routing/`](routing/) directory.

## Including the Bundle
Expand Down Expand Up @@ -129,7 +137,7 @@ Per-delegate `model_role` overrides (e.g. `delegate(agent="...", model_role="res

Levels 1 and 2 stay strictly above level 3, so an author who deliberately pinned a specialist model still gets one. Inheritance is a default, never a ceiling on explicit intent.

**Scoped by matrix, off unless the resolver can determine the caller's own model.** A matrix must carry a `preset:` block *and* the resolution must be able to determine the caller's own model. Only `openai.yaml` carries one. Every OTHER matrix shipped with this bundle still has no `preset:` key and resolves byte-identically — asserted, not asserted-to, by `tests/test_default_resolution_unchanged.py`, which replays a recording taken from the commit immediately before the feature landed, plus `modules/hooks-routing/tests/test_knob_consistent_routing.py::TestDefaultBehaviourUnchanged::test_anthropic_matrix_has_no_preset_block` and `::test_anthropic_root_resolution_unchanged_vs_pre_50_matrix` naming the Anthropic guardrail directly.
**Scoped by matrix, off unless the resolver can determine the caller's own model.** A matrix must carry a `preset:` block *and* the resolution must be able to determine the caller's own model. Of the matrices a user can land on by picking a name with no other setup, only `openai.yaml` carries one — every other **default-selectable** matrix (`anthropic`, `balanced`, `quality`, `economy`, `gemini`, `copilot`, `ollama`) still has no `preset:` key and resolves byte-identically, asserted, not asserted-to, by `tests/test_default_resolution_unchanged.py`, which replays a recording taken from the commit immediately before the feature landed, plus `modules/hooks-routing/tests/test_knob_consistent_routing.py::TestDefaultBehaviourUnchanged::test_anthropic_matrix_has_no_preset_block` and `::test_anthropic_root_resolution_unchanged_vs_pre_50_matrix` naming the Anthropic guardrail directly. The **opt-in** `openai-gpt6-canary.yaml` (see [Opt-in canary matrices](#opt-in-canary-matrices)) also carries a `preset:` block on purpose — it exists to canary knob-consistent delegation onto GPT-6 — but it postdates that recording (it never existed at the pre-feature commit) so it is excluded from it by name rather than silently passing, and it is separately allow-listed alongside `openai.yaml` in `test_knob_consistent_routing.py`'s own preset-bearing guard.

```yaml
# routing/openai.yaml -- shipped, DEFAULT ON
Expand Down
59 changes: 50 additions & 9 deletions docs/MATRIX_CURATOR_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -593,20 +593,61 @@ first candidate and nothing else. Real fallback needs a candidate on a
*different* provider; the caller's `model_role` list falls back across roles,
not across models.

### `reasoning_effort` Is Validated Per Provider, Not Per Model
**Consequence for a brand-new model family: an exact pin is never a proxy for
"has this model landed in the provider yet."** `resolve_model_role()` never
calls `list_models()` for an exact (non-glob) candidate at all -- it hands
back the literal string the instant the named provider family is installed,
whether or not the mounted provider *software* actually recognises that model
id. This makes exact-pin candidates for an unreleased-in-the-provider model a
strictly WORSE guarantee than a glob against a stale catalogue: a glob at
least fails closed (resolves to nothing, or falls through) when the id truly
isn't there yet, where an exact pin resolves regardless and pushes the
failure to request time, uncontrolled by this bundle.
[`routing/openai-gpt6-canary.yaml`](../routing/openai-gpt6-canary.yaml) is the
shipped example: GPT-6's model names have no numeric sub-version for a glob
to match, so its candidates are necessarily exact pins, and two of its three
roles (`fast`, `coding`) are consequently NOT protected by routing at all
against the companion provider PR being unmerged -- see that file's own "NO
CATALOG/CAPABILITY PREFLIGHT, THEREFORE NO RUNTIME FALLBACK" note before
copying this pattern for another pre-release model family. Do not describe an
exact pin on an unreleased model as "degrading gracefully" or "falling back"
in a PR description or matrix comment -- say plainly that it has no routing-
level protection, and name what (if anything) protects it instead (usually:
the companion provider PR must land and be verified first).

### `reasoning_effort` Is Validated Per Provider By Default, Per Model Where Declared

A matrix `config.reasoning_effort` is an **operator default** — a
caller-supplied effort wins.

`tests/matrix_validation_rules.yaml` checks the value against a **provider-wide**
vocabulary, keyed on `provider` alone. It cannot see whether the pinned *model*
advertises that level, so an effort that is legal for the provider but
unsupported by the model passes validation here and is left to the provider to
resolve at request time.

That file's header also records which of those lists are provider-derived and
which are only "the set already in use across the shipped matrices". Read the
levels from the installed provider before pinning one.
vocabulary, keyed on `provider` alone (`reasoning_effort_values`). By itself,
that check cannot see whether the pinned *model* advertises the level: two
models on the same provider can have genuinely different allowed sets (see
GPT-6 below), so a value that is legal for the provider but unsupported by
that specific model would otherwise pass validation here and be left to the
provider to resolve at request time.

**For a model with real per-model differences, add it to
`reasoning_effort_values_by_model` instead of relying on the provider-wide
list alone.** That second map is keyed by the *exact* model id (glob
candidates are not checked against it — a glob can resolve to more than one
concrete model, so only an exact id has one set to check against) and, when a
model has an entry there, `tests/test_matrix_config_validation.py` enforces it
as the authoritative, narrower check for every shipped matrix that names that
id. The GPT-6 family is the reason this map exists: `gpt-6-astra` accepts
`low|medium|high|xhigh|max` and **rejects `none`**, while `gpt-6-sol` and
`gpt-6-luna` accept `none` too — same provider (`openai`), three different
per-model sets, none of which the provider-wide `openai` entry alone can tell
apart. See [`routing/openai-gpt6-canary.yaml`](../routing/openai-gpt6-canary.yaml)
for the shipped example.

That file's header also records which of the provider-wide lists are
provider-derived and which are only "the set already in use across the
shipped matrices", and which of the per-model lists are sourced (with dates)
from the vendor's own model pages. Read the levels from the installed
provider (or its first-party docs, if a model is not yet fully live in the
provider package) before pinning one.

> **github-copilot only.** Once the model catalogue has loaded, an unadvertised
> effort is omitted from the request and reported as a `ChatResponse.degradation`
Expand Down
105 changes: 104 additions & 1 deletion modules/hooks-routing/tests/test_knob_consistent_routing.py
Original file line number Diff line number Diff line change
Expand Up @@ -163,7 +163,18 @@ def test_every_pre_existing_matrix_has_no_preset_block(self) -> None:
no measured win there yet, so no default change there."""
import yaml

allowed_with_preset = {"openai.yaml"}
# `openai-gpt6-canary.yaml` (amplifier-support#524, 2026-09-24) is a
# SEPARATE opt-in matrix, not a default: it is never selected unless
# a user names it explicitly (`amplifier routing use
# openai-gpt6-canary` / `routing.matrix:` in settings.yaml), and
# `behaviors/routing.yaml`'s `default_matrix` and hooks-routing's own
# code default both still say `balanced` -- untouched by this file.
# It carries a `preset:` block on purpose (a 4-rung strict ladder,
# see the file header) precisely because it exists to canary
# knob-consistent delegation for GPT-6, so it belongs on this
# allow-list for the same reason `openai.yaml` does: intentional,
# reviewed, not a default-behaviour regression.
allowed_with_preset = {"openai.yaml", "openai-gpt6-canary.yaml"}
offenders = []
for path in sorted(ROUTING_DIR.glob("*.yaml")):
data = yaml.safe_load(path.read_text(encoding="utf-8"))
Expand Down Expand Up @@ -963,3 +974,95 @@ async def test_a_top_tier_root_is_not_downgraded(self) -> None:
# ALONE. Both arms are now one file (`openai-knob-consistent.yaml` deleted
# 2026-09-07), so there is no longer a pair to hold identical -- the
# property is structural rather than tested.


# ---------------------------------------------------------------------------
# The opt-in GPT-6 canary matrix (amplifier-support#524), end to end
# ---------------------------------------------------------------------------


class TestGpt6CanaryMatrixPreset:
"""`routing/openai-gpt6-canary.yaml`'s `preset:` block, validated the
same way `TestShippedKnobConsistentMatrix` validates `openai.yaml`'s.

This matrix carries its own preset deliberately -- see that file's "WHY A
FOURTH RUNG" note -- and is allow-listed alongside `openai.yaml` in
`TestDefaultBehaviourUnchanged.test_every_pre_existing_matrix_has_no_preset_block`
above. What is asserted here is that the preset is well-formed, strict,
and four-rung, and that it behaves exactly like every other preset with
respect to the two guarantees the rest of this module establishes:
inert without a caller, and load-bearing with one.
"""

@staticmethod
def _load() -> dict[str, Any]:
import yaml

return yaml.safe_load(
(ROUTING_DIR / "openai-gpt6-canary.yaml").read_text(encoding="utf-8")
)

def test_preset_block_is_strict_four_rung(self) -> None:
data = self._load()
preset = parse_preset(data)
assert preset is not None
assert preset.inherit == "strict"
rungs = preset.tier_ladder["openai"]
assert len(rungs) == 4
assert rungs[-1] == ["gpt-6-astra"]

@pytest.mark.asyncio
async def test_cold_resolution_is_unaffected_by_the_preset(self) -> None:
"""The same "preset is opt-in twice over" invariant
`test_preset_bearing_matrix_is_stock_without_a_caller` checks for
`openai.yaml`, run directly against this matrix rather than through
the golden fixture (which excludes this file -- it postdates the
pre-feature recording; see `EXCLUDED_FROM_RECORDING` in
`tests/test_default_resolution_unchanged.py`)."""
data = self._load()
roles = data["roles"]
preset = parse_preset(data)
assert preset is not None
providers = {"openai": _openai_provider()}
providers["openai"].list_models = AsyncMock(
return_value=[
"gpt-6-astra",
"gpt-6-sol",
"gpt-6-luna",
"gpt-5.6-terra",
"gpt-5.6-luna",
]
)
for role in roles:
cold = await resolve_model_role([role], roles, providers)
with_preset = await resolve_model_role(
[role], roles, providers, preset=preset, caller_context=None
)
assert with_preset == cold, (
f"role {role}: attaching the parsed preset with no caller "
"context changed cold resolution"
)

@pytest.mark.asyncio
async def test_terra_root_delegate_gets_clamped_off_astra(self) -> None:
"""End to end, through `mount()`, against the REAL shipped
`openai-gpt6-canary.yaml`: a terra-tier root's `reasoning` sub-agent
must not land on `gpt-6-astra` -- it is clamped to the matrix's
second `reasoning` candidate, `gpt-5.6-terra`, by strict inherit.

`_mount_coordinator`'s default provider spec (`id: "terra"`,
`default_model: "gpt-5.6-terra"`) is what derives the terra-tier
caller context here; its default `_providers()` roster already
contains `gpt-5.6-terra` (for the glob candidate), and `gpt-6-astra`
needs no catalog entry at all (exact id -- see the matrix file's "NO
CATALOG/CAPABILITY PREFLIGHT" note)."""
agents = {"explorer": {"model_role": "reasoning"}}
coordinator = _mount_coordinator(agents)
await mount(
coordinator,
{"default_matrix": "openai-gpt6-canary", "_bundle_root": str(REPO_ROOT)},
)
await _run_session_start(coordinator)
assert agents["explorer"]["provider_preferences"][0]["model"] == (
"gpt-5.6-terra"
)
Loading
Loading