diff --git a/README.md b/README.md index 68db10f..89b0ff3 100644 --- a/README.md +++ b/README.md @@ -22,6 +22,14 @@ Eight curated matrices ship with this bundle, plus one explicit-name alias > `openai-knob-consistent` was **removed on 2026-09-07**. Once the `preset:` block became `openai`'s default on 2026-09-02, the two files were the same matrix under two names. Select **`openai`** instead -- it is byte-for-byte what `openai-knob-consistent` used to give you. +### Opt-in canary matrices + +Not "curated" in the same permanent sense as the eight above, and not selected by anyone automatically -- these exist so a rollout can be tried, measured and dropped without touching a default: + +| Matrix | When to use | +|--------|-------------| +| **openai-gpt6-canary** | **Opt-in only** -- select explicitly (`amplifier routing use openai-gpt6-canary` or `routing.matrix: openai-gpt6-canary` in settings.yaml). GPT-6 (exact `gpt-6-luna` / `gpt-6-sol` / `gpt-6-astra` IDs, no globs) on `fast`/`coding`/`reasoning` only; every other role is byte-for-byte identical to `openai.yaml`. Strict knob-consistent delegation on a four-rung ladder (`luna` < `terra` < `sol` < `astra`). **Exact IDs are not preflighted against any model catalog or provider capability declaration** -- `fast`/`coding` have no runtime fallback and depend entirely on the companion `microsoft/amplifier-module-provider-openai` PR (GPT-6 Sol/Luna support) landing FIRST; do not enable this canary for those two roles before that PR merges. Tracks `microsoft-amplifier/amplifier-support#524` -- see the file header in [`routing/openai-gpt6-canary.yaml`](routing/openai-gpt6-canary.yaml) for the full rationale, the dependency note, and its documented limitations. | + Browse the matrix files directly in the [`routing/`](routing/) directory. ## Including the Bundle @@ -129,7 +137,7 @@ Per-delegate `model_role` overrides (e.g. `delegate(agent="...", model_role="res Levels 1 and 2 stay strictly above level 3, so an author who deliberately pinned a specialist model still gets one. Inheritance is a default, never a ceiling on explicit intent. -**Scoped by matrix, off unless the resolver can determine the caller's own model.** A matrix must carry a `preset:` block *and* the resolution must be able to determine the caller's own model. Only `openai.yaml` carries one. Every OTHER matrix shipped with this bundle still has no `preset:` key and resolves byte-identically — asserted, not asserted-to, by `tests/test_default_resolution_unchanged.py`, which replays a recording taken from the commit immediately before the feature landed, plus `modules/hooks-routing/tests/test_knob_consistent_routing.py::TestDefaultBehaviourUnchanged::test_anthropic_matrix_has_no_preset_block` and `::test_anthropic_root_resolution_unchanged_vs_pre_50_matrix` naming the Anthropic guardrail directly. +**Scoped by matrix, off unless the resolver can determine the caller's own model.** A matrix must carry a `preset:` block *and* the resolution must be able to determine the caller's own model. Of the matrices a user can land on by picking a name with no other setup, only `openai.yaml` carries one — every other **default-selectable** matrix (`anthropic`, `balanced`, `quality`, `economy`, `gemini`, `copilot`, `ollama`) still has no `preset:` key and resolves byte-identically, asserted, not asserted-to, by `tests/test_default_resolution_unchanged.py`, which replays a recording taken from the commit immediately before the feature landed, plus `modules/hooks-routing/tests/test_knob_consistent_routing.py::TestDefaultBehaviourUnchanged::test_anthropic_matrix_has_no_preset_block` and `::test_anthropic_root_resolution_unchanged_vs_pre_50_matrix` naming the Anthropic guardrail directly. The **opt-in** `openai-gpt6-canary.yaml` (see [Opt-in canary matrices](#opt-in-canary-matrices)) also carries a `preset:` block on purpose — it exists to canary knob-consistent delegation onto GPT-6 — but it postdates that recording (it never existed at the pre-feature commit) so it is excluded from it by name rather than silently passing, and it is separately allow-listed alongside `openai.yaml` in `test_knob_consistent_routing.py`'s own preset-bearing guard. ```yaml # routing/openai.yaml -- shipped, DEFAULT ON diff --git a/docs/MATRIX_CURATOR_GUIDE.md b/docs/MATRIX_CURATOR_GUIDE.md index 99a1f31..4c45d96 100644 --- a/docs/MATRIX_CURATOR_GUIDE.md +++ b/docs/MATRIX_CURATOR_GUIDE.md @@ -593,20 +593,61 @@ first candidate and nothing else. Real fallback needs a candidate on a *different* provider; the caller's `model_role` list falls back across roles, not across models. -### `reasoning_effort` Is Validated Per Provider, Not Per Model +**Consequence for a brand-new model family: an exact pin is never a proxy for +"has this model landed in the provider yet."** `resolve_model_role()` never +calls `list_models()` for an exact (non-glob) candidate at all -- it hands +back the literal string the instant the named provider family is installed, +whether or not the mounted provider *software* actually recognises that model +id. This makes exact-pin candidates for an unreleased-in-the-provider model a +strictly WORSE guarantee than a glob against a stale catalogue: a glob at +least fails closed (resolves to nothing, or falls through) when the id truly +isn't there yet, where an exact pin resolves regardless and pushes the +failure to request time, uncontrolled by this bundle. +[`routing/openai-gpt6-canary.yaml`](../routing/openai-gpt6-canary.yaml) is the +shipped example: GPT-6's model names have no numeric sub-version for a glob +to match, so its candidates are necessarily exact pins, and two of its three +roles (`fast`, `coding`) are consequently NOT protected by routing at all +against the companion provider PR being unmerged -- see that file's own "NO +CATALOG/CAPABILITY PREFLIGHT, THEREFORE NO RUNTIME FALLBACK" note before +copying this pattern for another pre-release model family. Do not describe an +exact pin on an unreleased model as "degrading gracefully" or "falling back" +in a PR description or matrix comment -- say plainly that it has no routing- +level protection, and name what (if anything) protects it instead (usually: +the companion provider PR must land and be verified first). + +### `reasoning_effort` Is Validated Per Provider By Default, Per Model Where Declared A matrix `config.reasoning_effort` is an **operator default** — a caller-supplied effort wins. `tests/matrix_validation_rules.yaml` checks the value against a **provider-wide** -vocabulary, keyed on `provider` alone. It cannot see whether the pinned *model* -advertises that level, so an effort that is legal for the provider but -unsupported by the model passes validation here and is left to the provider to -resolve at request time. - -That file's header also records which of those lists are provider-derived and -which are only "the set already in use across the shipped matrices". Read the -levels from the installed provider before pinning one. +vocabulary, keyed on `provider` alone (`reasoning_effort_values`). By itself, +that check cannot see whether the pinned *model* advertises the level: two +models on the same provider can have genuinely different allowed sets (see +GPT-6 below), so a value that is legal for the provider but unsupported by +that specific model would otherwise pass validation here and be left to the +provider to resolve at request time. + +**For a model with real per-model differences, add it to +`reasoning_effort_values_by_model` instead of relying on the provider-wide +list alone.** That second map is keyed by the *exact* model id (glob +candidates are not checked against it — a glob can resolve to more than one +concrete model, so only an exact id has one set to check against) and, when a +model has an entry there, `tests/test_matrix_config_validation.py` enforces it +as the authoritative, narrower check for every shipped matrix that names that +id. The GPT-6 family is the reason this map exists: `gpt-6-astra` accepts +`low|medium|high|xhigh|max` and **rejects `none`**, while `gpt-6-sol` and +`gpt-6-luna` accept `none` too — same provider (`openai`), three different +per-model sets, none of which the provider-wide `openai` entry alone can tell +apart. See [`routing/openai-gpt6-canary.yaml`](../routing/openai-gpt6-canary.yaml) +for the shipped example. + +That file's header also records which of the provider-wide lists are +provider-derived and which are only "the set already in use across the +shipped matrices", and which of the per-model lists are sourced (with dates) +from the vendor's own model pages. Read the levels from the installed +provider (or its first-party docs, if a model is not yet fully live in the +provider package) before pinning one. > **github-copilot only.** Once the model catalogue has loaded, an unadvertised > effort is omitted from the request and reported as a `ChatResponse.degradation` diff --git a/modules/hooks-routing/tests/test_knob_consistent_routing.py b/modules/hooks-routing/tests/test_knob_consistent_routing.py index 6597406..ee53bbd 100644 --- a/modules/hooks-routing/tests/test_knob_consistent_routing.py +++ b/modules/hooks-routing/tests/test_knob_consistent_routing.py @@ -163,7 +163,18 @@ def test_every_pre_existing_matrix_has_no_preset_block(self) -> None: no measured win there yet, so no default change there.""" import yaml - allowed_with_preset = {"openai.yaml"} + # `openai-gpt6-canary.yaml` (amplifier-support#524, 2026-09-24) is a + # SEPARATE opt-in matrix, not a default: it is never selected unless + # a user names it explicitly (`amplifier routing use + # openai-gpt6-canary` / `routing.matrix:` in settings.yaml), and + # `behaviors/routing.yaml`'s `default_matrix` and hooks-routing's own + # code default both still say `balanced` -- untouched by this file. + # It carries a `preset:` block on purpose (a 4-rung strict ladder, + # see the file header) precisely because it exists to canary + # knob-consistent delegation for GPT-6, so it belongs on this + # allow-list for the same reason `openai.yaml` does: intentional, + # reviewed, not a default-behaviour regression. + allowed_with_preset = {"openai.yaml", "openai-gpt6-canary.yaml"} offenders = [] for path in sorted(ROUTING_DIR.glob("*.yaml")): data = yaml.safe_load(path.read_text(encoding="utf-8")) @@ -963,3 +974,95 @@ async def test_a_top_tier_root_is_not_downgraded(self) -> None: # ALONE. Both arms are now one file (`openai-knob-consistent.yaml` deleted # 2026-09-07), so there is no longer a pair to hold identical -- the # property is structural rather than tested. + + +# --------------------------------------------------------------------------- +# The opt-in GPT-6 canary matrix (amplifier-support#524), end to end +# --------------------------------------------------------------------------- + + +class TestGpt6CanaryMatrixPreset: + """`routing/openai-gpt6-canary.yaml`'s `preset:` block, validated the + same way `TestShippedKnobConsistentMatrix` validates `openai.yaml`'s. + + This matrix carries its own preset deliberately -- see that file's "WHY A + FOURTH RUNG" note -- and is allow-listed alongside `openai.yaml` in + `TestDefaultBehaviourUnchanged.test_every_pre_existing_matrix_has_no_preset_block` + above. What is asserted here is that the preset is well-formed, strict, + and four-rung, and that it behaves exactly like every other preset with + respect to the two guarantees the rest of this module establishes: + inert without a caller, and load-bearing with one. + """ + + @staticmethod + def _load() -> dict[str, Any]: + import yaml + + return yaml.safe_load( + (ROUTING_DIR / "openai-gpt6-canary.yaml").read_text(encoding="utf-8") + ) + + def test_preset_block_is_strict_four_rung(self) -> None: + data = self._load() + preset = parse_preset(data) + assert preset is not None + assert preset.inherit == "strict" + rungs = preset.tier_ladder["openai"] + assert len(rungs) == 4 + assert rungs[-1] == ["gpt-6-astra"] + + @pytest.mark.asyncio + async def test_cold_resolution_is_unaffected_by_the_preset(self) -> None: + """The same "preset is opt-in twice over" invariant + `test_preset_bearing_matrix_is_stock_without_a_caller` checks for + `openai.yaml`, run directly against this matrix rather than through + the golden fixture (which excludes this file -- it postdates the + pre-feature recording; see `EXCLUDED_FROM_RECORDING` in + `tests/test_default_resolution_unchanged.py`).""" + data = self._load() + roles = data["roles"] + preset = parse_preset(data) + assert preset is not None + providers = {"openai": _openai_provider()} + providers["openai"].list_models = AsyncMock( + return_value=[ + "gpt-6-astra", + "gpt-6-sol", + "gpt-6-luna", + "gpt-5.6-terra", + "gpt-5.6-luna", + ] + ) + for role in roles: + cold = await resolve_model_role([role], roles, providers) + with_preset = await resolve_model_role( + [role], roles, providers, preset=preset, caller_context=None + ) + assert with_preset == cold, ( + f"role {role}: attaching the parsed preset with no caller " + "context changed cold resolution" + ) + + @pytest.mark.asyncio + async def test_terra_root_delegate_gets_clamped_off_astra(self) -> None: + """End to end, through `mount()`, against the REAL shipped + `openai-gpt6-canary.yaml`: a terra-tier root's `reasoning` sub-agent + must not land on `gpt-6-astra` -- it is clamped to the matrix's + second `reasoning` candidate, `gpt-5.6-terra`, by strict inherit. + + `_mount_coordinator`'s default provider spec (`id: "terra"`, + `default_model: "gpt-5.6-terra"`) is what derives the terra-tier + caller context here; its default `_providers()` roster already + contains `gpt-5.6-terra` (for the glob candidate), and `gpt-6-astra` + needs no catalog entry at all (exact id -- see the matrix file's "NO + CATALOG/CAPABILITY PREFLIGHT" note).""" + agents = {"explorer": {"model_role": "reasoning"}} + coordinator = _mount_coordinator(agents) + await mount( + coordinator, + {"default_matrix": "openai-gpt6-canary", "_bundle_root": str(REPO_ROOT)}, + ) + await _run_session_start(coordinator) + assert agents["explorer"]["provider_preferences"][0]["model"] == ( + "gpt-5.6-terra" + ) diff --git a/routing/openai-gpt6-canary.yaml b/routing/openai-gpt6-canary.yaml new file mode 100644 index 0000000..0596cfd --- /dev/null +++ b/routing/openai-gpt6-canary.yaml @@ -0,0 +1,344 @@ +# Routing Matrix: openai-gpt6-canary +# +# OPT-IN CANARY. Not the default, not referenced by behaviors/routing.yaml, +# and not swapped in for anyone automatically. A user selects it explicitly: +# +# amplifier routing use openai-gpt6-canary +# # or, in settings.yaml: +# routing: +# matrix: openai-gpt6-canary +# +# WHAT THIS IS. `openai.yaml` (the shipped default OpenAI matrix) is +# untouched by this file. This is a SEPARATE, ninth matrix that carries the +# stage-1 opt-in canary shape recommended by the routing study attached to +# `microsoft-amplifier/amplifier-support#524`: "Canary `gpt-6-luna` for `fast` and +# `gpt-6-sol` for `coding`; offer `gpt-6-astra` only as an opt-in `reasoning` +# candidate until cost, latency, error, safety-stop, and task-quality +# evidence meets promotion gates." Every role other than `fast`/`coding`/ +# `reasoning` is copied unchanged from `openai.yaml` (still `gpt-5.6-terra`). +# +# ATTRIBUTION, STATED PRECISELY. The study recommends the three role/model +# canary pairings above and warns, generally, against reusing the existing +# Luna/Terra/Sol tier ladder for GPT-6 rungs "without policy" (see "WHY A +# FOURTH RUNG" below). It does NOT specify a four-rung-vs-three-rung ladder +# shape, and it does NOT call for a `gpt-5.6-terra` fallback candidate on +# `reasoning` -- neither appears in the study text. Both are THIS ROLLOUT'S +# OWN design decisions, made to satisfy this task's requirement to keep the +# canary narrow and reversible; they are consistent with the study's +# guidance but not a direct quote of it, and are documented as such below +# rather than attributed to the study. +# +# DEPENDENCY -- PROVIDER PR MUST LAND FIRST. `gpt-6-sol` and `gpt-6-luna` +# request/response handling (reasoning controls, per-model effort sets, +# streaming, structured outputs, cost accounting) ships in a companion PR to +# `microsoft/amplifier-module-provider-openai`, tracked by +# `microsoft/amplifier/amplifier-support#524`. +# Provider PR: https://github.com/microsoft/amplifier-module-provider-openai/pull/112 +# +# THIS ROUTING PR MUST MERGE AFTER THAT ONE. Selecting this matrix for +# `fast`/`coding` before the provider PR lands is NOT a safe, silently- +# degrading configuration -- see "NO CATALOG/CAPABILITY PREFLIGHT, +# THEREFORE NO RUNTIME FALLBACK" immediately below for exactly why, and do +# not enable this canary for those two roles until the provider PR has +# merged and been verified. +# +# NO CATALOG/CAPABILITY PREFLIGHT, THEREFORE NO RUNTIME FALLBACK. The study +# observes this about the resolver directly (its section 3, citing +# `resolver.py#L453-L465`): "Exact IDs are accepted without a catalog/ +# capability check ... Consequently, an exact account-gated ID can resolve +# and fail only at request time; candidate fallback is not a general +# runtime failover mechanism." Confirmed by reading that code path in this +# tree: for any candidate whose `model:` is NOT a glob (no `*`/`?`/`[`), +# `resolve_model_role` never calls `list_models()` and never checks the +# model against any capability table -- it sets `resolved_model = +# model_pattern` unconditionally, the instant `find_provider_by_type` finds +# ANY installed instance of the named provider family. Every GPT-6 +# candidate in this file is an exact ID (see "EXACT MODEL IDS ONLY" below), +# so this applies to all three canaried roles: +# +# - `fast` and `coding` each carry exactly ONE candidate +# (`gpt-6-luna`/`gpt-6-sol`). Because exact-ID resolution never fails at +# the routing layer for the reason above, there is no second candidate +# to fall through to and no mechanism that would trigger a fallback even +# if there were: routing hands back the literal id as "resolved" +# whenever any `openai`-family provider is mounted, REGARDLESS of +# whether the mounted provider software actually implements +# `gpt-6-luna`/`gpt-6-sol` yet. Whatever happens next -- accept, +# generically default, or reject the model id -- happens at the +# provider/request layer, entirely outside this bundle's control, and +# depends entirely on whether the companion provider PR has merged. If +# it has not, do not expect a graceful degrade: expect whatever +# `microsoft/amplifier-module-provider-openai`'s CURRENT (pre-#524) code +# does with an id it does not recognise, which this routing PR does not +# verify and does not control. +# - `reasoning` carries `gpt-6-astra` first and `gpt-5.6-terra` second. +# `gpt-6-astra` is ALSO an exact ID and is therefore ALSO never +# preflighted, so the same "resolves regardless of true provider +# support" fact applies to it. In the common case (no caller context -- +# a top-level session, or `amplifier routing show`), the second +# candidate is consequently NEVER REACHED: `resolve_model_role` returns +# as soon as the first candidate resolves, and the first candidate +# always "resolves" once `openai` is mounted. The only path that +# actually reaches `gpt-5.6-terra` is the STRICT knob-consistent +# delegation clamp (see "STRICT, FOUR-RUNG LADDER" below): +# `plan_candidates()` in `knob_consistency.py` rewrites the candidate +# list itself, before the ordinary loop runs, for a sub-agent whose +# caller is resolved below the `astra` rung -- filtering `gpt-6-astra` +# out and leaving `gpt-5.6-terra` as the new top candidate. That is a +# real, tested mechanism (see +# `tests/test_gpt6_canary_routing.py::test_reasoning_clamps_to_terra_for_a_terra_tier_caller`), +# but it is a DELEGATION-TIER clamp, not a capability/outage/unmerged-PR +# degrade path. Astra's provider support is already shipped (unlike +# Sol/Luna), which narrows this gap for `reasoning` relative to +# `fast`/`coding`, but the underlying mechanism -- and the underlying +# absence of any capability check -- is identical. +# +# EXACT MODEL IDS ONLY -- NO GLOBS FOR GPT-6. Per the study (cross-checked +# against live `openai.com` docs and a live `GET /v1/models` call, +# 2026-09-24): the GPT-6 family is exactly three model IDs -- `gpt-6-astra`, +# `gpt-6-sol`, `gpt-6-luna` -- with no aliases and no dated snapshots. A +# version glob (`gpt-?.?-*`) cannot express these names (they carry no +# numeric sub-version), so every GPT-6 candidate below pins the bare id. +# This also means there is no "-fast" ChatGPT-backend suffix risk for GPT-6 +# the way there is for `gpt-5.6-terra*` (see openai.yaml) -- nothing here +# ends in `*`. See "NO CATALOG/CAPABILITY PREFLIGHT" above for what pinning +# an exact id actually costs: no catalog check, ever, for these three. +# +# PER-MODEL EFFORT VALIDATION. The three GPT-6 models do NOT all accept the +# same `reasoning_effort` set (source: the study's section on OpenAI's +# first-party model pages, verified 2026-09-24): Sol and Luna accept +# `none|low|medium|high|xhigh|max` (default `medium`); Astra accepts +# `low|medium|high|xhigh|max` and REJECTS `none`. This matrix requests only +# low/medium/high across its GPT-6 candidates, all valid for every GPT-6 +# model, so no candidate here can hit an invalid combination. Enforcement +# lives in the repository's own per-model validation data, +# `tests/matrix_validation_rules.yaml`'s `reasoning_effort_values_by_model` +# map, exercised by `tests/test_matrix_config_validation.py` against every +# shipped matrix (including this one) and, redundantly, by +# `tests/test_gpt6_canary_routing.py` for this file specifically -- so a +# future edit that moves a role to `none` on `gpt-6-astra` fails loudly +# instead of going inert. +# +# WHY A FOURTH RUNG (this rollout's own choice, not a study quote -- see +# "ATTRIBUTION" above). The study's only ladder guidance is: "add GPT-6 +# Luna/Sol/Astra rungs deliberately... Astra and Sol pricing/quality tiers +# must not be guessed into the old Luna/Terra/Sol ladder without policy." +# Folding `gpt-6-astra` onto the existing `sol` rung would be exactly that +# kind of unpoliced guess, and it has a concrete bad consequence for +# `strict` delegation: same rung means no clamp, so a `sol`-tier caller's +# sub-agents could escalate onto `astra` -- the escalation `strict` exists +# to prevent. Astra is also already live and unpaused (unlike `sol`, which +# is PAUSED in `openai.yaml`'s own ladder), so treating it as sol-equivalent +# would misclassify its actual status. Astra gets its own top rung instead: +# a caller has to be AT that rung (i.e. already resolved to `gpt-6-astra`) +# for `reasoning` sub-delegation to land on it unclamped. +# +# STRICT, FOUR-RUNG LADDER (cheapest -> most expensive): +# 1. luna (gpt-?.?-luna* / gpt-?.?-mini* / gpt-?.?-nano* / gpt-6-luna) +# 2. terra (gpt-?.?-terra*) +# 3. sol (gpt-?.?-sol* / gpt-[0-9].[0-9] / gpt-6-sol) +# 4. astra (gpt-6-astra) -- its own rung, per "WHY A FOURTH RUNG" above. +# `inherit: strict` (same semantics as openai.yaml -- see that file's and +# the bundle README's "Knob-consistent delegation" section): a sub-agent's +# model_role resolution is clamped to the caller's own rung and effort +# ceiling, with escalation denied outright. A `gpt-6-sol @ medium` root's +# sub-agents cannot land on `gpt-6-astra`; a `gpt-6-astra @ high` root's +# sub-agents can. +# +# PER-ROLE CANARY SCOPE, STATED EXACTLY (only these three roles change vs. +# openai.yaml; every other role below is copied unchanged, description text +# included -- character-for-character equal to `openai.yaml`'s): +# - fast: gpt-6-luna @ low (single candidate -- no fallback of +# any kind; see "NO CATALOG/CAPABILITY PREFLIGHT" above) +# - coding: gpt-6-sol @ medium (single candidate -- no fallback of +# any kind; see "NO CATALOG/CAPABILITY PREFLIGHT" above) +# - reasoning: gpt-6-astra @ high, THEN gpt-5.6-terra @ xhigh as a second +# candidate. This is this rollout's own choice (not the +# study's, see "ATTRIBUTION" above), kept as the strict- +# ladder's landing target for a below-`astra`-rung caller's +# delegated `reasoning` sub-agents -- it is exercised ONLY by +# that clamp path, never by a cold/no-caller-context +# resolution and never by any capability or outage +# condition; see "NO CATALOG/CAPABILITY PREFLIGHT" above. +# `fast` and `coding` were deliberately left WITHOUT a second +# candidate at all -- one would be equally inert for the same +# reason and would misleadingly read as a safety net that +# does not exist. +# +# DOCUMENTED LIMITATIONS: +# 1. `fast` and `coding` have NO runtime protection of any kind against +# the companion provider PR being unmerged, absent, reverted, or +# incomplete. See "NO CATALOG/CAPABILITY PREFLIGHT, THEREFORE NO +# RUNTIME FALLBACK" above. Do not enable this canary for those two +# roles before that PR has merged and been verified. +# 2. ChatGPT-backend (`openai-chatgpt`) availability of `gpt-6-sol` / +# `gpt-6-luna` is UNVERIFIED as of 2026-09-24 -- the study confirmed +# `gpt-6-astra` live on the API-key backend via a read-only `GET +# /v1/models` call, but explicitly marks whether the ChatGPT backend +# serves Sol/Luna as UNKNOWN and warns it "must not be inferred from +# API-key availability." Every candidate below still says `provider: +# openai`, so `PROVIDER_FAMILY_ALIASES` will route a ChatGPT-only user +# to whatever that backend actually serves -- and per limitation 1, +# there is no routing-layer signal either way if it does not yet serve +# these ids. +# 3. This matrix does not alter `openai.yaml`, `behaviors/routing.yaml`'s +# `default_matrix`, or any resolver code -- it is additive and inert +# until a user opts in. `tests/test_default_resolution_unchanged.py` +# asserts every OTHER matrix's cold resolution is unaffected by this +# file's existence, and this file is explicitly excluded from that +# golden-fixture recording (it is new, so it was never in the +# pre-feature recording to begin with -- see that test's +# `EXCLUDED_FROM_RECORDING` comment for why exclusion, not silence, is +# the correct declaration). +# 4. `ui-coding`, `security-audit`, `critique`, `creative`, `writing`, +# `research`, `vision`, `image-gen`, `critical-ops` and `general` are +# unchanged from `openai.yaml` (still `gpt-5.6-terra`/`gpt-?.?-luna*`) +# -- the study's evidence table names them Possible at best, not +# Recommended, for GPT-6 today, so they are out of scope for this +# canary, not overlooked. +# +# --- Review log --- +# 2026-09-24: Created for microsoft-amplifier/amplifier-support#524, per its +# linked Amplifier model-routing study. Per-model effort data verified +# directly against OpenAI's first-party model pages (see the study, +# section citing developers.openai.com), independent of any provider-repo +# companion PR -- this matrix does not depend on that PR's code, only on +# its shipping (see DEPENDENCY above). +# 2026-09-24: Reviewed and corrected -- removed false claims that exact-ID +# GPT-6 candidates degrade gracefully via routing-level fallback (they do +# not: see "NO CATALOG/CAPABILITY PREFLIGHT" above) and corrected the +# attribution of the four-rung ladder and the `reasoning` fallback +# candidate, neither of which is a literal recommendation of the routing +# study (see "ATTRIBUTION" above). + +name: openai-gpt6-canary +description: "OPT-IN canary (amplifier-support#524): exact-ID GPT-6 (gpt-6-luna/gpt-6-sol/gpt-6-astra, no globs) on fast/coding/reasoning only; every other role is unchanged from openai.yaml. Strict knob-consistent delegation on a 4-rung ladder (luna < terra < sol < astra). Exact IDs are NOT preflighted against any catalog or provider capability declaration: fast/coding have no runtime fallback and require the companion microsoft/amplifier-module-provider-openai PR to merge FIRST -- do not enable this canary for those two roles before then." +updated: "2026-09-24" + +# --------------------------------------------------------------------------- +# Knob-consistent delegation preset -- same `strict` semantics as +# openai.yaml, extended to a fourth rung for gpt-6-astra. See this file's +# header ("WHY A FOURTH RUNG", "STRICT, FOUR-RUNG LADDER") for the rationale. +# --------------------------------------------------------------------------- +preset: + axis: [cheap, fast] + + tier_ladder: + openai: + - ["gpt-?.?-luna*", "gpt-?.?-mini*", "gpt-?.?-nano*", "gpt-6-luna"] + - ["gpt-?.?-terra*"] + - ["gpt-?.?-sol*", "gpt-[0-9].[0-9]", "gpt-6-sol"] + - ["gpt-6-astra"] + + delegation: + inherit: strict + report_unhonored: true + +roles: + general: + description: "Versatile catch-all, no specialization needed" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: high + + fast: + description: "Quick utility tasks — parsing, classification, file ops, bulk work" + candidates: + - provider: openai + model: gpt-6-luna + config: + reasoning_effort: low + + coding: + description: "Code generation, implementation, debugging" + candidates: + - provider: openai + model: gpt-6-sol + config: + reasoning_effort: medium + + ui-coding: + description: "Frontend/UI code — components, layouts, styling, spatial reasoning" + candidates: + - provider: openai + model: gpt-?.?-luna + config: + reasoning_effort: max + + security-audit: + description: "Vulnerability assessment, attack surface analysis, code auditing" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: xhigh + + reasoning: + description: "Deep architectural reasoning, system design, complex multi-step analysis" + candidates: + - provider: openai + model: gpt-6-astra + config: + reasoning_effort: high + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: xhigh + + critique: + description: "Analytical evaluation — finding flaws in existing work, not generating solutions" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: xhigh + + creative: + description: "Design direction, aesthetic judgment, high-quality creative output" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: high + + writing: + description: "Long-form content — documentation, marketing, case studies, storytelling" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: high + + research: + description: "Deep investigation, information synthesis across multiple sources" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: xhigh + + vision: + description: "Understanding visual input — screenshots, diagrams, UI mockups" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: medium + + image-gen: + description: "Image generation (limited — gpt-image not available via chat API)" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: high + + critical-ops: + description: "High-reliability operational tasks — infrastructure, orchestration, coordination" + candidates: + - provider: openai + model: gpt-?.?-terra + config: + reasoning_effort: xhigh diff --git a/tests/matrix_validation_rules.yaml b/tests/matrix_validation_rules.yaml index 40804ff..5100d40 100644 --- a/tests/matrix_validation_rules.yaml +++ b/tests/matrix_validation_rules.yaml @@ -9,6 +9,15 @@ # The test FAILS on providers missing from this map — that failure is the # prompt to look up the provider's valid set and record it here. # - Provider gains/loses effort levels? Edit its list below. +# - A specific MODEL (not just its provider) has a narrower or different +# allowed set than the provider-wide superset above? Add an entry under +# `reasoning_effort_values_by_model`, keyed by the exact model id as it +# appears in a candidate's `model:` field. This only applies to EXACT +# (non-glob) model ids -- a glob candidate is checked only against the +# provider-wide set, since it can resolve to more than one concrete +# model. When a model has a per-model entry here, it is the +# authoritative check for that model (narrower than, and enforced in +# addition to, the provider-wide `reasoning_effort_values` check). # # NOTE: this file must live OUTSIDE routing/ — any *.yaml in routing/ is # loaded as a matrix by the routing hook. @@ -52,3 +61,25 @@ reasoning_effort_values: openai: [none, minimal, low, medium, high, xhigh, max] gemini: [] # provider reads no effort key from mount config -- see note above github-copilot: [none, low, medium, high, xhigh, max] + +# Per-MODEL valid values for `reasoning_effort`, narrower than (and checked +# in addition to) the provider-wide superset above. Only exact model ids are +# checked here -- see HOW TO EXTEND. +# +# Sources: +# gpt-6-astra: OpenAI's gpt-6-astra model page (developers. +# openai.com), per the amplifier-support#524 +# routing study: "low|medium|high|xhigh|max... +# rejects `none`." Confirmed independently in +# microsoft/amplifier-module-provider-openai's +# `_GPT_6_ALLOWED_EFFORTS` (companion, +# unmerged PR -- see routing/openai-gpt6-canary.yaml +# DEPENDENCY note), both verified 2026-09-24. +# gpt-6-sol / gpt-6-luna: Same study, same model pages: "Sol/Luna accept +# `none|low|medium|high|xhigh|max` (default +# `medium`)." Also mirrored in the same +# provider-repo constant. +reasoning_effort_values_by_model: + gpt-6-astra: [low, medium, high, xhigh, max] + gpt-6-sol: [none, low, medium, high, xhigh, max] + gpt-6-luna: [none, low, medium, high, xhigh, max] diff --git a/tests/test_default_resolution_unchanged.py b/tests/test_default_resolution_unchanged.py index 40ea464..55bd74d 100644 --- a/tests/test_default_resolution_unchanged.py +++ b/tests/test_default_resolution_unchanged.py @@ -104,7 +104,17 @@ # `test_preset_bearing_matrix_is_stock_without_a_caller` for the same # invariant asserted directly. `openai.yaml` therefore stays recorded and # stays checked, preset or no preset. -EXCLUDED_FROM_RECORDING = {"openai.yaml"} +# +# `openai-gpt6-canary.yaml` (added 2026-09-24 for +# amplifier-support#524) is excluded for a different, simpler reason: it +# is a brand-new opt-in matrix that never existed at the pre-feature +# commit this fixture was recorded from, so there is no historical +# "before" resolution for it to be byte-identical to. Its own cold-vs- +# preset invariant (the same one `test_preset_bearing_matrix_is_stock_ +# without_a_caller` checks for `openai.yaml`) is asserted independently, +# without depending on this golden fixture, in +# `tests/test_gpt6_canary_routing.py`. +EXCLUDED_FROM_RECORDING = {"openai.yaml", "openai-gpt6-canary.yaml"} # Every matrix present in the recording as it stands. A frozen manifest, so a # matrix vanishing from the fixture -- by a bad `--regenerate`, a bad merge, a diff --git a/tests/test_gpt6_canary_routing.py b/tests/test_gpt6_canary_routing.py new file mode 100644 index 0000000..5365f2d --- /dev/null +++ b/tests/test_gpt6_canary_routing.py @@ -0,0 +1,509 @@ +"""Focused tests for the opt-in `routing/openai-gpt6-canary.yaml` matrix. + +Companion to `microsoft-amplifier/amplifier-support#524`. This matrix is +intentionally NOT covered by `tests/test_default_resolution_unchanged.py`'s +golden-fixture byte-identity check (it is new; there is no pre-feature +recording for a file that did not exist yet -- see that test's +`EXCLUDED_FROM_RECORDING` comment). Everything that check would normally give +us for free is instead asserted here, directly, against this one file: + +1. Structural shape: opt-in only, 4-rung strict ladder, exact GPT-6 ids. +2. Per-model `reasoning_effort` validation (astra vs. sol/luna allowed sets + differ). The actual per-model rule DATA lives in the shared, repository- + level `tests/matrix_validation_rules.yaml` (`reasoning_effort_values_by_model`) + and is enforced for every shipped matrix by + `tests/test_matrix_config_validation.py`; this file re-runs that same + check scoped to just this one matrix, so a failure here points straight + at the canary rather than at the parametrized sweep over every matrix. +3. No-catalog/no-capability-preflight behaviour for exact GPT-6 ids: they + resolve the same way whether the provider's model catalog is missing, + errors, or lists only near-miss strings that are NOT the clean id -- + because, per the matrix file's own "NO CATALOG/CAPABILITY PREFLIGHT" + note, an exact (non-glob) candidate never calls `list_models()` at all. + Near-miss strings in a catalog must never be selected in place of an + exact id, and must never be picked up by this matrix's GLOB candidates + either. +4. The delegation-preset cold-resolution invariant: with no caller context, + resolution is identical whether or not the preset is attached (the same + invariant `test_preset_bearing_matrix_is_stock_without_a_caller` checks + for `openai.yaml`, reproduced here without touching the golden fixture). +5. The escalation-clamp behaviour the whole file exists to exercise: a + `reasoning` sub-delegation from a caller below the `astra` rung is + clamped down to the matrix's own second (`gpt-5.6-terra`) candidate; a + caller already AT the `astra` rung is not. This clamp path is the ONLY + way that second candidate is ever reached -- see the matrix file's "NO + CATALOG/CAPABILITY PREFLIGHT" note: it is not a capability/outage + fallback, and cold resolution (no caller context) never reaches it. +""" + +from __future__ import annotations + +import asyncio +import sys +from functools import lru_cache +from pathlib import Path +from typing import Any +from unittest.mock import AsyncMock, MagicMock + +import pytest +import yaml + +REPO_ROOT = Path(__file__).resolve().parent.parent +ROUTING_DIR = REPO_ROOT / "routing" +MATRIX_PATH = ROUTING_DIR / "openai-gpt6-canary.yaml" +RULES_PATH = Path(__file__).parent / "matrix_validation_rules.yaml" + +sys.path.insert(0, str(REPO_ROOT / "modules" / "hooks-routing")) + +from amplifier_module_hooks_routing.knob_consistency import ( # noqa: E402 + CANONICAL_EFFORT_KEY, + CallerContext, + EscalationState, + parse_preset, +) +from amplifier_module_hooks_routing.resolver import ( # noqa: E402 + _is_glob, + resolve_model_role, +) + +# --------------------------------------------------------------------------- +# Per-model effort validation data -- READ from the repository-level rules +# file, not duplicated here. `reasoning_effort_values_by_model` in +# tests/matrix_validation_rules.yaml is the single source of truth (also +# enforced, for every shipped matrix, by +# tests/test_matrix_config_validation.py::test_reasoning_effort_values_valid_per_model). +# That file's own header documents the source (OpenAI's first-party model +# pages, per the amplifier-support#524 routing study, verified 2026-09-24). +# --------------------------------------------------------------------------- +_RULES = yaml.safe_load(RULES_PATH.read_text(encoding="utf-8")) +GPT_6_ALLOWED_EFFORTS: dict[str, frozenset[str]] = { + model: frozenset(values) + for model, values in _RULES["reasoning_effort_values_by_model"].items() +} + + +@lru_cache(maxsize=1) +def _load_matrix() -> dict[str, Any]: + """The file is static for the whole test run -- read and parse it once. + + No test in this module mutates the returned dict, so callers safely + share the cached object. + """ + return yaml.safe_load(MATRIX_PATH.read_text(encoding="utf-8")) + + +def _roles() -> dict[str, Any]: + return _load_matrix()["roles"] + + +def _gpt6_candidates() -> list[tuple[str, int, dict[str, Any]]]: + """Every (role, index, candidate) whose model is an exact GPT-6 id.""" + out = [] + for role, role_def in _roles().items(): + for i, cand in enumerate(role_def["candidates"]): + if cand.get("model") in GPT_6_ALLOWED_EFFORTS: + out.append((role, i, cand)) + return out + + +def _providers(models: dict[str, list[str]]) -> dict[str, Any]: + providers: dict[str, Any] = {} + for name, model_list in models.items(): + provider = MagicMock() + provider.list_models = AsyncMock(return_value=list(model_list)) + providers[name] = provider + return providers + + +def _providers_with_erroring_catalog(names: list[str]) -> dict[str, Any]: + """Providers whose `list_models()` raises. Used to prove exact-id + candidates never call it (so it can never be the reason they fail).""" + providers: dict[str, Any] = {} + for name in names: + provider = MagicMock() + provider.list_models = AsyncMock( + side_effect=RuntimeError(f"{name}: simulated catalog fetch failure") + ) + providers[name] = provider + return providers + + +# A roster that mounts every model this matrix can select, so a cold +# resolution always finds its top candidate rather than falling through. +FULL_OPENAI_ROSTER = [ + "gpt-6-astra", + "gpt-6-sol", + "gpt-6-luna", + "gpt-5.6-sol", + "gpt-5.6-terra", + "gpt-5.6-luna", +] + +# A catalog that lists near-miss strings for every GPT-6 id -- close enough +# to look like a stale/wrong entry -- but never the three clean ids +# themselves, and never any 5.6-family id either. Used to prove (a) exact +# GPT-6 candidates resolve to the clean id regardless (no catalog dependency +# at all) and (b) this matrix's glob candidates (`gpt-?.?-terra`, +# `gpt-?.?-luna`) never accidentally match a GPT-6-shaped near-miss. +NEAR_MISS_ONLY_ROSTER = [ + "gpt-6-astra-preview", + "gpt-6-astra-fast", + "gpt-60-astra", + "gpt-6-sol-fast", + "gpt-6-sol-mini", + "gpt-6-luna-fast", + "gpt-6-luna-mini", +] + +# A roster that has NO models at all for the provider -- distinct from a +# catalog error: this is what an empty-but-successful list_models() call +# returns, per openai.yaml's own documented "if that call fails, the glob +# resolves to nothing" behaviour for GLOB candidates. Exact ids must not +# share that fate. +EMPTY_ROSTER: list[str] = [] + + +# --------------------------------------------------------------------------- +# 1. Structural shape +# --------------------------------------------------------------------------- + + +def test_matrix_file_exists_and_parses() -> None: + assert MATRIX_PATH.exists(), f"missing {MATRIX_PATH}" + data = _load_matrix() + assert data is not None + assert "roles" in data + + +def test_is_not_the_default_matrix() -> None: + """This canary must never become anyone's default silently.""" + behaviour = yaml.safe_load( + (REPO_ROOT / "behaviors" / "routing.yaml").read_text(encoding="utf-8") + ) + declared = [ + hook.get("config", {}).get("default_matrix") + for hook in behaviour.get("hooks", []) + if hook.get("module") == "hooks-routing" + ] + assert "openai-gpt6-canary" not in declared, ( + "openai-gpt6-canary must stay opt-in; behaviors/routing.yaml must " + "not default to it" + ) + + +def test_ships_a_preset_with_strict_inherit_and_four_rungs() -> None: + data = _load_matrix() + assert "preset" in data, "openai-gpt6-canary.yaml must carry a preset: block" + parsed = parse_preset(data) + assert parsed is not None + assert parsed.inherit == "strict" + rungs = parsed.tier_ladder["openai"] + assert len(rungs) == 4, f"expected a 4-rung ladder, got {len(rungs)}: {rungs}" + # Astra sits ALONE on the top rung -- the deviation from the study's + # guidance the matrix file documents under "WHY A FOURTH RUNG". + assert rungs[-1] == ["gpt-6-astra"] + assert "gpt-6-sol" in rungs[2] + assert "gpt-6-luna" in rungs[0] + + +def test_gpt6_model_ids_are_exact_no_globs() -> None: + """No GPT-6 candidate may use a glob -- see the file header: the family + has no numeric sub-version and no dated snapshots to glob over.""" + for role, i, cand in [ + (role, i, cand) + for role, role_def in _roles().items() + for i, cand in enumerate(role_def["candidates"]) + ]: + model = cand.get("model", "") + if model.startswith("gpt-6"): + assert model in GPT_6_ALLOWED_EFFORTS, ( + f"{role}[{i}]: {model!r} is a gpt-6* id but not one of the " + f"three known exact ids {sorted(GPT_6_ALLOWED_EFFORTS)}" + ) + assert not _is_glob(model), ( + f"{role}[{i}]: GPT-6 candidate {model!r} must be an exact id, " + "not a glob" + ) + + +def test_canary_scope_is_exactly_fast_coding_reasoning() -> None: + """Only these three roles may select a GPT-6 model; every other role + stays on the pre-existing gpt-5.6-terra/-luna candidates.""" + roles = _roles() + gpt6_roles = { + role + for role, role_def in roles.items() + if any(c.get("model") in GPT_6_ALLOWED_EFFORTS for c in role_def["candidates"]) + } + assert gpt6_roles == {"fast", "coding", "reasoning"} + + +def test_fast_and_coding_have_exactly_one_candidate_reasoning_has_two() -> None: + """The documented asymmetry: `reasoning` keeps the pre-existing + gpt-5.6-terra candidate as an explicit second candidate (reachable only + via the strict clamp -- see module docstring point 5, and + TestEscalationClamp below); `fast` and `coding` do not have a second + candidate of any kind, because one would be equally unreachable by + ordinary resolution and would misleadingly read as a fallback that does + not exist -- see the matrix file's "NO CATALOG/CAPABILITY PREFLIGHT" + note.""" + roles = _roles() + assert len(roles["fast"]["candidates"]) == 1 + assert roles["fast"]["candidates"][0]["model"] == "gpt-6-luna" + + assert len(roles["coding"]["candidates"]) == 1 + assert roles["coding"]["candidates"][0]["model"] == "gpt-6-sol" + + reasoning_models = [c["model"] for c in roles["reasoning"]["candidates"]] + assert reasoning_models[0] == "gpt-6-astra" + assert "gpt-?.?-terra" in reasoning_models[1:] + + +# --------------------------------------------------------------------------- +# 2. Per-model effort validation +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "role,index,candidate", + _gpt6_candidates(), + ids=lambda v: v if isinstance(v, str) else "", +) +def test_gpt6_candidates_use_a_model_valid_effort( + role: str, index: int, candidate: dict[str, Any] +) -> None: + model = candidate["model"] + effort = (candidate.get("config") or {}).get(CANONICAL_EFFORT_KEY) + allowed = GPT_6_ALLOWED_EFFORTS[model] + assert effort is not None, f"{role}[{index}] ({model}): no {CANONICAL_EFFORT_KEY} set" + assert effort in allowed, ( + f"{role}[{index}]: reasoning_effort={effort!r} is not valid for " + f"{model!r} (allowed: {sorted(allowed)})" + ) + + +def test_astra_candidate_does_not_use_none() -> None: + """Regression guard named directly: `none` is valid for sol/luna but NOT + for astra -- a future edit that moves `reasoning` to `none` on astra must + fail here, not go silently inert on the wire.""" + roles = _roles() + astra_candidate = roles["reasoning"]["candidates"][0] + assert astra_candidate["model"] == "gpt-6-astra" + assert astra_candidate["config"][CANONICAL_EFFORT_KEY] != "none" + + +# --------------------------------------------------------------------------- +# 3. No catalog / no capability preflight -- missing, erroring, and +# near-miss-only catalogs all resolve exact GPT-6 ids identically. +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "role,expected_model", + [("fast", "gpt-6-luna"), ("coding", "gpt-6-sol"), ("reasoning", "gpt-6-astra")], +) +def test_exact_gpt6_ids_resolve_with_an_empty_catalog( + role: str, expected_model: str +) -> None: + """`list_models()` returning an empty (but successful) list must not + affect an exact-id candidate -- unlike a glob, which would resolve to + nothing (see openai.yaml's own documented degrade path for globs).""" + roles = _roles() + providers = _providers({"openai": EMPTY_ROSTER}) + result = asyncio.run(resolve_model_role([role], roles, providers)) + assert result, f"role {role!r} resolved to nothing with an empty catalog" + assert result[0]["model"] == expected_model + + +@pytest.mark.parametrize( + "role,expected_model", + [("fast", "gpt-6-luna"), ("coding", "gpt-6-sol"), ("reasoning", "gpt-6-astra")], +) +def test_exact_gpt6_ids_resolve_when_list_models_raises( + role: str, expected_model: str +) -> None: + """`list_models()` raising must not affect an exact-id candidate. This + is the direct proof of the matrix file's central claim: `list_models()` + is never even CALLED for a non-glob candidate, so its failure mode is + structurally unreachable for `fast`/`coding`/`reasoning`'s top + candidate. If this ever starts calling `list_models()` for an exact id, + this test fails with the injected RuntimeError instead of silently + passing.""" + roles = _roles() + providers = _providers_with_erroring_catalog(["openai"]) + result = asyncio.run(resolve_model_role([role], roles, providers)) + assert result, f"role {role!r} resolved to nothing when list_models() raised" + assert result[0]["model"] == expected_model + providers["openai"].list_models.assert_not_called() + + +@pytest.mark.parametrize( + "role,expected_model", + [("fast", "gpt-6-luna"), ("coding", "gpt-6-sol"), ("reasoning", "gpt-6-astra")], +) +def test_exact_gpt6_ids_resolve_to_the_clean_id_with_a_near_miss_only_catalog( + role: str, expected_model: str +) -> None: + """A catalog full of near-miss strings for every GPT-6 id (`-preview`, + `-fast`, `-mini`, a digit near-miss) but never the clean id itself must + not change the result: the exact candidate resolves to the CLEAN id it + names, never to a near-miss, because no matching against the catalog + happens for it at all.""" + roles = _roles() + providers = _providers({"openai": NEAR_MISS_ONLY_ROSTER}) + result = asyncio.run(resolve_model_role([role], roles, providers)) + assert result, f"role {role!r} resolved to nothing with a near-miss-only catalog" + assert result[0]["model"] == expected_model + assert result[0]["model"] not in NEAR_MISS_ONLY_ROSTER + + +def test_glob_candidates_never_match_a_gpt6_near_miss() -> None: + """The OTHER half of the near-miss guarantee: this matrix's glob + candidates (`gpt-?.?-terra`, `gpt-?.?-luna`, used by every non-canaried + role) must not accidentally select a GPT-6-shaped near-miss string if + one happens to be present in the catalog alongside legitimate 5.6-family + ids. `gpt-?.?-terra`/`gpt-?.?-luna` require a literal '.' after a single + version-digit character; none of the GPT-6 near-miss ids have that + shape, so none should ever satisfy the glob.""" + roles = _roles() + mixed_roster = NEAR_MISS_ONLY_ROSTER + ["gpt-5.6-terra", "gpt-5.6-luna"] + providers = _providers({"openai": mixed_roster}) + + async def _run() -> dict[str, Any]: + return { + role: await resolve_model_role([role], roles, providers) + for role in ("general", "ui-coding", "critique") + } + + resolved = asyncio.run(_run()) + for role, result in resolved.items(): + assert result, f"role {role!r} resolved to nothing" + assert result[0]["model"] in ("gpt-5.6-terra", "gpt-5.6-luna"), ( + f"role {role!r} resolved to {result[0]['model']!r}, " + "which looks like it matched a GPT-6 near-miss string" + ) + assert result[0]["model"] not in NEAR_MISS_ONLY_ROSTER + + +def test_reasoning_candidates_never_call_list_models_cold() -> None: + """Cold resolution (no caller context) of `reasoning` must not touch + `list_models()` at all: BOTH its candidates -- `gpt-6-astra` (the one + that always wins cold) and `gpt-5.6-terra` (never reached cold, see + module docstring point 5) -- are either an exact id or never evaluated, + so the erroring mock must never be invoked.""" + roles = _roles() + providers = _providers_with_erroring_catalog(["openai"]) + result = asyncio.run(resolve_model_role(["reasoning"], roles, providers)) + assert result and result[0]["model"] == "gpt-6-astra" + providers["openai"].list_models.assert_not_called() + + +# --------------------------------------------------------------------------- +# 4. Cold-resolution / preset-is-opt-in-twice invariant +# --------------------------------------------------------------------------- + + +def test_cold_resolution_is_unchanged_by_the_preset() -> None: + """No caller context: resolving with the parsed preset attached must + equal resolving with no preset at all. Reproduces, for this one new + matrix, the same invariant + test_preset_bearing_matrix_is_stock_without_a_caller checks for + openai.yaml -- without depending on that test's golden fixture, since + this matrix was never in the pre-feature recording.""" + data = _load_matrix() + roles = data["roles"] + parsed = parse_preset(data) + assert parsed is not None + + async def _resolve_all(preset: Any) -> dict[str, Any]: + return { + role: await resolve_model_role( + [role], + roles, + _providers({"openai": FULL_OPENAI_ROSTER}), + preset=preset, + caller_context=None, + ) + for role in roles + } + + without = asyncio.run(_resolve_all(None)) + with_preset = asyncio.run(_resolve_all(parsed)) + assert with_preset == without + + +def test_every_non_exempt_role_resolves_with_the_full_roster() -> None: + """Sanity: with every GPT-6/5.6 model mounted, no role silently yields + nothing (mirrors tests/test_single_provider_coverage.py's intent for + this one opt-in matrix, which that file does not cover since it is + parametrized over the shipped matrix names, not a glob).""" + roles = _roles() + providers = _providers({"openai": FULL_OPENAI_ROSTER}) + + async def _run() -> dict[str, Any]: + return { + role: await resolve_model_role([role], roles, providers) for role in roles + } + + resolved = asyncio.run(_run()) + unresolved = {r for r, got in resolved.items() if not got} + assert not unresolved, f"roles with no resolution: {sorted(unresolved)}" + + +# --------------------------------------------------------------------------- +# 5. Escalation clamp -- the point of the exercise +# --------------------------------------------------------------------------- + +TERRA_CALLER = CallerContext( + family="openai", model="gpt-5.6-terra", effort="medium", provider_key="terra" +) +SOL_CALLER = CallerContext( + family="openai", model="gpt-5.6-sol", effort="high", provider_key="sol" +) +ASTRA_CALLER = CallerContext( + family="openai", model="gpt-6-astra", effort="high", provider_key="astra" +) + + +def _resolve_reasoning(caller: CallerContext | None) -> list[dict[str, Any]]: + data = _load_matrix() + roles = data["roles"] + preset = parse_preset(data) + providers = _providers({"openai": FULL_OPENAI_ROSTER}) + return asyncio.run( + resolve_model_role( + ["reasoning"], + roles, + providers, + caller_context=caller, + preset=preset, + escalations=EscalationState(), + ) + ) + + +def test_reasoning_clamps_to_terra_for_a_terra_tier_caller() -> None: + """A caller resolved to gpt-5.6-terra (rung 2 of 4) must not be able to + delegate a sub-agent onto gpt-6-astra (rung 4) -- strict inherit denies + the escalation and the sub-agent lands on the matrix's own second + (`gpt-5.6-terra`) candidate instead. This is the one and only path that + reaches that second candidate at all -- see module docstring point 5.""" + result = _resolve_reasoning(TERRA_CALLER) + assert result + assert result[0]["model"] == "gpt-5.6-terra" + + +def test_reasoning_clamps_to_sol_rung_for_a_sol_tier_caller() -> None: + """A caller at rung 3 (sol) is still below astra's rung 4 and must also + be denied the escalation.""" + result = _resolve_reasoning(SOL_CALLER) + assert result + assert result[0]["model"] != "gpt-6-astra" + + +def test_reasoning_lands_on_astra_for_an_astra_tier_caller() -> None: + """A caller already AT the top rung is not a clamp target: `reasoning` + resolves to its own top candidate, gpt-6-astra, unchanged.""" + result = _resolve_reasoning(ASTRA_CALLER) + assert result + assert result[0]["model"] == "gpt-6-astra" diff --git a/tests/test_matrix_config_validation.py b/tests/test_matrix_config_validation.py index 711b68f..f54909a 100644 --- a/tests/test_matrix_config_validation.py +++ b/tests/test_matrix_config_validation.py @@ -12,6 +12,7 @@ tests/matrix_validation_rules.yaml -- extend them there, not here. """ +import sys import yaml import pytest from pathlib import Path @@ -20,9 +21,13 @@ ROUTING_DIR = REPO_ROOT / "routing" RULES_PATH = Path(__file__).parent / "matrix_validation_rules.yaml" +sys.path.insert(0, str(REPO_ROOT / "modules" / "hooks-routing")) +from amplifier_module_hooks_routing.resolver import _is_glob # noqa: E402 + RULES = yaml.safe_load(RULES_PATH.read_text()) CANONICAL_EFFORT_KEY = RULES["canonical_effort_key"] EFFORT_VALUES_BY_PROVIDER = RULES["reasoning_effort_values"] +EFFORT_VALUES_BY_MODEL = RULES.get("reasoning_effort_values_by_model", {}) MATRIX_FILES = sorted(ROUTING_DIR.glob("*.yaml")) @@ -94,3 +99,37 @@ def test_reasoning_effort_values_valid_per_provider(self, matrix_path): f"{matrix_path.name}: invalid reasoning_effort values:\n " + "\n ".join(violations) ) + + def test_reasoning_effort_values_valid_per_model(self, matrix_path): + """(c) A narrower, per-MODEL allowed set overrides the provider-wide + superset for any EXACT (non-glob) model id listed in + tests/matrix_validation_rules.yaml's `reasoning_effort_values_by_model`. + + The provider-wide check above cannot catch this class: e.g. + `gpt-6-astra` and `gpt-6-sol` are both `provider: openai` and both + pass the provider-wide superset for `none`, but `gpt-6-astra` + rejects `none` at the provider while `gpt-6-sol` accepts it. Only a + model-keyed table can tell them apart. Glob candidates are + deliberately excluded -- a glob can resolve to more than one + concrete model, so only an exact id is checked here. + """ + violations = [] + for role, i, cand in iter_candidates(matrix_path): + model = cand.get("model") + if not isinstance(model, str) or _is_glob(model): + continue # glob candidate -- not checked per-model + allowed = EFFORT_VALUES_BY_MODEL.get(model) + if allowed is None: + continue # no per-model rule for this id -- provider-wide check covers it + value = (cand.get("config") or {}).get(CANONICAL_EFFORT_KEY) + if value is not None and value not in allowed: + violations.append( + f"{role}[{i}] ({cand.get('provider')}/{model}): " + f"reasoning_effort={value!r} not in the model-specific " + f"allowed set {allowed} for {model!r} " + f"(reasoning_effort_values_by_model in {RULES_PATH.name})" + ) + assert not violations, ( + f"{matrix_path.name}: invalid per-model reasoning_effort values:\n " + + "\n ".join(violations) + ) diff --git a/tests/test_single_provider_coverage.py b/tests/test_single_provider_coverage.py index dea8420..5887bd0 100644 --- a/tests/test_single_provider_coverage.py +++ b/tests/test_single_provider_coverage.py @@ -275,3 +275,91 @@ def test_no_matrix_names_the_chatgpt_backend_directly() -> None: "`provider: openai` already reaches that backend via " "PROVIDER_FAMILY_ALIASES -- one candidate serves both bills." ) + + +# --------------------------------------------------------------------------- +# openai-gpt6-canary.yaml: single-provider coverage, including a roster with +# GPT-6 near-miss ids alongside the real ones. +# +# The opt-in canary is not swept by `test_provider_alone_routes_every_role_of_ +# the_default_matrix` above (that test is parametrized on `DEFAULT_MATRIX` +# only, i.e. `balanced`). It gets its own equivalent here: mount ONLY +# `openai`, with a roster carrying every GPT-6 near-miss id a live catalog +# could plausibly contain (a `-preview`, a `-fast`, a `-mini`, a digit +# near-miss) ALONGSIDE the real gpt-5.x/-6 ids, and require every role to +# still resolve to the matrix's own declared candidate -- never to a +# near-miss, and never to nothing. +# --------------------------------------------------------------------------- + +GPT6_CANARY_MATRIX = "openai-gpt6-canary" + +# A roster containing every model this matrix's roles can select, PLUS +# near-miss ids for each GPT-6 model that a stale or wrong catalog entry +# might plausibly contain. None of the near-miss strings are valid GPT-6 +# ids; a correct resolution never selects one. +OPENAI_ROSTER_WITH_GPT6_NEAR_MISSES: list[str] = [ + # Real ids this matrix's roles can select. + "gpt-6-astra", + "gpt-6-sol", + "gpt-6-luna", + "gpt-5.6-terra", + "gpt-5.6-luna", + # Near-miss ids -- must never be selected in place of the real id above. + "gpt-6-astra-preview", + "gpt-6-astra-fast", + "gpt-60-astra", + "gpt-6-sol-fast", + "gpt-6-sol-mini", + "gpt-6-luna-fast", + "gpt-6-luna-mini", +] + + +def test_gpt6_canary_matrix_resolves_every_role_with_openai_alone() -> None: + """With ONLY `openai` mounted and a roster that also contains GPT-6 + near-miss ids, every role of `openai-gpt6-canary.yaml` must still + resolve -- and never to one of the near-miss ids.""" + roles = _roles(GPT6_CANARY_MATRIX) + providers = _providers("openai") + providers["openai"].list_models = AsyncMock( + return_value=list(OPENAI_ROSTER_WITH_GPT6_NEAR_MISSES) + ) + + async def _run() -> dict[str, Any]: + return { + role: await resolve_model_role([role], roles, providers) for role in roles + } + + resolved = asyncio.run(_run()) + unresolved = {r for r, got in resolved.items() if not got} + assert not unresolved, ( + f"with ONLY openai configured (roster includes GPT-6 near-misses), " + f"{GPT6_CANARY_MATRIX}.yaml leaves these roles unrouted: " + f"{sorted(unresolved)}" + ) + + near_misses = {m for m in OPENAI_ROSTER_WITH_GPT6_NEAR_MISSES if m not in ( + "gpt-6-astra", "gpt-6-sol", "gpt-6-luna", "gpt-5.6-terra", "gpt-5.6-luna", + )} + for role, got in resolved.items(): + assert got[0]["model"] not in near_misses, ( + f"role {role!r} resolved to near-miss id {got[0]['model']!r}" + ) + + +@pytest.mark.parametrize( + "role,expected_model", + [("fast", "gpt-6-luna"), ("coding", "gpt-6-sol"), ("reasoning", "gpt-6-astra")], +) +def test_gpt6_canary_canaried_roles_land_on_the_exact_clean_id( + role: str, expected_model: str +) -> None: + """The three canaried roles specifically: with the near-miss roster + mounted, each must resolve to its own declared clean id, exactly.""" + roles = _roles(GPT6_CANARY_MATRIX) + providers = _providers("openai") + providers["openai"].list_models = AsyncMock( + return_value=list(OPENAI_ROSTER_WITH_GPT6_NEAR_MISSES) + ) + result = asyncio.run(resolve_model_role([role], roles, providers)) + assert result and result[0]["model"] == expected_model