Skip to content

feat(routing): add opt-in GPT-6 canary matrix - #78

Closed
Ken Chau (kenotron-ms) wants to merge 1 commit into
mainfrom
feat/gpt-6-canary-routing
Closed

Ken Chau (kenotron-ms) wants to merge 1 commit into
mainfrom
feat/gpt-6-canary-routing

Conversation

@kenotron-ms

Copy link
Copy Markdown
Contributor

What changed

Adds routing/openai-gpt6-canary.yaml, an opt-in Stage-1 canary matrix for the OpenAI GPT-6 family, along with its tests and docs. Tracks microsoft-amplifier/amplifier-support#524 and implements the routing-study comment on that issue.

Important

Depends on microsoft/amplifier-module-provider-openai#112 (GPT-6 Sol/Luna provider support) and must merge after it. The matrix names exact GPT-6 IDs that only #112 teaches the provider to serialize and validate correctly. Do not merge this PR, or tell users to select the matrix, until #112 has merged.

Stage-1 canary scope (exact)

Role Candidate(s)
fast gpt-6-luna @ reasoning_effort: low
coding gpt-6-sol @ reasoning_effort: medium
reasoning gpt-6-astra @ reasoning_effort: high, then gpt-5.6-terra @ xhigh as the strict tier-clamp landing candidate
every other role Identical to openai.yaml
  • Uses exact model IDs only (gpt-6-luna, gpt-6-sol, gpt-6-astra). No globs, aliases or snapshots.
  • Puts no GPT-6 on creative or image-gen.
  • Leaves existing matrices, defaults and golden behavior unchanged. No shipped matrix is edited, behaviors/ and resolver/hook source are untouched, and the default stays balanced. The golden replay in tests/test_default_resolution_unchanged.py excludes the new file by name (it did not exist at the recorded commit) rather than re-recording.
  • Opt-in: amplifier routing use openai-gpt6-canary (or routing.matrix: openai-gpt6-canary in settings.yaml).
  • Rollback: amplifier routing use openai (or whatever matrix you were on before), or remove the routing.matrix override.

Four-rung strict ladder

The preset: block uses inherit: strict with a four-rung tier_ladder: luna < terra < sol < astra. Astra sits alone on the top rung. The study says not to guess GPT-6 into the existing ladder "without policy"; this ladder is that policy. It is this rollout's own design choice, not a requirement of the study, and the rationale is documented inline under "WHY A FOURTH RUNG". Strict clamping means a delegating caller can never be routed above its own rung. A candidate above the caller's rung is clamped down to the highest candidate at or below it, which is why reasoning carries gpt-5.6-terra as a landing candidate.

Exact-ID semantics: no catalog check, no runtime failover

  • The resolver short-circuits exact (non-glob) IDs. They are never checked against list_models() or any capability table. They resolve as soon as any openai-family provider is mounted, whether or not that provider's code actually implements the ID. That is why #112 must land first.
  • Candidate order is selection-time priority, not runtime failover. If a GPT-6 request fails at call time, the next candidate is not retried. fast and coding have no routing-level fallback. The Terra candidate on reasoning is reached only through the strict clamp path, never as an outage or capability fallback.
  • Upstream Honor mounted provider account availability during routing #77 is a different mechanism. Honor mounted provider account availability during routing #77 (fix(routing): honor mounted account availability and scoped catalogs, already on main at 8541929, which this branch is rebased on) filters on whether a provider/account is mounted and on its scoped catalog. It does not check whether an exact model ID is authorized or served. An unauthorized or unavailable exact GPT-6 ID still resolves and then fails at request time.

Why

Issue #524 asks for GPT-6 support plus the routing integration the study identified. The study recommends a canary placement rather than a global default change. Shipping the canary as a separate, inert-until-selected matrix makes it reviewable, testable and linkable, and it touches no existing user. A settings override on openai cannot extend the ladder, because preset: is read only from the base matrix, so strict clamping would silently drop GPT-6.

Verification

Offline / synthetic (deterministic fixtures and mocked providers, no network):

  • Root tests/: 89 passed on each of Python 3.11, 3.12 and 3.13.
  • modules/hooks-routing tests: 648 passed.
  • ruff: clean.
  • Bundle structure check: 9 routing files, 11 parsed.
  • New tests/test_gpt6_canary_routing.py covers the file being opt-in only and non-default, exact IDs only, no GPT-6 on creative/image-gen, non-canary roles identical to openai.yaml, per-model effort validity (matrix_validation_rules.yaml → reasoning_effort_values_by_model), ladder/rung matching, and strict clamp outcomes. It also extends the existing single-provider coverage, config validation, knob-consistency and golden-unchanged guards.

LOCAL DTU (live, bounded):

  • Ran under amplifier-digital-twin on Incus with current upstream Amplifier main, with the #112 provider worktree loaded via source: file://….
  • Because this PR is unmerged, the matrix was loaded as a user override, and the matrix file hash matched between host and DTU.
  • Astra-root session: routing selected Luna / Sol / Astra for fast / coding / reasoning. The live calls returned FAST-OK, CODING-OK and REASONING-OK.
  • Luna-root session: the strict clamp selected gpt-5.6-luna for coding and reasoning, as the ladder requires.
  • The per-model evidence on the provider side (actual served model, usage and cost) is in #112 and is not repeated here.
  • Teardown was verified: the DTU was destroyed and no Incus containers remain. Unrelated containers were not touched.

Compatibility, rollout, risks and residuals

  • Breaking changes: none. The file is additive and inert until selected, and no global default changes.
  • ChatGPT backend (openai-chatgpt) is unverified. GPT-6 availability was confirmed only on the API-key backend.
  • Unauthorized or unavailable exact IDs fail at request time. Routing does not preflight them.
  • Routing candidate order is not runtime failover (see above).
  • A GPT-6 misalignment-policy 403 must not be routed around. It surfaces as a non-retryable ContentFilterError (handled in #112), and this matrix deliberately provides no fallback that would bypass it.
  • Stage-2 promotion is deferred: GPT-6 in openai.yaml or any default is out of scope and needs separate evidence.

Add routing/openai-gpt6-canary.yaml, an opt-in Stage-1 canary matrix
for the GPT-6 family (microsoft-amplifier/amplifier-support#524):

- fast      -> gpt-6-luna  @ low
- coding    -> gpt-6-sol   @ medium
- reasoning -> gpt-6-astra @ high (gpt-5.6-terra @ xhigh as the strict
  tier-clamp landing candidate)

All other roles are unchanged from openai.yaml. Exact model IDs only; no
GPT-6 creative/image-gen. Uses a 4-rung strict tier ladder
(luna < terra < sol < astra). Existing matrices, defaults, and golden
resolution behavior are unchanged; the matrix is inert until selected.

Depends on microsoft/amplifier-module-provider-openai#112 and must merge
after it.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@kenotron-ms

Copy link
Copy Markdown
Contributor Author

Closing at product direction: dedicated GPT-6 routing is not yet proven and is not wanted at this stage. We should first validate GPT-6 through wildcard/model-pattern handling in the existing OpenAI matrix after provider support lands, preserving current roles, fallbacks, and defaults. No replacement routing PR is planned now.

@kenotron-ms

Copy link
Copy Markdown
Contributor Author

Fresh linked-stack review at dfed3afa858d0b6f539d1003134ae1da8608e579: the canary remains correctly inert/opt-in and its checks are green, but this closed PR should not be reopened as-is. Its repeated merge-order gate names provider #112, which is now a stale/conflicting duplicate; GPT-6 Sol/Luna support reached provider main through #113 instead. If routing is revisited after product evidence, update the dependency to the actual landed provider implementation and revalidate account/backend availability and request-time failure behavior. Per the existing product-direction closure, no dedicated replacement matrix is approved here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants