feat(routing): add opt-in GPT-6 canary matrix - #78
Ken Chau (kenotron-ms) wants to merge 1 commit into
Conversation
Add routing/openai-gpt6-canary.yaml, an opt-in Stage-1 canary matrix for the GPT-6 family (microsoft-amplifier/amplifier-support#524): - fast -> gpt-6-luna @ low - coding -> gpt-6-sol @ medium - reasoning -> gpt-6-astra @ high (gpt-5.6-terra @ xhigh as the strict tier-clamp landing candidate) All other roles are unchanged from openai.yaml. Exact model IDs only; no GPT-6 creative/image-gen. Uses a 4-rung strict tier ladder (luna < terra < sol < astra). Existing matrices, defaults, and golden resolution behavior are unchanged; the matrix is inert until selected. Depends on microsoft/amplifier-module-provider-openai#112 and must merge after it. Generated with Amplifier Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
|
Closing at product direction: dedicated GPT-6 routing is not yet proven and is not wanted at this stage. We should first validate GPT-6 through wildcard/model-pattern handling in the existing OpenAI matrix after provider support lands, preserving current roles, fallbacks, and defaults. No replacement routing PR is planned now. |
|
Fresh linked-stack review at |
What changed
Adds
routing/openai-gpt6-canary.yaml, an opt-in Stage-1 canary matrix for the OpenAI GPT-6 family, along with its tests and docs. Tracks microsoft-amplifier/amplifier-support#524 and implements the routing-study comment on that issue.Important
Depends on microsoft/amplifier-module-provider-openai#112 (GPT-6 Sol/Luna provider support) and must merge after it. The matrix names exact GPT-6 IDs that only #112 teaches the provider to serialize and validate correctly. Do not merge this PR, or tell users to select the matrix, until #112 has merged.
Stage-1 canary scope (exact)
fastgpt-6-luna@reasoning_effort: lowcodinggpt-6-sol@reasoning_effort: mediumreasoninggpt-6-astra@reasoning_effort: high, thengpt-5.6-terra@xhighas the strict tier-clamp landing candidateopenai.yamlgpt-6-luna,gpt-6-sol,gpt-6-astra). No globs, aliases or snapshots.creativeorimage-gen.behaviors/and resolver/hook source are untouched, and the default staysbalanced. The golden replay intests/test_default_resolution_unchanged.pyexcludes the new file by name (it did not exist at the recorded commit) rather than re-recording.amplifier routing use openai-gpt6-canary(orrouting.matrix: openai-gpt6-canaryinsettings.yaml).amplifier routing use openai(or whatever matrix you were on before), or remove therouting.matrixoverride.Four-rung strict ladder
The
preset:block usesinherit: strictwith a four-rungtier_ladder:luna<terra<sol<astra. Astra sits alone on the top rung. The study says not to guess GPT-6 into the existing ladder "without policy"; this ladder is that policy. It is this rollout's own design choice, not a requirement of the study, and the rationale is documented inline under "WHY A FOURTH RUNG". Strict clamping means a delegating caller can never be routed above its own rung. A candidate above the caller's rung is clamped down to the highest candidate at or below it, which is whyreasoningcarriesgpt-5.6-terraas a landing candidate.Exact-ID semantics: no catalog check, no runtime failover
list_models()or any capability table. They resolve as soon as anyopenai-family provider is mounted, whether or not that provider's code actually implements the ID. That is why #112 must land first.fastandcodinghave no routing-level fallback. The Terra candidate onreasoningis reached only through the strict clamp path, never as an outage or capability fallback.fix(routing): honor mounted account availability and scoped catalogs, already onmainat8541929, which this branch is rebased on) filters on whether a provider/account is mounted and on its scoped catalog. It does not check whether an exact model ID is authorized or served. An unauthorized or unavailable exact GPT-6 ID still resolves and then fails at request time.Why
Issue #524 asks for GPT-6 support plus the routing integration the study identified. The study recommends a canary placement rather than a global default change. Shipping the canary as a separate, inert-until-selected matrix makes it reviewable, testable and linkable, and it touches no existing user. A settings override on
openaicannot extend the ladder, becausepreset:is read only from the base matrix, so strict clamping would silently drop GPT-6.Verification
Offline / synthetic (deterministic fixtures and mocked providers, no network):
tests/: 89 passed on each of Python 3.11, 3.12 and 3.13.modules/hooks-routingtests: 648 passed.ruff: clean.tests/test_gpt6_canary_routing.pycovers the file being opt-in only and non-default, exact IDs only, no GPT-6 on creative/image-gen, non-canary roles identical toopenai.yaml, per-model effort validity (matrix_validation_rules.yaml→reasoning_effort_values_by_model), ladder/rung matching, and strict clamp outcomes. It also extends the existing single-provider coverage, config validation, knob-consistency and golden-unchanged guards.LOCAL DTU (live, bounded):
amplifier-digital-twinon Incus with current upstream Amplifiermain, with the #112 provider worktree loaded viasource: file://….fast/coding/reasoning. The live calls returnedFAST-OK,CODING-OKandREASONING-OK.gpt-5.6-lunaforcodingandreasoning, as the ladder requires.Compatibility, rollout, risks and residuals
openai-chatgpt) is unverified. GPT-6 availability was confirmed only on the API-key backend.ContentFilterError(handled in #112), and this matrix deliberately provides no fallback that would bypass it.openai.yamlor any default is out of scope and needs separate evidence.