Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,7 @@ jobs:
- name: Run root tests
run: >
uv run --no-project --python ${{ matrix.python-version }}
--with pytest --with pytest-asyncio --with pyyaml
--with pytest --with pytest-asyncio --with pyyaml --with click
python -m pytest tests -q --tb=short

# ---------------------------------------------------------------------------
Expand Down Expand Up @@ -137,6 +137,13 @@ jobs:
--with git+https://github.com/microsoft/amplifier-foundation
python -m pytest -q --tb=short

- name: Run catalog refresh runtime seam tests with real dependencies
working-directory: modules/${{ matrix.module }}
run: >
uv run --frozen --extra dev
--with git+https://github.com/microsoft/amplifier-foundation
python -c "import amplifier_core, amplifier_foundation.spawn_utils, pytest; raise SystemExit(pytest.main(['../../tests/test_catalog_refresh.py', '-q', '--tb=short']))"

# ---------------------------------------------------------------------------
# Bundle structure — a cheap YAML parse of bundle.md's frontmatter, every
# behaviors/*.yaml and every routing/*.yaml. No network, no keys, ~1 second.
Expand Down
36 changes: 36 additions & 0 deletions .github/workflows/model-catalogs.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
name: Live model catalogs

# No inference. Disabled until dedicated credentials and the enable variable
# are provisioned; missing credentials on a requested check fail, not skip.
on:
workflow_dispatch:
schedule:
- cron: "23 15 * * 1"

permissions:
contents: read

jobs:
catalogs:
if: github.event_name == 'workflow_dispatch' || vars.MODEL_CATALOG_CHECKS_ENABLED == 'true'
runs-on: ubuntu-latest
timeout-minutes: 10
strategy:
fail-fast: false
matrix:
provider: [openai, anthropic, gemini, github-copilot]
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- uses: astral-sh/setup-uv@ae62891fec2bb8e7d6c99fc78c9fec3a63790f8d # v10.0.0
- name: Check fresh provider catalog
env:
OPENAI_API_KEY: ${{ matrix.provider == 'openai' && secrets.OPENAI_API_KEY || '' }}
ANTHROPIC_API_KEY: ${{ matrix.provider == 'anthropic' && secrets.ANTHROPIC_API_KEY || '' }}
GOOGLE_API_KEY: ${{ matrix.provider == 'gemini' && secrets.GOOGLE_API_KEY || '' }}
COPILOT_GITHUB_TOKEN: ${{ matrix.provider == 'github-copilot' && secrets.COPILOT_GITHUB_TOKEN || '' }}
run: >
uv run --no-project --python 3.13
--with ./modules/hooks-routing --with click
--with git+https://github.com/microsoft/amplifier-module-provider-${{ matrix.provider }}@main
--with git+https://github.com/microsoft/amplifier-core@main
python scripts/check_model_catalogs.py --provider '${{ matrix.provider }}'
52 changes: 49 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ Eight curated matrices ship with this bundle, plus one explicit-name alias
| **quality** | Maximum capability. Uses the strongest models for every role, regardless of cost. |
| **economy** | Cost-optimized. Prefers free tiers, smaller models, and local providers like Ollama. |
| **anthropic** | Anthropic Claude models exclusively. No knob-consistent delegation -- no measured win for this family yet (the "Anthropic guardrail"). |
| **openai** | OpenAI models exclusively. **Ships knob-consistent delegation ON by default** (2026-09-02, a measured win -- see below). Covers **both OpenAI backends** — the pay-per-use API (`openai`) and the ChatGPT subscription (`openai-chatgpt`) — with a single candidate per role: the resolver treats them as one family, API key first. The flagship `sol` tier is PAUSED as of 2026-09-07 pending further evals; roles that used it now run `terra` one effort notch higher, and `ui-coding` runs `luna` at `max`. |
| **openai** | OpenAI models exclusively, API-first across both backends. Knob-consistent delegation is ON by default. The interim October 2 update selects Sol 6.1 instead of Terra and current Luna for fast/UI roles; `ui-coding` keeps `max` effort. Comparative selection evals follow separately. |
| **gemini** | Google Gemini models exclusively. |
| **copilot** | Strongest model per role across the Claude, GPT and Gemini families Copilot serves, through a single provider. |
| **ollama** | A **template**, deliberately minimal: only the two required roles (`general`, `fast`), both `model: "*"`. Every Ollama user has pulled a different set of models, so there is no useful curation to ship -- copy it and pin what you actually have. |
Expand All @@ -24,6 +24,52 @@ Eight curated matrices ship with this bundle, plus one explicit-name alias

Browse the matrix files directly in the [`routing/`](routing/) directory.

### October 2026 catalog refresh

Copilot pins use Sonnet/Opus 5.5, Luna 6 and Sol 6.1; Copilot vision uses
advertised Sonnet 5.5 instead of the unadvertised Gemini 3.5 Flash pin. OpenAI
Luna globs now accept both dotted and whole-generation GPT-6+ IDs without `-fast`
siblings. The subsequent approved interim patch replaces every shipped Terra
candidate (including mixed-matrix Copilot fallbacks) with exact `gpt-6.1-sol`.
This supersedes the September Sol pause. Role ordering and efforts are retained;
this is an opportunistic model refresh, not evidence of a comparative eval win.
Legacy Terra remains only in caller classification and historical examples.

The `openai` preset canonicalizes the ChatGPT backend through the same provider
family aliases as model selection. When no curated candidate fits the caller's
ceiling, in-family substitution preserves the caller's exact model and effort,
instead of turning a broad classification glob into a `-fast` selection. Explicit
fast callers remain fast. The ladder recognizes GPT-6 Luna, Sol and Astra;
Astra shares the upper ordinal rung for ceiling classification only, not as a
cost-equivalence claim or a new role candidate.

### Catalog freshness checks

`scripts/check_model_catalogs.py` is a catalog-only Click wrapper around reusable
checks in the routing module. Install the routing hook, Click and the requested
provider modules, then run:

```bash
python scripts/check_model_catalogs.py --provider openai --provider anthropic \
--provider gemini --provider github-copilot
```

It fails on missing pins/glob matches, newer same-family standard IDs, or vision
without advertised support. It checks lower-priority candidates too, preserves
same-family comparisons, disables Copilot disk fallback, and never generates
content or edits model settings. Missing/empty/failed catalogs are failures, not
freshness success. Output contains coverage counts and repo-declared patterns,
not credentials, exception bodies or account-specific inventories.

`.github/workflows/model-catalogs.yml` provides manual dispatch and a weekly
schedule. The schedule is **disabled** unless the repository variable
`MODEL_CATALOG_CHECKS_ENABLED=true` is set and dedicated provider secrets are
provisioned (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GOOGLE_API_KEY`,
`COPILOT_GITHUB_TOKEN`). The Actions token alone does not grant Copilot access.
ChatGPT OAuth and local Ollama are outside this hosted workflow; their omission
is coverage not provided, not proof of freshness. Ordinary PR tests remain
secret-free. Fresh catalogs do not replace inference smoke tests or quality evals.

## Including the Bundle

**Foundation already includes this bundle** — no extra configuration needed if you use Foundation.
Expand Down Expand Up @@ -136,9 +182,9 @@ Levels 1 and 2 stay strictly above level 3, so an author who deliberately pinned
preset:
tier_ladder: # cheapest -> most expensive, declared
openai:
- ["gpt-?.?-luna*", "gpt-?.?-mini*", "gpt-?.?-nano*"]
- ["gpt-[6-9]*-luna", "gpt-[0-9]*-luna*", "gpt-?.?-mini*", "gpt-?.?-nano*"]
- ["gpt-?.?-terra*"]
- ["gpt-?.?-sol*", "gpt-[0-9].[0-9]"]
- ["gpt-[0-9]*-sol*", "gpt-[0-9]*-astra*", "gpt-[0-9].[0-9]"]
delegation:
inherit: strict # none | effort | tier-and-effort | strict
report_unhonored: true
Expand Down
19 changes: 13 additions & 6 deletions docs/MATRIX_CURATOR_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -459,7 +459,10 @@ After reviewing benchmark data and weather report alignment, follow this 3-step

### Pin Model Names

Always use exact, versioned model names in matrix files. Globs are for user overrides and local providers only.
Use exact versioned names for deliberate pins, or bounded class-scoped globs
when live discovery is supported. Never use a broad glob as a tier policy.
The October 2026 interim update pins Sol 6.1, replacing Terra selections while
retaining role efforts. This approved opportunistic choice is not a quality eval.

**Good ✅ — class-scoped globs** for providers whose `list_models()` is backed
by a live API. These auto-track new releases within a class without silently
Expand All @@ -484,6 +487,10 @@ suffixed sibling, a class glob will find it.
> glob **selects** one model (narrow is right — it is how `-fast` is excluded);
> a ladder glob **classifies** a model already chosen (broad is right — a user
> who hand-pins `gpt-5.6-terra-fast` must still land on the terra rung).
> In-family ceiling substitution uses the caller's exact model, not that broad
> classification glob. This prevents an unrequested fast-sibling substitution.
> ChatGPT caller contexts use the canonical OpenAI ladder unless a custom preset
> explicitly declares a backend-specific ladder.

**`openai-chatgpt` is `openai` to the resolver.** It is a separate provider
MODULE (the OAuth/ChatGPT-subscription backend) — same models, different bill —
Expand Down Expand Up @@ -528,8 +535,8 @@ reject `thinking_level`.
> the base alias explicitly.
| `gpt-?.?-sol*` | any dotted-version sol / flagship tier (e.g. `gpt-5.6-sol`) | base, terra, mini, nano, luna, pro |
| `gpt-?.?-terra*` | any dotted-version terra / mid tier (e.g. `gpt-5.6-terra`) | base, sol, mini, nano, luna, pro |
| `gpt-?.?-terra` | the standard terra id ONLY — **the shipped form** | everything above, **plus `-fast` variants and dated snapshots** |
| `gpt-?.?-luna` | the standard luna id ONLY — **the shipped form** | everything above, **plus `-fast` variants and dated snapshots** |
| `gpt-?.?-terra` | legacy standard Terra IDs for custom overrides; no longer a stock candidate | everything above, **plus `-fast` variants and dated snapshots** |
| `gpt-[6-9]*-luna` | standard dotted or whole-generation Luna IDs, generation 6–9 — **the shipped form** | pre-6 and other named tiers, **plus `-fast` variants and dated snapshots** |
| `gpt-?.?-luna*` | any dotted-version luna / cheap-fast tier (e.g. `gpt-5.6-luna`) | base, terra, mini, nano, sol, pro |
| `gpt-?.?-mini*` | any dotted-version mini (e.g. `gpt-5.4-mini`) | base, pro, nano, sol, luna, `gpt-5-mini` (no dot) |
| `gpt-?.?-nano*` | any dotted-version nano | base, mini, pro, sol, luna |
Expand Down Expand Up @@ -627,10 +634,10 @@ Different providers use different naming conventions for the **same underlying m
| Claude Sonnet 4.x | `claude-sonnet-*` (glob) | — | — | `claude-sonnet-4.6` (pin) |
| Claude Opus 4.x | `claude-opus-*` (glob) | — | — | `claude-opus-4.8` (pin) |
| Claude Haiku 4.x | `claude-haiku-*` (glob) | — | — | `claude-haiku-4.5` (pin) |
| GPT mid-tier (terra) | — | `gpt-?.?-terra` (glob, **no `*`**) | — | pinned, e.g. `gpt-5.6-terra` |
| GPT mid-tier (terra) | — | legacy caller classification only | — | no stock Terra pin |
| GPT base / pre-5.6 migration fallback | — | `gpt-[0-9].[0-9]` (glob) | — | — |
| GPT flagship (sol) | — | `gpt-?.?-sol*` (glob) | — | pinned, e.g. `gpt-5.6-sol` |
| GPT cheap-fast (luna) | — | `gpt-?.?-luna` (glob, **no `*`**) | — | pinned, e.g. `gpt-5.6-luna` |
| GPT flagship (sol) | — | exact `gpt-6.1-sol` | — | exact `gpt-6.1-sol` |
| GPT cheap-fast (luna) | — | `gpt-[6-9]*-luna` (glob, **no trailing `*`**) | — | exact `gpt-6-luna` |
| GPT-5.x mini | — | `gpt-?.?-mini*` (glob) | — | pinned, e.g. `gpt-5.4-mini` |
| Gemini Pro | — | — | `gemini-[3-9]*-pro-preview` (glob) | — |
| Gemini Flash | — | — | `gemini-[3-9]*-flash` (glob) | — |
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
"""Catalog freshness checks; no inference, configuration writes or model policy."""

from __future__ import annotations

import asyncio
import fnmatch
import re
from typing import Any

from .resolver import _is_glob, _version_sort_key


def _family(model: str) -> str | None:
"""Only compare standard IDs within a named class, never cross tiers."""
match = re.fullmatch(r"gpt-\d+(?:\.\d+)?-(sol|terra|luna)", model)
if match:
return "gpt-" + match[1]
match = re.fullmatch(r"claude-(sonnet|opus|haiku)-\d+(?:[.-]\d+)*", model)
if match:
return "claude-" + match[1]
match = re.fullmatch(
r"gemini-\d+(?:\.\d+)?-(flash|flash-lite|pro|pro-preview|pro-image|flash-image)", model
)
if match:
return "gemini-" + ("pro" if match[1] == "pro-preview" else match[1])
return None


def _freshness_key(model: str) -> tuple:
"""Compare Pro GA/preview as peers; prefer GA on an equal release version."""
if _family(model) == "gemini-pro":
return (_version_sort_key(model.removesuffix("-preview")), not model.endswith("-preview"))
return (_version_sort_key(model), True)


async def collect_catalogs(providers: dict[str, Any], timeout: float = 60) -> dict[str, dict]:
"""Fetch actual provider menus; callers must disable static/cache fallback.

Serialize only public model IDs/capabilities and exception class names.
Never serialize provider objects, config, exception text or request headers.
"""
catalogs = {}
for name, provider in providers.items():
try:
models = await asyncio.wait_for(provider.list_models(), timeout)
catalogs[name] = {
"status": "ok" if models else "empty",
"models": [
{"id": m.id, "capabilities": list(getattr(m, "capabilities", []))}
for m in models
],
}
except Exception as error:
catalogs[name] = {"status": "error", "error_type": type(error).__name__, "models": []}
return catalogs


def audit_matrix_catalogs(matrices: dict[str, dict], catalogs: dict[str, dict]) -> dict:
"""Validate all candidates on requested backends, including lower fallbacks.

Unrequested providers are outside this check, not evidence of availability.
Exact-pin acceptance in the runtime resolver intentionally remains unchanged.
"""
issues, checked = [], 0
for backend, catalog in catalogs.items():
if catalog.get("status") != "ok" or not catalog.get("models"):
issues.append({"kind": "catalog_unavailable", "provider": backend})
continue
models = {m["id"]: m for m in catalog["models"]}
names = list(models)
for matrix_name, matrix in matrices.items():
for role, definition in matrix["roles"].items():
for index, candidate in enumerate(definition["candidates"]):
wanted = candidate["provider"]
if backend != wanted and not (wanted == "openai" and backend == "openai-chatgpt"):
continue
checked += 1
pattern = candidate["model"]
context = {"provider": backend, "matrix": matrix_name, "role": role,
"candidate": index + 1, "pattern": pattern}
matches = [m for m in names if fnmatch.fnmatch(m.lower(), pattern.lower())]
if not matches:
issues.append({"kind": "missing_glob" if _is_glob(pattern) else "missing_pin",
**context})
continue
selected = max(matches, key=_version_sort_key)
family = _family(selected)
peers = [m for m in names if family and _family(m) == family]
if peers:
latest = max(peers, key=_freshness_key)
if _freshness_key(latest) > _freshness_key(selected):
issues.append({"kind": "stale_model", **context,
"selected": selected, "latest": latest})
if role == "vision" and "vision" not in models[selected].get("capabilities", []):
issues.append({"kind": "vision_not_advertised", **context, "selected": selected})
return {"checked_candidates": checked, "issues": issues,
"status": "fail" if issues or not checked else "pass"}
Original file line number Diff line number Diff line change
Expand Up @@ -605,13 +605,23 @@ def derive_caller_context(

module = str(best.get("module") or "")
family = module.replace("provider-", "") or str(best.get("id") or "")
# Model selection and tier inheritance must agree about backend aliases.
# Preserve an explicitly declared backend ladder in a custom preset.
if preset is not None and family not in preset.tier_ladder:
from .resolver import PROVIDER_FAMILY_ALIASES

family = next(
(canonical for canonical, aliases in PROVIDER_FAMILY_ALIASES.items()
if family in aliases and canonical in preset.tier_ladder),
family,
)
effort_key = (preset or Preset()).effort_key_for(family)
effort = cfg.get(effort_key)
return CallerContext(
family=family,
model=model,
effort=str(effort) if effort is not None else None,
provider_key=str(best.get("id") or module),
provider_key=str(best.get("instance_id") or best.get("id") or module),
)


Expand Down Expand Up @@ -840,10 +850,20 @@ def plan_candidates(
return candidates, record if preset.report_unhonored else None

rung_index = min(caller_rung, len(sub_ladder) - 1)
# A ladder classifies suffix variants; it is not a safe selection glob.
# In-family fallback can honor the caller's exact chosen model, including
# an explicitly chosen -fast variant, without inventing a different sibling.
substitute_model = sub_ladder[rung_index][0]
substitute_provider = sub_family
if sub_family == caller.family and rung_of(caller.model, sub_ladder) == rung_index:
substitute_model = caller.model
# Backend-specific models (for example ChatGPT -fast variants) must
# remain on the mount that selected them, not API-first family lookup.
substitute_provider = caller.provider_key or sub_family
substitute = _with_effort(
{
"provider": sub_family,
"model": sub_ladder[rung_index][0],
"provider": substitute_provider,
"model": substitute_model,
"config": dict(original_top.get("config") or {}),
},
preset,
Expand Down
Loading
Loading