Skip to content

feat: add routing catalog and maintained eval program - #81

Merged
Brian Krabach (bkrabach) merged 2 commits into
mainfrom
feat/routing-catalog-v1
Oct 3, 2026
Merged

Brian Krabach (bkrabach) merged 2 commits into
mainfrom
feat/routing-catalog-v1

Conversation

@bkrabach

@bkrabach Brian Krabach (bkrabach) commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Add the accepted distinct-ID routing catalog first slice: openai-api and openai-chatgpt are additive profiles; the original eight matrix policies and golden entries remain unchanged.
  • Add the bounded v1 catalog library, metadata/schema, resolver integration, tests, and catalog API guide.
  • Relocate and maintain routing-specific evaluation configurations, scenarios, graders, reuse rules, and promotion policy in evals/, with policy in docs/EVALUATION-PROGRAM.md. Generic execution bricks remain in the separate evaluation library as optional development dependencies; no evaluator dependency is added to the production hook.

Evaluation tooling boundary

The relocated evidence.py library and thin Click CLI implement offline plan, readiness, and analyze only. Every CLI command requires an explicit --benchmark-root; there is no ambient sibling lookup, download, network access, model execution, or paid benchmark runner. Historical sample bytes and task/source/grader locks are preserved rather than rewritten to imply evidence for the current routing revision. The CI job acquires the pinned public task assets from evaluation revision 7c3646796eea2b042bbe6ffb2de3906d31daa381, without installing its evaluator runtime, executing upstream scripts, using secrets, or uploading results.

Catalog scope and limits

The catalog API provides list, describe, assess, and discover_provider_only over explicit policy directories and immutable caller-supplied snapshots. It does not call providers/models or modify settings. Exact module constraints apply only to the bundle's routed paths, not account consent or universal dispatch containment. All reports say enforcement.status: not_enforced and execution_ready: false; explicit-preference assessment and next-dispatch prediction remain unsupported. Consumer adoption, principal/account admission, dynamic-provider evidence, and the broader DRAFT contract remain unqualified. No model preferences or production routing semantics are changed by the evaluation relocation.

Verification

  • Routing and catalog suites: 858 passed, 0 skipped.
  • Relocated evaluation suite: 104 passed, 0 skipped (93 preserved contract cases plus 11 relocation regressions). Combined: 962 passed, 0 skipped against real Core/Foundation dependencies in the same pytest process.
  • Verified actual pinned TaskSpec/GraderConfig task IDs, timeouts, criterion identities and bounds for all five portfolio tasks.
  • Pinned Ruff 0.15.11 checks, scoped formatting, and bundle structure checks passed. Offline CLI plan/readiness behavior was checked from unrelated working directories; omitted benchmark root fails before input reads.
  • No benchmark subject execution, model calls, or network attempts were made by the planner/tests.

Privacy and direction

A fresh-context relocation/source review, semantic stranger/leak review, and publication checks passed. No private raw results, captures, account bindings, local paths, source maps, or reviewer receipts are included. The contract and evaluation program remain DRAFT: this is a bounded implementation, not a frozen contract, live quality finding, recommendation, or automatic promotion system. Existing rubric qualification and future execution/evidence adapters remain separate gates.

Breaking changes

None intended. The evaluation library and its main branch are unchanged; this moves only the routing-owned offline tooling into its owning repository and adds a separate pinned-assets CI job. Existing routing policy remains unchanged apart from the previously described additive profiles.

Add a distinct-ID catalog planning slice with source-aware resolution, exact module scope planning, synthetic contract tests, and draft interface documentation. Preserve existing shipped policy and make enforcement and unsupported prediction boundaries explicit.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Move the routing-owned offline evaluation program into this repository, keeping the generic evaluation library unchanged. Preserve the historical sample and source locks, require explicit benchmark inputs, and add isolated pinned-asset offline CI across Python 3.11-3.13.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach Brian Krabach (bkrabach) changed the title feat: add bounded routing catalog v1 feat: add routing catalog and maintained eval program Oct 2, 2026
@bkrabach
Brian Krabach (bkrabach) merged commit f8bd250 into main Oct 3, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants