Repository navigation
feat: add routing catalog and maintained eval program - #81
Merged
Merged
Conversation
Add a distinct-ID catalog planning slice with source-aware resolution, exact module scope planning, synthetic contract tests, and draft interface documentation. Preserve existing shipped policy and make enforcement and unsupported prediction boundaries explicit. Generated with Amplifier Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Move the routing-owned offline evaluation program into this repository, keeping the generic evaluation library unchanged. Preserve the historical sample and source locks, require explicit benchmark inputs, and add isolated pinned-asset offline CI across Python 3.11-3.13. Generated with Amplifier Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
openai-apiandopenai-chatgptare additive profiles; the original eight matrix policies and golden entries remain unchanged.evals/, with policy indocs/EVALUATION-PROGRAM.md. Generic execution bricks remain in the separate evaluation library as optional development dependencies; no evaluator dependency is added to the production hook.Evaluation tooling boundary
The relocated
evidence.pylibrary and thin Click CLI implement offlineplan,readiness, andanalyzeonly. Every CLI command requires an explicit--benchmark-root; there is no ambient sibling lookup, download, network access, model execution, or paid benchmark runner. Historical sample bytes and task/source/grader locks are preserved rather than rewritten to imply evidence for the current routing revision. The CI job acquires the pinned public task assets from evaluation revision7c3646796eea2b042bbe6ffb2de3906d31daa381, without installing its evaluator runtime, executing upstream scripts, using secrets, or uploading results.Catalog scope and limits
The catalog API provides
list,describe,assess, anddiscover_provider_onlyover explicit policy directories and immutable caller-supplied snapshots. It does not call providers/models or modify settings. Exact module constraints apply only to the bundle's routed paths, not account consent or universal dispatch containment. All reports sayenforcement.status: not_enforcedandexecution_ready: false; explicit-preference assessment and next-dispatch prediction remain unsupported. Consumer adoption, principal/account admission, dynamic-provider evidence, and the broader DRAFT contract remain unqualified. No model preferences or production routing semantics are changed by the evaluation relocation.Verification
Privacy and direction
A fresh-context relocation/source review, semantic stranger/leak review, and publication checks passed. No private raw results, captures, account bindings, local paths, source maps, or reviewer receipts are included. The contract and evaluation program remain DRAFT: this is a bounded implementation, not a frozen contract, live quality finding, recommendation, or automatic promotion system. Existing rubric qualification and future execution/evidence adapters remain separate gates.
Breaking changes
None intended. The evaluation library and its
mainbranch are unchanged; this moves only the routing-owned offline tooling into its owning repository and adds a separate pinned-assets CI job. Existing routing policy remains unchanged apart from the previously described additive profiles.