Skip to content

feat: reusable judge definitions and independent rubric ranges - #1003

Merged
kfallah merged 4 commits into
mainfrom
codex/judge-api
Sep 23, 2026
Merged

kfallah merged 4 commits into
mainfrom
codex/judge-api

Conversation

@kfallah

@kfallah kfallah commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Reusable judge definitions and calibration with independently bounded rubric axes, separate from model evaluation execution and Platform hosting.

  • Immutable JudgeDefinition with a content-derived revision, named syllabus and a meaning for every score. The default remains binary task success.
  • Public Python judge setup and calibration APIs reuse the existing implementation and exact prompt identity.
  • Signed, independent axis ranges flow through LM parsing, human labels, draft/final review, grouped calibration and normalized judgments. Each axis is limited to 101 score points before allocation.
  • Authored criteria, human approval and empirical calibration remain distinct. No calibration provenance is fabricated.

Current restack

Integrated current main 6bb4e9618, including the merged capture changes and simulated tool responses. No additional capture implementation, Platform changes, dependency bump, publication or deployment is included. Score contracts now document signed bounds and provenance under the current repository documentation rules.

Current head: f9a114d11c438e6243dbadef9df52abb15f3cb29.

  • Full-repository Ruff, format and ty checks pass.
  • All 171 focused judging, calibration and judge CLI tests pass.
  • Full local Python suite: 5,723 passed, 6 skipped, in 141.31 seconds. This includes installed-wheel validation on the clean committed head.
  • Fresh Greptile 5/5 on this exact head, with no new actionable findings.
  • All current-head remote CI checks pass, including the full gate, package build, five-platform native wheels, Darwin evidence, latency, CodeQL and security. Publication jobs are intentionally skipped. No paid provider calls or model-quality measurement is claimed.

Scope

Follows the latest Evals product direction. Experiential owns judges, calibration and evaluation behavior; Platform owns authorization, credit admission and hosting. No CLI expansion, UX redesign, README changes, paid calls or merge is included. #1004 remains the stacked evaluation API.

@superagent-security superagent-security Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Superagent found 1 security concern(s).

Comment thread exp/optimize/router/judging/contracts.py
@greptile-apps

greptile-apps Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The current restacked head appears safe to merge, with no outstanding findings or newly introduced actionable defects.

Summary

The PR adds reusable, content-addressed judge definitions and carries independently bounded signed rubric axes through setup, labeling, judging, calibration, and normalized judgments.

  • Exposes judge-definition and manual calibration APIs through the public Python package.
  • Binds authored judge syllabi and dimensions to exact prompt identities without conflating authored criteria with empirical calibration.
  • Supports independent signed score ranges while limiting each axis and score map to 101 integer points.
  • Validates human labels, model judgments, review drafts, and calibration observations against their exact rubric axes.
  • The changes since the previous review clarify signed-bound and provenance contracts; the previously reported direct-write bounds issue is fixed.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  D[JudgeDefinition<br/>name, syllabus, axes] --> T[Bound prompt template<br/>content-derived identity]
  D --> R[Persisted rubric<br/>independent axis ranges]
  T --> J[Structured judge proposals]
  R --> J
  R --> H[Human labels and corrections]
  J --> C[Grouped calibration]
  H --> C
  C --> M[Per-axis score maps]
  J --> F[Final judgments]
  M --> F
  F --> N[Equal-weight normalized score]
Loading

Reviews (4) · Last reviewed commit: "Restack judge API on released capture ma..."

Comment thread exp/common/judging/labels.py
@kfallah kfallah changed the title feat: reusable judge definitions and independent calibration ranges feat: reusable judge definitions and independently ranged calibration axes Sep 16, 2026
@kfallah

kfallah commented Sep 16, 2026

Copy link
Copy Markdown
Contributor Author

@greptile please re-review the latest commit; both label-validation findings are addressed with regression tests.

@superagent-security

Copy link
Copy Markdown

Superagent didn't find any vulnerabilities or security issues in this PR.

@kfallah kfallah changed the title feat: reusable judge definitions and independently ranged calibration axes feat: reusable judge definitions and independent rubric ranges Sep 16, 2026

kfallah commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Please re-review the restacked head on main bbfc71a. Full-repository Ruff/format/ty and focused feature tests pass. The latest Platform UX is authoritative, with the underlying judge/evaluation implementation kept in Experiential. Full-suite results are being refreshed; no merge or release authorized.

@kfallah

kfallah commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current restacked head, including the current-main integration and signed rubric contract documentation.

@kfallah
kfallah merged commit ce2de6a into main Sep 23, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant