You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reusable judge definitions and calibration with independently bounded rubric axes, separate from model evaluation execution and Platform hosting.
Immutable JudgeDefinition with a content-derived revision, named syllabus and a meaning for every score. The default remains binary task success.
Public Python judge setup and calibration APIs reuse the existing implementation and exact prompt identity.
Signed, independent axis ranges flow through LM parsing, human labels, draft/final review, grouped calibration and normalized judgments. Each axis is limited to 101 score points before allocation.
Authored criteria, human approval and empirical calibration remain distinct. No calibration provenance is fabricated.
Current restack
Integrated current main 6bb4e9618, including the merged capture changes and simulated tool responses. No additional capture implementation, Platform changes, dependency bump, publication or deployment is included. Score contracts now document signed bounds and provenance under the current repository documentation rules.
Current head: f9a114d11c438e6243dbadef9df52abb15f3cb29.
Full-repository Ruff, format and ty checks pass.
All 171 focused judging, calibration and judge CLI tests pass.
Full local Python suite: 5,723 passed, 6 skipped, in 141.31 seconds. This includes installed-wheel validation on the clean committed head.
Fresh Greptile 5/5 on this exact head, with no new actionable findings.
All current-head remote CI checks pass, including the full gate, package build, five-platform native wheels, Darwin evidence, latency, CodeQL and security. Publication jobs are intentionally skipped. No paid provider calls or model-quality measurement is claimed.
Scope
Follows the latest Evals product direction. Experiential owns judges, calibration and evaluation behavior; Platform owns authorization, credit admission and hosting. No CLI expansion, UX redesign, README changes, paid calls or merge is included. #1004 remains the stacked evaluation API.
The current restacked head appears safe to merge, with no outstanding findings or newly introduced actionable defects.
Summary
The PR adds reusable, content-addressed judge definitions and carries independently bounded signed rubric axes through setup, labeling, judging, calibration, and normalized judgments.
Exposes judge-definition and manual calibration APIs through the public Python package.
Binds authored judge syllabi and dimensions to exact prompt identities without conflating authored criteria with empirical calibration.
Supports independent signed score ranges while limiting each axis and score map to 101 integer points.
Validates human labels, model judgments, review drafts, and calibration observations against their exact rubric axes.
The changes since the previous review clarify signed-bound and provenance contracts; the previously reported direct-write bounds issue is fixed.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
D[JudgeDefinition<br/>name, syllabus, axes] --> T[Bound prompt template<br/>content-derived identity]
D --> R[Persisted rubric<br/>independent axis ranges]
T --> J[Structured judge proposals]
R --> J
R --> H[Human labels and corrections]
J --> C[Grouped calibration]
H --> C
C --> M[Per-axis score maps]
J --> F[Final judgments]
M --> F
F --> N[Equal-weight normalized score]
@greptileai Please re-review the restacked head on main bbfc71a. Full-repository Ruff/format/ty and focused feature tests pass. The latest Platform UX is authoritative, with the underlying judge/evaluation implementation kept in Experiential. Full-suite results are being refreshed; no merge or release authorized.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reusable judge definitions and calibration with independently bounded rubric axes, separate from model evaluation execution and Platform hosting.
JudgeDefinitionwith a content-derived revision, named syllabus and a meaning for every score. The default remains binary task success.Current restack
Integrated current main
6bb4e9618, including the merged capture changes and simulated tool responses. No additional capture implementation, Platform changes, dependency bump, publication or deployment is included. Score contracts now document signed bounds and provenance under the current repository documentation rules.Current head:
f9a114d11c438e6243dbadef9df52abb15f3cb29.Scope
Follows the latest Evals product direction. Experiential owns judges, calibration and evaluation behavior; Platform owns authorization, credit admission and hosting. No CLI expansion, UX redesign, README changes, paid calls or merge is included. #1004 remains the stacked evaluation API.