Hi — first, this is a genuinely well-designed evaluation kernel. DatasetManifest / DatasetCase / BusinessRule freezing the prompt+schema+rules contract, and AttemptResult / EvaluationResult / CandidateMetrics / DecisionReport giving explicit cost provenance (actual / estimated / unknown, never silently zero) is a level of rigor most VLM eval tools skip.
I maintain EvalPort, a small open JSON Schema + Python SDK (evalport-sdk) for a portable "Suite" (test cases) / "ResultSet" (results) interchange format, so eval datasets and results aren't locked to one tool's internal schema. openeval.validate.validate_suite() and validate_result_set() check conformance.
I'm not proposing to touch VLMForge's public types, vlmforge/domain/models.py, or any algorithm — CONTRIBUTING is right to protect those. I'm raising whether an optional, separate interop module (e.g. vlmforge/interop/evalport.py) that maps your existing types to EvalPort's schema would be worth having, purely as opt-in export for teams that want to move a DatasetBundle or a completed run's results out of VLMForge into a tool-agnostic format, or archive them that way.
Sketching what such a mapping could look like — this is a proposal for discussion, not code I've built or tested against your kernel:
# vlmforge/interop/evalport.py (proposed, illustrative only)
from vlmforge.domain.models import (
AttemptResult, BusinessRule, DatasetCase, DatasetManifest, EvaluationResult,
)
def suite_from_dataset(
manifest: DatasetManifest,
cases: tuple[DatasetCase, ...],
rules: tuple[BusinessRule, ...],
) -> dict:
"""DatasetBundle -> EvalPort Suite.
Suite.version must be semver; manifest.version is a free-form string today
(e.g. could default to "1.0.0" and carry manifest.version in Suite metadata
instead, open question).
"""
return {
"version": "1.0.0",
"id": manifest.name,
"test_cases": [
{
"id": case.case_id,
"input": {"images": case.images, "variables": case.variables},
"graders": [
{"id": rule.id, "type": "jsonlogic", "config": rule.logic}
for rule in rules
]
+ [{"id": "schema", "type": "json_schema"}],
}
for case in cases
],
}
def result_set_from_attempts(
run_id: str,
suite_id: str,
started_at: str,
attempts: list[AttemptResult],
evaluations: dict[str, list[EvaluationResult]], # keyed by attempt_id
) -> dict:
"""attempts + their EvaluationResults -> EvalPort ResultSet."""
return {
"version": "1.0.0",
"suite_id": suite_id,
"run_id": run_id,
"started_at": started_at,
"results": [
{
"test_case_id": attempt.case_id,
"passed": all(
e.passed for e in evaluations.get(attempt.attempt_id, [])
if e.passed is not None
),
"grader_results": [
{
"grader_id": e.evaluator,
"passed": e.passed,
"score": e.score,
"details": e.details,
}
for e in evaluations.get(attempt.attempt_id, [])
],
}
for attempt in attempts
],
}
Open questions I'd want feedback on before anyone writes real code for this:
- How should
RuleSeverity.CRITICAL failures vs. ordinary passed=False round-trip into EvalPort's grader results — a severity field on grader_results, or is that VLMForge-specific and better left out of the interop layer entirely?
- Do
CandidateMetrics / DecisionReport (Pareto frontier, gates, recommendation) belong anywhere in ResultSet metadata, or should they stay VLMForge-only since they're about candidate selection rather than a single result set?
DatasetManifest.version is a free string; EvalPort's Suite.version requires semver — worth changing on either side, or just documenting the adaptation in the interop module?
Per CONTRIBUTING's process of clarifying terminology/boundaries in an issue before coding: happy to help build and test this as a PR if there's appetite, but wanted to check interest and get your read on the mapping first rather than show up with an unsolicited PR touching the kernel. No pressure either way if this isn't a priority — flagging it because your typed kernel is one of the few VLM eval tools I've looked at where an interchange mapping would actually be this straightforward, thanks to how disciplined the domain models already are.
Hi — first, this is a genuinely well-designed evaluation kernel.
DatasetManifest/DatasetCase/BusinessRulefreezing the prompt+schema+rules contract, andAttemptResult/EvaluationResult/CandidateMetrics/DecisionReportgiving explicit cost provenance (actual/estimated/unknown, never silently zero) is a level of rigor most VLM eval tools skip.I maintain EvalPort, a small open JSON Schema + Python SDK (
evalport-sdk) for a portable "Suite" (test cases) / "ResultSet" (results) interchange format, so eval datasets and results aren't locked to one tool's internal schema.openeval.validate.validate_suite()andvalidate_result_set()check conformance.I'm not proposing to touch VLMForge's public types,
vlmforge/domain/models.py, or any algorithm — CONTRIBUTING is right to protect those. I'm raising whether an optional, separate interop module (e.g.vlmforge/interop/evalport.py) that maps your existing types to EvalPort's schema would be worth having, purely as opt-in export for teams that want to move aDatasetBundleor a completed run's results out of VLMForge into a tool-agnostic format, or archive them that way.Sketching what such a mapping could look like — this is a proposal for discussion, not code I've built or tested against your kernel:
Open questions I'd want feedback on before anyone writes real code for this:
RuleSeverity.CRITICALfailures vs. ordinarypassed=Falseround-trip into EvalPort's grader results — a severity field ongrader_results, or is that VLMForge-specific and better left out of the interop layer entirely?CandidateMetrics/DecisionReport(Pareto frontier, gates, recommendation) belong anywhere in ResultSet metadata, or should they stay VLMForge-only since they're about candidate selection rather than a single result set?DatasetManifest.versionis a free string; EvalPort'sSuite.versionrequires semver — worth changing on either side, or just documenting the adaptation in the interop module?Per CONTRIBUTING's process of clarifying terminology/boundaries in an issue before coding: happy to help build and test this as a PR if there's appetite, but wanted to check interest and get your read on the mapping first rather than show up with an unsolicited PR touching the kernel. No pressure either way if this isn't a priority — flagging it because your typed kernel is one of the few VLM eval tools I've looked at where an interchange mapping would actually be this straightforward, thanks to how disciplined the domain models already are.