Skip to content

Proposal: optional EvalPort interop mapping for DatasetBundle / AttemptResult (not a public-type change) #1

Description

@adhabnr-ux

Hi — first, this is a genuinely well-designed evaluation kernel. DatasetManifest / DatasetCase / BusinessRule freezing the prompt+schema+rules contract, and AttemptResult / EvaluationResult / CandidateMetrics / DecisionReport giving explicit cost provenance (actual / estimated / unknown, never silently zero) is a level of rigor most VLM eval tools skip.

I maintain EvalPort, a small open JSON Schema + Python SDK (evalport-sdk) for a portable "Suite" (test cases) / "ResultSet" (results) interchange format, so eval datasets and results aren't locked to one tool's internal schema. openeval.validate.validate_suite() and validate_result_set() check conformance.

I'm not proposing to touch VLMForge's public types, vlmforge/domain/models.py, or any algorithm — CONTRIBUTING is right to protect those. I'm raising whether an optional, separate interop module (e.g. vlmforge/interop/evalport.py) that maps your existing types to EvalPort's schema would be worth having, purely as opt-in export for teams that want to move a DatasetBundle or a completed run's results out of VLMForge into a tool-agnostic format, or archive them that way.

Sketching what such a mapping could look like — this is a proposal for discussion, not code I've built or tested against your kernel:

# vlmforge/interop/evalport.py (proposed, illustrative only)
from vlmforge.domain.models import (
    AttemptResult, BusinessRule, DatasetCase, DatasetManifest, EvaluationResult,
)


def suite_from_dataset(
    manifest: DatasetManifest,
    cases: tuple[DatasetCase, ...],
    rules: tuple[BusinessRule, ...],
) -> dict:
    """DatasetBundle -> EvalPort Suite.

    Suite.version must be semver; manifest.version is a free-form string today
    (e.g. could default to "1.0.0" and carry manifest.version in Suite metadata
    instead, open question).
    """
    return {
        "version": "1.0.0",
        "id": manifest.name,
        "test_cases": [
            {
                "id": case.case_id,
                "input": {"images": case.images, "variables": case.variables},
                "graders": [
                    {"id": rule.id, "type": "jsonlogic", "config": rule.logic}
                    for rule in rules
                ]
                + [{"id": "schema", "type": "json_schema"}],
            }
            for case in cases
        ],
    }


def result_set_from_attempts(
    run_id: str,
    suite_id: str,
    started_at: str,
    attempts: list[AttemptResult],
    evaluations: dict[str, list[EvaluationResult]],  # keyed by attempt_id
) -> dict:
    """attempts + their EvaluationResults -> EvalPort ResultSet."""
    return {
        "version": "1.0.0",
        "suite_id": suite_id,
        "run_id": run_id,
        "started_at": started_at,
        "results": [
            {
                "test_case_id": attempt.case_id,
                "passed": all(
                    e.passed for e in evaluations.get(attempt.attempt_id, [])
                    if e.passed is not None
                ),
                "grader_results": [
                    {
                        "grader_id": e.evaluator,
                        "passed": e.passed,
                        "score": e.score,
                        "details": e.details,
                    }
                    for e in evaluations.get(attempt.attempt_id, [])
                ],
            }
            for attempt in attempts
        ],
    }

Open questions I'd want feedback on before anyone writes real code for this:

  • How should RuleSeverity.CRITICAL failures vs. ordinary passed=False round-trip into EvalPort's grader results — a severity field on grader_results, or is that VLMForge-specific and better left out of the interop layer entirely?
  • Do CandidateMetrics / DecisionReport (Pareto frontier, gates, recommendation) belong anywhere in ResultSet metadata, or should they stay VLMForge-only since they're about candidate selection rather than a single result set?
  • DatasetManifest.version is a free string; EvalPort's Suite.version requires semver — worth changing on either side, or just documenting the adaptation in the interop module?

Per CONTRIBUTING's process of clarifying terminology/boundaries in an issue before coding: happy to help build and test this as a PR if there's appetite, but wanted to check interest and get your read on the mapping first rather than show up with an unsolicited PR touching the kernel. No pressure either way if this isn't a priority — flagging it because your typed kernel is one of the few VLM eval tools I've looked at where an interchange mapping would actually be this straightforward, thanks to how disciplined the domain models already are.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions