Skip to content

[FEATURE] Failure cohort analysis for evaluation results #348

Description

@max-rattray-aws

[FEATURE] Group failed cases by evaluator (failure cohort analysis)

Problem Statement

After running a multi-evaluator experiment, EvaluationReport gives you a flat list of pass/fail results. If 30 out of 100 cases failed, you can't immediately tell whether those are 30 separate problems or 3 problems that each hit 10 cases.

Right now you have to loop through report.cases, check report.test_passes, and group by the "evaluator" key yourself. Everyone who runs experiments with multiple evaluators writes some version of this grouping code. It belongs in the library.

The grouping is simple. Take all failed cases, bucket them by evaluator name, sort buckets by size (largest first). A bucket of 15 Faithfulness failures is a systemic problem. A bucket with 1 failure is an edge case. This tells you where to focus.

Proposed Solution

Add a function that takes an EvaluationReport and returns grouped failure data. Something like strands_evals.analysis.group_failures_by_evaluator (new module) or a method on EvaluationReport itself.

The result types should follow the project's Pydantic convention:

from pydantic import BaseModel


class FailureCohort(BaseModel):
    """A group of test cases that all failed the same evaluator."""

    evaluator_name: str
    failed_case_indices: list[int]
    failed_case_names: list[str]
    count: int

    @property
    def is_systemic(self) -> bool:
        """Two or more failures suggests a shared root cause."""
        return self.count >= 2


class CohortAnalysis(BaseModel):
    """Result of grouping failed cases by evaluator."""

    cohorts: list[FailureCohort]
    total_failures: int
    total_cases: int

    @property
    def systemic_cohorts(self) -> list[FailureCohort]:
        return [c for c in self.cohorts if c.is_systemic]

    @property
    def one_off_failures(self) -> list[FailureCohort]:
        return [c for c in self.cohorts if c.count == 1]

The function itself:

from strands_evals.types.evaluation_report import EvaluationReport


def analyze_failure_cohorts(report: EvaluationReport) -> CohortAnalysis:
    """Group failed cases by evaluator, sorted largest-first.

    Each case in the report has an "evaluator" key identifying which evaluator
    produced that row. Failed cases get bucketed by that key.
    """
    failures_by_evaluator: dict[str, list[tuple[int, str]]] = {}

    for i, (case, passed) in enumerate(zip(report.cases, report.test_passes)):
        if passed:
            continue
        eval_name = case.get("evaluator", "unknown")
        case_name = case.get("name", f"case_{i}")
        failures_by_evaluator.setdefault(eval_name, []).append((i, case_name))

    cohorts = []
    for eval_name, members in failures_by_evaluator.items():
        indices = [m[0] for m in members]
        names = [m[1] for m in members]
        cohorts.append(
            FailureCohort(
                evaluator_name=eval_name,
                failed_case_indices=indices,
                failed_case_names=names,
                count=len(members),
            )
        )

    cohorts.sort(key=lambda c: (-c.count, c.evaluator_name))

    return CohortAnalysis(
        cohorts=cohorts,
        total_failures=sum(1 for p in report.test_passes if not p),
        total_cases=len(report.test_passes),
    )

Usage:

from strands_evals import Experiment

experiment = Experiment(cases=my_cases, evaluators=[correctness, faithfulness, harmfulness])
report = experiment.run_evaluations(task=my_task)

analysis = analyze_failure_cohorts(report)

for cohort in analysis.systemic_cohorts:
    print(f"{cohort.evaluator_name}: {cohort.count} failures")
    print(f"  Cases: {', '.join(cohort.failed_case_names)}")

The function is pure. No side effects, no model calls. Works on any EvaluationReport including ones loaded via EvaluationReport.from_file().

Optional display helper using the existing Rich dependency:

def print_cohort_summary(analysis: CohortAnalysis) -> None:
    from rich.console import Console
    from rich.table import Table

    console = Console()
    table = Table(title=f"Failure Cohorts ({analysis.total_failures}/{analysis.total_cases} failed)")
    table.add_column("Evaluator", style="bold")
    table.add_column("Count", justify="right")
    table.add_column("Cases")

    for cohort in analysis.cohorts:
        names = ", ".join(cohort.failed_case_names[:5])
        if cohort.count > 5:
            names += f" (+{cohort.count - 5} more)"
        table.add_row(cohort.evaluator_name, str(cohort.count), names)

    console.print(table)

Use Case

  1. You run 50 cases with Correctness, Faithfulness, and Harmfulness evaluators. 20 fail. Cohort analysis shows 14 of those are all Faithfulness. That's one problem (hallucination), not 20.

  2. You're comparing two prompt versions. Both score similarly overall. But version A has 8 failures concentrated in one evaluator while version B has 8 spread across 6. The cohort view makes this obvious.

  3. CI gating. Run cohort analysis on nightly evals. Flag any cohort larger than N as a regression. Scattered one-off failures are noise. A cohort of 10 in the same evaluator is signal.

Alternatives Considered

You can do this in 10 lines with collections.Counter. But having it in the library means consistent naming, consistent sort order, and a typed object that other analysis tools can build on. The existing detectors module (failure_detector, root_cause_analyzer) works at the Session/trace level. This sits one level above, operating on report-level pass/fail data.

Additional Context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesarea-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or request

    Fields

    Language

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions