[FEATURE] Group failed cases by evaluator (failure cohort analysis)
Problem Statement
After running a multi-evaluator experiment, EvaluationReport gives you a flat list of pass/fail results. If 30 out of 100 cases failed, you can't immediately tell whether those are 30 separate problems or 3 problems that each hit 10 cases.
Right now you have to loop through report.cases, check report.test_passes, and group by the "evaluator" key yourself. Everyone who runs experiments with multiple evaluators writes some version of this grouping code. It belongs in the library.
The grouping is simple. Take all failed cases, bucket them by evaluator name, sort buckets by size (largest first). A bucket of 15 Faithfulness failures is a systemic problem. A bucket with 1 failure is an edge case. This tells you where to focus.
Proposed Solution
Add a function that takes an EvaluationReport and returns grouped failure data. Something like strands_evals.analysis.group_failures_by_evaluator (new module) or a method on EvaluationReport itself.
The result types should follow the project's Pydantic convention:
from pydantic import BaseModel
class FailureCohort(BaseModel):
"""A group of test cases that all failed the same evaluator."""
evaluator_name: str
failed_case_indices: list[int]
failed_case_names: list[str]
count: int
@property
def is_systemic(self) -> bool:
"""Two or more failures suggests a shared root cause."""
return self.count >= 2
class CohortAnalysis(BaseModel):
"""Result of grouping failed cases by evaluator."""
cohorts: list[FailureCohort]
total_failures: int
total_cases: int
@property
def systemic_cohorts(self) -> list[FailureCohort]:
return [c for c in self.cohorts if c.is_systemic]
@property
def one_off_failures(self) -> list[FailureCohort]:
return [c for c in self.cohorts if c.count == 1]
The function itself:
from strands_evals.types.evaluation_report import EvaluationReport
def analyze_failure_cohorts(report: EvaluationReport) -> CohortAnalysis:
"""Group failed cases by evaluator, sorted largest-first.
Each case in the report has an "evaluator" key identifying which evaluator
produced that row. Failed cases get bucketed by that key.
"""
failures_by_evaluator: dict[str, list[tuple[int, str]]] = {}
for i, (case, passed) in enumerate(zip(report.cases, report.test_passes)):
if passed:
continue
eval_name = case.get("evaluator", "unknown")
case_name = case.get("name", f"case_{i}")
failures_by_evaluator.setdefault(eval_name, []).append((i, case_name))
cohorts = []
for eval_name, members in failures_by_evaluator.items():
indices = [m[0] for m in members]
names = [m[1] for m in members]
cohorts.append(
FailureCohort(
evaluator_name=eval_name,
failed_case_indices=indices,
failed_case_names=names,
count=len(members),
)
)
cohorts.sort(key=lambda c: (-c.count, c.evaluator_name))
return CohortAnalysis(
cohorts=cohorts,
total_failures=sum(1 for p in report.test_passes if not p),
total_cases=len(report.test_passes),
)
Usage:
from strands_evals import Experiment
experiment = Experiment(cases=my_cases, evaluators=[correctness, faithfulness, harmfulness])
report = experiment.run_evaluations(task=my_task)
analysis = analyze_failure_cohorts(report)
for cohort in analysis.systemic_cohorts:
print(f"{cohort.evaluator_name}: {cohort.count} failures")
print(f" Cases: {', '.join(cohort.failed_case_names)}")
The function is pure. No side effects, no model calls. Works on any EvaluationReport including ones loaded via EvaluationReport.from_file().
Optional display helper using the existing Rich dependency:
def print_cohort_summary(analysis: CohortAnalysis) -> None:
from rich.console import Console
from rich.table import Table
console = Console()
table = Table(title=f"Failure Cohorts ({analysis.total_failures}/{analysis.total_cases} failed)")
table.add_column("Evaluator", style="bold")
table.add_column("Count", justify="right")
table.add_column("Cases")
for cohort in analysis.cohorts:
names = ", ".join(cohort.failed_case_names[:5])
if cohort.count > 5:
names += f" (+{cohort.count - 5} more)"
table.add_row(cohort.evaluator_name, str(cohort.count), names)
console.print(table)
Use Case
-
You run 50 cases with Correctness, Faithfulness, and Harmfulness evaluators. 20 fail. Cohort analysis shows 14 of those are all Faithfulness. That's one problem (hallucination), not 20.
-
You're comparing two prompt versions. Both score similarly overall. But version A has 8 failures concentrated in one evaluator while version B has 8 spread across 6. The cohort view makes this obvious.
-
CI gating. Run cohort analysis on nightly evals. Flag any cohort larger than N as a regression. Scattered one-off failures are noise. A cohort of 10 in the same evaluator is signal.
Alternatives Considered
You can do this in 10 lines with collections.Counter. But having it in the library means consistent naming, consistent sort order, and a typed object that other analysis tools can build on. The existing detectors module (failure_detector, root_cause_analyzer) works at the Session/trace level. This sits one level above, operating on report-level pass/fail data.
Additional Context
[FEATURE] Group failed cases by evaluator (failure cohort analysis)
Problem Statement
After running a multi-evaluator experiment,
EvaluationReportgives you a flat list of pass/fail results. If 30 out of 100 cases failed, you can't immediately tell whether those are 30 separate problems or 3 problems that each hit 10 cases.Right now you have to loop through
report.cases, checkreport.test_passes, and group by the"evaluator"key yourself. Everyone who runs experiments with multiple evaluators writes some version of this grouping code. It belongs in the library.The grouping is simple. Take all failed cases, bucket them by evaluator name, sort buckets by size (largest first). A bucket of 15 Faithfulness failures is a systemic problem. A bucket with 1 failure is an edge case. This tells you where to focus.
Proposed Solution
Add a function that takes an
EvaluationReportand returns grouped failure data. Something likestrands_evals.analysis.group_failures_by_evaluator(new module) or a method onEvaluationReportitself.The result types should follow the project's Pydantic convention:
The function itself:
Usage:
The function is pure. No side effects, no model calls. Works on any
EvaluationReportincluding ones loaded viaEvaluationReport.from_file().Optional display helper using the existing Rich dependency:
Use Case
You run 50 cases with Correctness, Faithfulness, and Harmfulness evaluators. 20 fail. Cohort analysis shows 14 of those are all Faithfulness. That's one problem (hallucination), not 20.
You're comparing two prompt versions. Both score similarly overall. But version A has 8 failures concentrated in one evaluator while version B has 8 spread across 6. The cohort view makes this obvious.
CI gating. Run cohort analysis on nightly evals. Flag any cohort larger than N as a regression. Scattered one-off failures are noise. A cohort of 10 in the same evaluator is signal.
Alternatives Considered
You can do this in 10 lines with
collections.Counter. But having it in the library means consistent naming, consistent sort order, and a typed object that other analysis tools can build on. The existingdetectorsmodule (failure_detector, root_cause_analyzer) works at the Session/trace level. This sits one level above, operating on report-level pass/fail data.Additional Context
overall_scorealone.detectorsmodule analyzes individual sessions for root causes. Cohort analysis tells you which group of cases to feed into root cause analysis together.rich>=14.0.0,<15.0.0) so the display helper adds no new deps.strands_evals/analysis/module or a method onEvaluationReport. Open to either.