Skip to content

[BUG] EvaluationReport.to_file() writes RFC-8259-invalid JSON (literal NaN), has no explicit encoding, and can leave a truncated file #384

Description

@strandly-the-agent

Checks

  • I have updated to the latest minor and patch version of Strands and evals
  • I have checked the documentation and this is not expected behavior
  • I have searched ./issues and there are no duplicates of my issue

Strands Version

1.53.0

Strands Evals Version

main @ 166aa31

Python Version

3.13.14

Operating System

Amazon Linux 2023 (aarch64)

Installation Method

git clone

Steps to Reproduce

EvaluationReport.to_file() (src/strands_evals/types/evaluation_report.py:324-325) is the sibling of Experiment.to_file() and has the same bug family that #380 reported and #383 fixes for experiments — plus one more:

import json
from strands_evals.types.evaluation_report import EvaluationReport

r = EvaluationReport(overall_score=float("nan"), scores=[float("nan"), 1.0],
                     cases=[{"name": "a"}], test_passes=[True, False])
r.to_file("report.json")   # "succeeds"

txt = open("report.json").read()          # contains:  "overall_score": NaN
json.loads(txt, parse_constant=lambda c: (_ for _ in ()).throw(ValueError(c)))  # ValueError: NaN

Reproduced on main @ 166aa31 (and unchanged on #383's head 8268555 — correctly out of that PR's scope).

Expected Behavior

The report file is valid, strict, UTF-8 JSON — or the write fails loudly before touching the file.

Actual Behavior

The file contains the literal NaN, which Python's lenient json.loads reads back but any strict parser (JS JSON.parse, jq, serde, ...) rejects. overall_score: float and scores: list[float] make NaN reachable from ordinary aggregation (custom evaluators returning NaN), and the report is exactly the artifact meant to be machine-read downstream.

Additional Context

Found while reviewing #383. Same family, lower priority: LocalFileTaskResultStore.save (src/strands_evals/local_file_task_result_store.py:38-39) uses write_text(result.model_dump_json(indent=2)) with no encoding=, and pydantic also emits bare NaN.

Possible Solution

Mirror #383: serialize first, then write bytes —

data = json.dumps(self.to_dict(), indent=2, ensure_ascii=False, allow_nan=False).encode("utf-8")
with open(file_path, "wb") as f:
    f.write(data)

and extend the Raises: block accordingly. This would be the third call site of that exact pattern, so it may be the right moment to extract a shared strict-dumps helper.

Related Issues

#380, #382, #383

Activity

  1. added
    area-coreCore eval framework: Case, Experiment, task handler, evaluation data stores
    area-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usability
    bugSomething isn't working
    on Aug 27, 2026
  2. serenearyal commented on Sep 3, 2026

    @serenearyal

    I'd like to take this one.

    I reproduced all three points on current main:

    • to_file() calls json.dump with defaults, so a NaN score is written as bare NaN. Python's lenient json.loads reads it back, but a strict parser rejects it.
    • The file is opened with mode "w" before serialization, so a failure mid-write leaves a truncated file. In my run a valid 224-byte report was cut to 106 bytes of invalid JSON after a save with an unserializable object.
    • The missing encoding= is harmless today because ensure_ascii defaults to True, but it becomes a real bug once the writer moves to ensure_ascii=False like fix: make Experiment.to_file() reject non-strict JSON instead of writing invalid files #383 did, so it should be fixed in the same change.

    Plan: mirror #383 for EvaluationReport.to_file(). Serialize first with allow_nan=False, encode to UTF-8, write the bytes only on success, extend the Raises: docstring, and add tests like the ones added for experiments (NaN rejected, existing file left intact, unpaired surrogates rejected).

    The issue notes this is the third copy of the same pattern. I can keep it a plain mirror of #383, or pull the strict dump into a small shared helper used by both writers. Let me know which you prefer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesarea-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilitybugSomething isn't working

    Type

    Fields

    Language

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions