Checks
Strands Version
1.53.0
Strands Evals Version
main @ 166aa31
Python Version
3.13.14
Operating System
Amazon Linux 2023 (aarch64)
Installation Method
git clone
Steps to Reproduce
EvaluationReport.to_file() (src/strands_evals/types/evaluation_report.py:324-325) is the sibling of Experiment.to_file() and has the same bug family that #380 reported and #383 fixes for experiments — plus one more:
import json
from strands_evals.types.evaluation_report import EvaluationReport
r = EvaluationReport(overall_score=float("nan"), scores=[float("nan"), 1.0],
cases=[{"name": "a"}], test_passes=[True, False])
r.to_file("report.json") # "succeeds"
txt = open("report.json").read() # contains: "overall_score": NaN
json.loads(txt, parse_constant=lambda c: (_ for _ in ()).throw(ValueError(c))) # ValueError: NaN
Reproduced on main @ 166aa31 (and unchanged on #383's head 8268555 — correctly out of that PR's scope).
Expected Behavior
The report file is valid, strict, UTF-8 JSON — or the write fails loudly before touching the file.
Actual Behavior
The file contains the literal NaN, which Python's lenient json.loads reads back but any strict parser (JS JSON.parse, jq, serde, ...) rejects. overall_score: float and scores: list[float] make NaN reachable from ordinary aggregation (custom evaluators returning NaN), and the report is exactly the artifact meant to be machine-read downstream.
Additional Context
Found while reviewing #383. Same family, lower priority: LocalFileTaskResultStore.save (src/strands_evals/local_file_task_result_store.py:38-39) uses write_text(result.model_dump_json(indent=2)) with no encoding=, and pydantic also emits bare NaN.
Possible Solution
Mirror #383: serialize first, then write bytes —
data = json.dumps(self.to_dict(), indent=2, ensure_ascii=False, allow_nan=False).encode("utf-8")
with open(file_path, "wb") as f:
f.write(data)
and extend the Raises: block accordingly. This would be the third call site of that exact pattern, so it may be the right moment to extract a shared strict-dumps helper.
Related Issues
#380, #382, #383
Checks
Strands Version
1.53.0
Strands Evals Version
main @ 166aa31
Python Version
3.13.14
Operating System
Amazon Linux 2023 (aarch64)
Installation Method
git clone
Steps to Reproduce
EvaluationReport.to_file()(src/strands_evals/types/evaluation_report.py:324-325) is the sibling ofExperiment.to_file()and has the same bug family that #380 reported and #383 fixes for experiments — plus one more:json.dumpis called with defaults, soallow_nan=True: NaN/Infinity scores are written as bareNaN— not valid JSON per RFC 8259 §6.open(file_path, "w")passes noencoding=, so the file's bytes are locale-dependent on non-UTF-8 systems.Reproduced on main @ 166aa31 (and unchanged on #383's head 8268555 — correctly out of that PR's scope).
Expected Behavior
The report file is valid, strict, UTF-8 JSON — or the write fails loudly before touching the file.
Actual Behavior
The file contains the literal
NaN, which Python's lenientjson.loadsreads back but any strict parser (JSJSON.parse,jq, serde, ...) rejects.overall_score: floatandscores: list[float]make NaN reachable from ordinary aggregation (custom evaluators returning NaN), and the report is exactly the artifact meant to be machine-read downstream.Additional Context
Found while reviewing #383. Same family, lower priority:
LocalFileTaskResultStore.save(src/strands_evals/local_file_task_result_store.py:38-39) useswrite_text(result.model_dump_json(indent=2))with noencoding=, and pydantic also emits bareNaN.Possible Solution
Mirror #383: serialize first, then write bytes —
and extend the
Raises:block accordingly. This would be the third call site of that exact pattern, so it may be the right moment to extract a shared strict-dumps helper.Related Issues
#380, #382, #383