Checks
Strands Version / Strands Evals Version / Python Version / OS
main @ 234eae6 · repo checkout (not a released version) · Python 3.13 · Linux (aarch64)
Installation Method
pip
Steps to Reproduce
Run CorrectnessEvaluator in reference mode (i.e. with Case.expected_assertion set) over any trajectory where a tool call happens before the final answer, then print the prompt the judge receives. Reproduced against the repo's own captured span fixture, tests/strands_evals/mappers/fixtures/claude_live_spans.json.
Expected Behavior
The judge prompt's USER QUERY section contains the user's request, so the judge can grade the response against what was actually asked.
Actual Behavior
USER QUERY is empty:
USER QUERY:
AGENT RESPONSE:
New York is 89°F and Seattle is 72°F. The temperature difference is 17°F...
EXPECTED RESPONSE:
The agent delegated to a math specialist and computed the correct result (17).
Mechanism:
_extract_user_prompt (src/strands_evals/evaluators/evaluator.py:138-153) returns "" unless session_history[-1] is a non-list message with text content.
_extract_trace_level (src/strands_evals/extractors/trace_extractor.py:54-89) always appends the owned tool-execution list to previous_turns before building that turn's TraceLevelInput (see the append at :69-82).
So for any trace with a tool call before the final response, session_history[-1] is that tool-execution list rather than the UserMessage, and _extract_user_prompt bails to "".
Additional Context
Not specific to one caller, and not new: tests_integ/test_adk_eval.py:288-311 (merged, on main) already runs CorrectnessEvaluator in reference mode on a trajectory with a tool call before the final answer, so this is live today.
Impact: the judge is asked to verify a claim against a query it cannot see. It also never sees the tool trace in this mode (correctness_evaluator.py:150-162 renders only USER QUERY / AGENT RESPONSE / EXPECTED RESPONSE), so an expected_assertion that references how the agent worked — which tool it used, whether it delegated — is ungradable here. The likely failure direction is leniency (false pass) rather than false failure, which is presumably why it hasn't surfaced as a flaky test: it makes reference-mode CorrectnessEvaluator quietly weaker than its contract implies whenever tool use precedes the final turn.
Suggested fix (either): keep the UserMessage reachable after appending tool executions in _extract_trace_level, or have _extract_user_prompt walk back past a trailing tool-execution list to the nearest actual UserMessage.
Found while reviewing #353 — it explains why a compound expected_assertion ("delegated to X and computed Y") behaves differently under GoalSuccessRateEvaluator, which does see tool calls, than under CorrectnessEvaluator, which doesn't.
Filed by strandly-the-agent. Verify before acting on it — I can be wrong.
Checks
Strands Version / Strands Evals Version / Python Version / OS
main@234eae6· repo checkout (not a released version) · Python 3.13 · Linux (aarch64)Installation Method
pip
Steps to Reproduce
Run
CorrectnessEvaluatorin reference mode (i.e. withCase.expected_assertionset) over any trajectory where a tool call happens before the final answer, then print the prompt the judge receives. Reproduced against the repo's own captured span fixture,tests/strands_evals/mappers/fixtures/claude_live_spans.json.Expected Behavior
The judge prompt's
USER QUERYsection contains the user's request, so the judge can grade the response against what was actually asked.Actual Behavior
USER QUERYis empty:Mechanism:
_extract_user_prompt(src/strands_evals/evaluators/evaluator.py:138-153) returns""unlesssession_history[-1]is a non-list message with text content._extract_trace_level(src/strands_evals/extractors/trace_extractor.py:54-89) always appends the owned tool-execution list toprevious_turnsbefore building that turn'sTraceLevelInput(see the append at:69-82).So for any trace with a tool call before the final response,
session_history[-1]is that tool-execution list rather than theUserMessage, and_extract_user_promptbails to"".Additional Context
Not specific to one caller, and not new:
tests_integ/test_adk_eval.py:288-311(merged, onmain) already runsCorrectnessEvaluatorin reference mode on a trajectory with a tool call before the final answer, so this is live today.Impact: the judge is asked to verify a claim against a query it cannot see. It also never sees the tool trace in this mode (
correctness_evaluator.py:150-162renders onlyUSER QUERY/AGENT RESPONSE/EXPECTED RESPONSE), so anexpected_assertionthat references how the agent worked — which tool it used, whether it delegated — is ungradable here. The likely failure direction is leniency (false pass) rather than false failure, which is presumably why it hasn't surfaced as a flaky test: it makes reference-modeCorrectnessEvaluatorquietly weaker than its contract implies whenever tool use precedes the final turn.Suggested fix (either): keep the
UserMessagereachable after appending tool executions in_extract_trace_level, or have_extract_user_promptwalk back past a trailing tool-execution list to the nearest actualUserMessage.Found while reviewing #353 — it explains why a compound
expected_assertion("delegated to X and computed Y") behaves differently underGoalSuccessRateEvaluator, which does see tool calls, than underCorrectnessEvaluator, which doesn't.Filed by
strandly-the-agent. Verify before acting on it — I can be wrong.