Skip to content

About

An eval suite for a customer-support agent in a regulated setting: durable Temporal workflows, an explicit state machine, guardrail classifiers and path-based scoring. Synthetic data.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

regulated-support-eval

An eval suite for a customer-support agent in a regulated setting: durable Temporal workflows, an explicit state machine, guardrail classifiers, path-based scoring and a per-conversation audit trail. All data is synthetic and written by us.

The mock agent follows the correct path in 79.8% of 104 scripted conversations. A local Qwen2.5 7B reaches 67.3% outcome accuracy but only 1.0% exact-path, because almost every path it takes is one tool call short; its paths are identical across runs only 60% of the time at temperature 0.7.

The finding worth your time is none of those. Plugging in a real model for the first time, after four sessions of building the eval, exposed seven defects at the model boundary — none of which 490 tests, mutation checks that provably bite, or a 104-conversation replay suite had ever touched, because the mock satisfied every contract by construction rather than by agreement.

Seven defects at the model boundary, found the first time a real model ran

The agent was driven by a hand-built keyword model for four sessions. The first time a real language model was plugged in -- a local Qwen2.5 7B through an OpenAI-compatible endpoint -- the model boundary turned out to have seven latent defects. None was caught by the 490 tests, by the mutation checks that provably bite, or by the 104-conversation replay suite.

defect what the mock did what a real model did
the prompt never listed the legal states used the State enum from code returned valid JSON with "next_state": "VERIFY_IDENTITY", a state that does not exist
the context was populated only by the mock filled customer_id itself with a regex proposed the right state every turn; the guard rejected it and the conversation sat in IDENTIFY forever
the prompt never named the tools called them by name from code walked the correct states and called nothing
tool_args is declared dict[str, str] only ever emitted strings emitted 2, null, true; deserialisation failed inside the workflow as "Failed decoding arguments", far from its cause
the prompt never contained the procedures knew all four from its own keyword table at CLASSIFY_INTENT the word "procedure" appeared zero times; it chose among ten state names with no description of any
guardrail states were offered as legal moves never proposes one; they are override-only self-routed into OUT_OF_SCOPE_REDIRECT and FRAUD_ROUTE, which the scorer counted as spurious guardrail fires
a tool attaches to the move into a state encoded that in its handlers emitted the right tool one turn late, on arrival; the move was rejected and the tool never ran

Three further setup bugs were hit before any of these -- a missing optional package, a doubled provider prefix in the model id, and a pasted key whose trailing newline is an illegal HTTP header. Those are ordinary first-run friction and are not in the table; the seven above are the ones where the mock had quietly satisfied a contract that was never written down.

Every one produced confident, well-formed, wrong output -- the failure mode Gradient Labs calls failing successfully. And every one was invisible under the mock for the same reason.

The mock satisfied every contract by construction rather than by agreement, so no contract was ever tested. 490 tests and mutation checks that provably bite, and not one of them touched the seam where a real model plugs in.

This is also the honest answer to why this took four sessions. The first three built an eval harness that had never been run against the kind of thing it exists to evaluate.

Fixing the seven took the local model from 0% to 100% outcome accuracy on a happy-path subset. What remains after them is a model result, not a harness one: the instruction to call get_account on verification is in the prompt, and the model returns "tool": null anyway.

Scorecard

scorecard

Every number traces to a file in reports/, produced by python scripts/run_all.py --model mock --gate. Night-1 figures come from reports/night1_baseline.json, captured before anything changed.

night 1 night 3 source
Trajectory accuracy, exact path match 0.538 0.798 reports/trajectories.json
Partial credit, mean longest-common-subsequence 0.828 0.931 reports/trajectories.json
Outcome accuracy 0.750 0.913 reports/trajectories.json
Wrong tool calls, whole suite 4 4 reports/trajectories.json
Conversations sent for human review 58 of 104 32 of 104 reports/review_queue.md
Guardrail spurious fires 4 3
Guardrail mutations caught 8 of 8 8 of 8 reports/mutation.md
Trajectory mutation caught 1 of 1 1 of 1 reports/mutation.md
Conversation resumes after the worker is killed not tested 2.45s reports/durability.md

Guardrail classifiers, on a held-out split frozen in guardrails/datasets/holdout.txt and pinned by a test. Night 3's split is a harder test set than night 1's: 62 to 65 per cent hand-written against 26 to 35 per cent, so the F1 columns are not strictly comparable. Hand-written recall is the number to read.

classifier F1 night 1 F1 night 3 hand-written recall night 1 hand-written recall night 3 active
regulated_advice 0.966 0.918 0.667 0.750 rule
out_of_scope 0.710 0.560 0.167 0.238 model
vulnerable_customer 0.824 0.800 0.100 0.500 model
fraud_topic 0.720 0.755 0.214 0.429 rule

Two of the four are short of the 0.60 target, and one moved the wrong way: out_of_scope hand-written recall fell from 0.667 to 0.238 when guardrails moved to a conservative calibrated firing threshold of 0.85. That is the trade the operating point buys — fewer wrong fires, more misses — and the system metrics above went up while this one went down. It is reported here rather than in a footnote because a reader comparing only the classifier columns would reasonably conclude the repository got worse. reports/guardrails.md lists exactly what they miss and why a keyword rule cannot reach it.

The hosted-model column is absent: EVAL_MODEL, EVAL_MODEL_BASE_URL and EVAL_MODEL_API_KEY were unset, so that path has still never run. NOTES/progress.md says so rather than the README implying otherwise.

Recovery after a spurious detour is used below and is our own metric, not a standard one: of the conversations where a guardrail fired that the expected path does not contain, the share that still reached their expected outcome. It answers "when a guardrail is wrong, can the conversation still finish?"

Mock against a local 7B, on the same 104 conversations

mock (keyword table) local Qwen2.5 7B
Exact-path accuracy 0.798 0.010
Partial credit, mean LCS 0.931 0.617
Outcome accuracy 0.913 0.673
Missing tool calls 0 175
Spurious guardrail fires 3 3
Latency p50 / p95 per turn sub-millisecond 3686 ms / 5976 ms

Read the first three rows together. Exact-path accuracy collapses to 0.010 while partial credit stays at 0.617 and outcome accuracy at 0.673. The cause is a single systematic omission: 175 missing tool calls, almost all of them get_account, which the model declines to call even though the prompt names it. Outcome scoring waves that through; a metric that scores the path does not. That is the argument for trajectory accuracy over final-answer scoring, made against our own data rather than asserted.

Latency is a local 7B on a laptop with a 606-token prompt. It is not comparable to a hosted small model on production infrastructure, and no comparison is drawn here.

Path consistency, and why one of these numbers would have been a lie

sampling consistency what it measures
temperature 0, fixed seed 100% that the harness is reproducible — greedy decoding repeating itself
temperature 0.7, seed unpinned 60% the model's actual path stability

At temperature 0.7, 8 of 20 scripts took a different path across three runs, and outcome accuracy moved 0.45, 0.50, 0.55 between them. Measured on a 20-script subset, not the full suite. Source: reports/local_model_temp07.json.

The first row is not a result and is shown only to explain why it is not one. This repository reported consistency as "not measured" for four sessions because the mock is deterministic; the first hosted run reproduced that mistake by sending temperature=0 with a fixed seed, which makes three identical runs the only possible outcome. Quoting that 100% would have been a claim about the harness dressed as a claim about the model. Its one legitimate use is the reason the partial-credit figure above needed only a single run rather than three.

Three commands

bash scripts/demo_durability.sh          # kill -9 a worker mid-conversation; it finishes anyway
bash scripts/demo_deploy_gate.sh         # a one-line change to the state machine, caught
python scripts/inspect.py card_lost_stolen-happy-001 --html out.html

The durability test behind that demo runs in CI as a separate, non-gating job. It has failed roughly one cold full-suite run in four. One real cause was found and fixed — a Temporal query during worker handover raises rather than returning a stale value — which reduced but did not eliminate it. It is split out rather than retried green or deleted, so its result stays visible; see the limitations for what has been ruled out.

The full suite, which is the same gate CI runs:

pip install -r requirements.txt
python scripts/run_all.py --model mock --gate

It exits non-zero if any gate in eval.config.yaml is not met.

Four things the evals caught that reading the code did not

A mutation check that could not fail. The trajectory mutation removed a transition from the state machine and replayed: baseline 43.5 per cent, mutated 43.5 per cent, zero scripts affected. Temporal runs workflow code in a sandbox that re-imports modules, so the mutation never reached the workflow. Run unsandboxed, accuracy fell to 0.0.

A vulnerability guardrail that fired on "I lost my card." A cue of lost my, added while widening the classifier, matched the commonest request a bank sees. The held-out split missed it because no example of that shape existed; an end-to-end replay caught it.

A close that was not a close. After a guardrail detour, a customer saying "OK, so back to what I called about" was read as ending the conversation, so 24 procedures never called their tool. Grouping failures by kind made it obvious; counting them did not.

Two merged fixes that did nothing at all. A guardrail override recorded which procedure it interrupted, and a separate fix set a one-time vulnerability flag. Both wrote to a context object the workflow rebuilt from a dict each turn and never wrote back, so both were discarded silently. Every test passed, review would have passed them, and neither moved a single number — which is how they were caught. A change that is supposed to alter behaviour and alters no metric has either no effect or no coverage, and both answers are worth knowing.

The finding: precision is worth four times recall here, and the gate was measuring the wrong thing

Three implementations of each guardrail were measured on a frozen holdout, then replayed through all 104 conversations. On the holdout the learned classifiers win every cell — hand-written recall, rule then TF-IDF then sentence embeddings:

classifier rule TF-IDF embeddings
regulated_advice 0.750 1.000 1.000
fraud_topic 0.429 0.905 0.952
vulnerable_customer 0.450 0.900 0.900
out_of_scope 0.190 0.667 0.762

fraud_topic and regulated_advice run on rules at 0.429 and 0.750 hand-written recall, below the 0.60 target for the first, even though TF-IDF reaches 0.905 and embeddings 0.952 on the same holdout. That is deliberate: every configuration that switched them to a learned classifier lost more trajectory accuracy in the conversation loop than it gained in recall, and the loop is what ships. The 21-configuration sweep behind that choice is in reports/cost_of_a_false_fire.md, and the misses it costs are listed under Known misses in reports/guardrails.md.

In the conversation loop they lose almost every configuration. Regressing conversations correct against guardrail behaviour across eight replayed configurations gives the reason:

cost per event R²
a guardrail that fires when it should not −0.87 conversations 0.92
a guardrail that fails to fire −3.82 conversations 0.46

The first is nearly one-for-one and the fit is tight. Checked directly rather than by regression: 4 of 4 conversations with a spurious fire failed, and none recovered.

The mechanism is in the transition table. FRAUD_ROUTE leads only to ESCALATE, so a spurious fraud fire ends the conversation. The other guardrail states return to CLASSIFY_INTENT, but the customer's next turn is a mid-procedure answer — a bare transaction reference — which intent routing cannot place, so the conversation looped until it ran out of turns. Three of the four failures were that stall.

Two changes followed, and both came from the measurement rather than from taste:

The agent now remembers the procedure a guardrail interrupted. No transition was added; the route back already existed. What was missing was memory. Finding this exposed a worse bug: decide() mutated a throwaway copy of the conversation context, so the write never reached the workflow. Two earlier fixes had been merged, changed no number, and were dead on arrival.

Guardrails now fire at a calibrated 0.85, not the default 0.5. The isotonic calibrators fitted on night 3 were being used for the review queue but not for the decision. Sweeping the operating point against conversations rather than against F1:

calibrated threshold trajectory accuracy fired wrongly missed
0.5, the library default 0.769 7 10
0.7 0.779 7 9
0.85 0.779 3 10
0.95 0.750 2 13

Embeddings do not survive this either. At 0.85 they fire zero times and miss 25: their calibrated probabilities are too compressed to separate at any threshold, which is a better description of the problem than "they over-fire."

And then the repository's own deploy gate failed. Moving to a conservative operating point drops holdout F1 for out_of_scope from 0.800 to 0.560, and the gate had a per-classifier F1 floor of 0.70. The gate would have blocked the change that improved the system. So the gate changed: the trajectory floor was tightened from 0.50 to 0.70, a direct cap on spurious fires was added, and the F1 floor was lowered to 0.50 and demoted to a report rather than a guard. That is the whole finding applied to this repository itself — an eval suite that gates on a classifier metric will block the work.

How this maps to Gradient Labs' published architecture

Each claim below carries the page it came from. Sources on gradient-labs.ai and blog.gradient-labs.ai are theirs; the OpenAI and Temporal case studies are third parties writing about them. NOTES/recon.md holds the full working.

Conversations are Temporal workflows, one per conversation, because that is how they describe their own system: "We represent each conversation in our platform as a Temporal Workflow" (building agentic workflows), and model, tool and classifier calls are activities, since "each of these skills is a Temporal Activity" (same page); they name the underlying property directly, "we use Temporal, a durable execution system" (building resilient agentic systems). The orchestrator is an explicit state machine because their agent is "built as a finite state machine" (safe AI agents in high-stakes industries), and the four flows are called procedures after "Procedures, the natural language instructions our agents follow" (headless AI agents). Guardrails run on every user turn, matching "20+ financial-services-specific guardrails on every turn of every conversation" (lending agent), and a firing guardrail re-routes the conversation to its own state because "the agent can be re-routed to a guardrail-specific procedure" (agent and customer guardrails). That same page splits guardrails into two suites, ones that "inspect what the customer is saying" and "a second suite of guardrails that inspect answers the AI agent is drafting", which is why the regulated-advice classifier also runs over the model's draft reply; the framing they give it is "Safety is a two-sided problem" (lessons from deploying AI agents in banking). The audit trail records the proposed transition next to the accepted one because "Every decision returns events for observability" (ZenML's summary of their talk, third party), and the aim is theirs: "That record turns an audit from an investigation into a lookup" (Collaborate). The synthetic bank exists because of a problem they name, evaluating without real tool access, since sensitive customer APIs cannot be called during an eval (same ZenML page). The scorecard has three panels because the Senior AI Engineer role asks for "structured evals measuring accuracy, safety, and latency" (job description).

Two words here are borrowed rather than theirs, and one idea is not theirs at all. Trajectory accuracy and replay reach us through OpenAI writing about them, which attributes rather than asserts: "what they call trajectory accuracy: whether the system follows the correct path", and "the team replays real customer conversations and compares the system's behavior against the expected procedure" (OpenAI case study, third party). Neither word appears in any of the 27 posts of theirs read for this repo; their own vocabulary is test cases, test suites, simulations, synthesised conversations and scorers. Mutation checking is not theirs in any form. Nothing read describes deliberately breaking a component to prove its eval fails; that idea is this repository's, and it is the part most worth arguing with.

Limitations

These are the reasons not to trust the scorecard, in the order they should worry you.

The review queue is a third of the suite, and cannot easily be smaller. It selects 32 of 104 conversations, down from 95 on night 2. Isotonic calibration fitted on the training split took classifier_uncertain from 91 conversations to 4, and Brier scores improved for every classifier and implementation; reports/calibration.md has the reliability tables. What remains is not noise: 30 of 104, or 28.8 per cent, are conversations that a mandatory rule selects — every trajectory mismatch, every escalation, every wrong tool call. The queue cannot go below that floor while trajectory accuracy is 0.769, and the remaining five are the deliberate 5 per cent random sample.

The queue's own precision is measured against our labels, not a human's. 20 of the 36 queue items carry a verdict, all of them assigned by this project and marked source: synthetic-reviewer in reports/review_verdicts.jsonl. On those 20, overall precision is 0.85: trajectory_mismatch, wrong_tool_called and classifier_uncertain each flagged only conversations that were genuinely mishandled, while random_sample came out at 0.50, which is what a random sample should look like. Read that as a sanity check on the reason codes, not as a validated accuracy figure.

The hosted-model path has still never run. Three nights, three times the environment variables were unset. The openai-compatible provider, the call cap, the backoff and the key-leak test described in the brief do not exist, because writing them untested would be worse than their absence. NOTES/progress.md records this each night.

The sentence-embedding classifier is optional and off by default. It needs requirements-optional.txt, which pulls torch and downloads a MiniLM model. Nothing in CI, pytest, or run_all.py --model mock requires it; with it absent the registry falls back to rules and says so on stderr.

The gate thresholds were set after measurement, as regression guards, and one was deliberately loosened. They sit just below what the system currently does, so the build fails when the agent gets worse than it is today. The per-classifier F1 floor was then lowered from 0.70 to 0.50 and demoted from a guard to a report, because a conservative firing threshold lowers holdout F1 while raising trajectory accuracy, and the old gate would have blocked that change. To keep the protection net positive, the trajectory floor was tightened from 0.50 to 0.70 and a direct cap on spurious fires was added. Loosening a gate to admit your own change is the classic way to turn a gate into a rubber stamp, so judge that swap on whether the replacement guards are tighter, which is the claim being made. Raise all of them as the agent improves; do not treat them as targets already met.

Two classifiers miss the hand-written recall target. fraud_topic at 0.429 and vulnerable_customer at 0.450, against a target of 0.60. Cues were not widened further to reach it, because every remaining candidate was either memorising one training example or broad enough to fire on ordinary banking traffic. reports/guardrails.md lists what they miss under Known misses.

The headline F1 is carried by the easier half of the test set. The labelled examples come in two kinds: hand-written and generated from templates with slot substitution. The templated ones are structurally regular, and the classifiers do far better on them. On the held-out split, recall on hand-written positives against templated positives:

classifier hand-written templated
regulated_advice 0.750 1.000
out_of_scope 0.667 1.000
vulnerable_customer 0.450 0.923
fraud_topic 0.429 0.917

The test split is majority templated, so the F1 in the scorecard overstates how these rules would do on free text. The hand-written recall is the more honest number, and it is poor. This is as much a property of the dataset as of the classifiers.

The active classifiers are mostly keyword rules. Three of the four are transparent rule classifiers, chosen because they win in the loop, not because they are the strongest classifiers available. Their precision on the holdout is 1.000, which mostly means they are conservative. out_of_scope runs the offline TF-IDF model.

The mock model is deliberately imperfect. It carries three named flaws, so trajectory accuracy could never reach 1.0 by construction. Without them the number would be a decoration. See agent/model.py.

The data is synthetic and small. 104 conversations, 880 labelled examples across four classifiers of which roughly half are hand-written, one invented bank. Nothing here has met a real customer, and no claim about production behaviour follows from any of it.

Consistency is not measured. With the mock model the agent is deterministic, so every script runs once and there is no variance to report. Running each script several times would only measure the same answer repeatedly.

One agent-side guardrail, not a suite. Only regulated-advice inspects the model's draft reply. Their published design has a second full suite; this has one classifier.

The agent reads identifiers only in its first state. If a customer gives their account number later in a conversation, it is ignored. A real agent should absorb identifiers from any turn.

Latency is measured, but means little here. Sub-millisecond figures are the cost of a keyword match and a dictionary lookup, not of a model call. The panel exists so the axis is present and honest, not because the number is interesting.

Method note

The labelled examples were written from the definitions by a pass that never saw the classifier rules, and the first version of the rules was written by a pass forbidden to open the data. The rules were then widened against the training split only; the published numbers are on the held-out split, which was never read during development. The train column in the scorecard is there so the gap is visible: a large positive gap means the rules memorised their training examples. reports/guardrails.md prints it per classifier.

The expected path of every scripted conversation is a specification of correct behaviour, derived from the procedure definitions. None of it was produced by running the agent and recording what it did. Had it been, trajectory accuracy would be 1.0 by construction.

Prior guardrail work

apexive/odoo-llm#264 — tool-consent enforcement gate, 13 tests, mutation-checked.

Licence

MIT. See LICENSE.

About

An eval suite for a customer-support agent in a regulated setting: durable Temporal workflows, an explicit state machine, guardrail classifiers and path-based scoring. Synthetic data.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages