Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,7 @@ Assertion:
| { type: "must_not_call_tool"; tool: string }
| { type: "tool_call_count"; tool?: string; op: Op; value: number }
| { type: "tool_call_order"; sequence: string[] }
| { type: "tool_call_collection"; tools: string[] } # order-insensitive multiset: each name must appear at least as many times as listed; extra/intervening calls are ignored

# ─── Text assertions (over the final assistant turn's text content)
| { type: "text_contains"; value: string; case_sensitive?: bool }
Expand Down
5 changes: 5 additions & 0 deletions docs/developer/TASKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,11 @@ Initial development queue to reach MVP for the three e2e projects under `e2e/`:

- **live tool-definition wiring is completely missing (all providers)** — Confirmed via live reproduction 2026-07-07 while doing Feature C's (ailly-evals project) OpenAI live-confirmation pass: `CompletionRequest` ([src/engine/engine.rs:24-27](../../src/engine/engine.rs)) has no `tools` field at all, and `RigEngine::complete` ([src/engine/rig_engine.rs](../../src/engine/rig_engine.rs)) hardcodes `tools: Vec::new()` on every outgoing `rig::completion::CompletionRequest`, for every provider — not an OpenAI-specific or streaming-specific gap (this was previously named only as one bullet inside "engine deferred decisions" below, with no live evidence; promoting it here now that it has real reproduction). Consequence: no live `ailly run` has ever been able to make a model emit a genuine native tool call, so `must_call_tool`/`must_not_call_tool`/`tool_call_count`/`tool_call_order`/`tool_call_collection` assertions can only ever be exercised today against a conversation whose tool_use blocks were pre-filled by hand or mined from elsewhere (e.g. the judge-calibration project's mined conversations) — never against something `ailly run` itself produced live. The `ContentBlock::ToolUse`/`AssistantContent::ToolCall` mapping in `rig_engine.rs` already exists and is unit-tested, so it is reachable once a real tool schema is actually sent; the gap is entirely on the outbound request side. **Reproduction:** a live conversation with a system message describing a `get_weather` tool and a user message explicitly asking the model to use it, run against `gpt-4o-mini`, returns plain prose with no tool_use content block and no `tool_call` in the raw API response. Also confirmed while investigating: `RigEngine::complete` only ever calls `self.model.completion(...)` (non-streaming), never `.stream()` — `ailly_two` does not use Rig's streaming path today at all, so any future research framing this as a streaming-specific risk is wrong; the blocker is upstream of streaming entirely. **Trigger:** a harness needs a live model to actually emit a tool call (as opposed to grading a pre-filled one), or a live-run confidence claim about `tool_call_*` assertions is made without this caveat.

- **ailly-skill-eval: recommend `tool_call_collection` by default** — Now that `tool_call_collection` exists ([src/content/evaluation.rs](../../src/content/evaluation.rs), [src/knowledge/assertions.rs](../../src/knowledge/assertions.rs)), update [skills/ailly-skill-eval/SKILL.md](../../skills/ailly-skill-eval/SKILL.md) and its `references/method.md` to recommend it over `tool_call_order` by default for new suites, per that design's own Alternatives section framing ("prefer the collection variant by default"; `tool_call_order` stays for genuine sequencing requirements). Named as deferred in `.ailly/developer/2026-07-06-A-ailly-evals/feature-f-report-stats/design.md`'s Summary rather than folded into that feature-step, because it touches a cross-cutting project skill the feature-step's own Specification does not otherwise read or write. **Trigger:** the next suite-authoring pass through `ailly-skill-eval`.

- **per-case paired-difference granularity** — `compute_comparison`'s new paired-difference test ([src/knowledge/report.rs](../../src/knowledge/report.rs)) reports one significance figure for the whole comparison, not one per `CaseComparison`. Not built now: a single case's `n` is typically a handful of assertions, too small for a test to have real power, and the parent brief's "wire it through, don't redesign the report UX" instruction argued for the minimal whole-comparison version first. **Trigger:** a real comparison's case-level detail turns out to need its own significance read. See `.ailly/developer/2026-07-06-A-ailly-evals/feature-f-report-stats/design.md` Summary.
- **round-trip serde test for `Assertion::ToolCallCollection`** — [src/content/evaluation.rs](../../src/content/evaluation.rs)'s test module has a `regression_suite_parses_round_trips_and_exposes_typed_variants` fixture covering every other `Assertion` variant's parse/serialize round trip; `ToolCallCollection` was not added to it. Natural Plan-phase follow-on for whichever change next touches that fixture. See `.ailly/developer/2026-07-06-A-ailly-evals/feature-f-report-stats/design.md` Summary.

## Deferred

- **content/conversation deferred decisions** — Five revisit-after items carried over from the conversation design doc: typed `ImageSource`, narrowing `tool_result.content`, closing `TraceEvent` into a named enum, promoting `Conversation` to an aggregate root, and lifting blank/filled `Message` into type-states. See [TASK-NOTES-conversation-deferred.md](TASK-NOTES-conversation-deferred.md) for trigger conditions per item.
Expand Down
255 changes: 0 additions & 255 deletions plan.md

This file was deleted.

3 changes: 3 additions & 0 deletions src/content/evaluation.rs
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,9 @@ pub enum Assertion {
ToolCallOrder {
sequence: Vec<String>,
},
ToolCallCollection {
tools: Vec<String>,
},

TextContains {
value: String,
Expand Down
Loading
Loading