Skip to content

Latest commit

 

History

History
203 lines (144 loc) · 8.99 KB

File metadata and controls

203 lines (144 loc) · 8.99 KB

Testing pi-loop

Setup

git clone https://github.com/trvon/pi-loop.git
cd pi-loop
npm install
npm run hooks:install

Contributions use focused branches and signed, thematic commits. Open the first pull request against master; dependent pull requests target their predecessor branch. Run the local gate at every PR tip and relevant checks at each intermediate commit. Push normally with hooks; do not rewrite published stack history. Contributions are MIT licensed.

Local gate

Run before opening or merging a code change:

npm run lint
npm run typecheck
npm test
npm run build
npm run test:package
npm audit --audit-level=moderate
git diff --check

Lint currently reports two established optional-chain warnings; new warnings are not accepted.

Test layers

Layer Command Purpose
Unit/integration npm test Reducers, stores, runtimes, tools, session wiring
Coverage npm run test:coverage Enforced statement/branch/function/line thresholds
Property npm run test:property Generated state-machine invariants
Fuzz campaign npm run test:fuzz 10,000 generated cases
Package smoke npm run test:package Tarball contents and public imports
Live E2E npm run test:e2e Pi RPC/model conformance
Benchmarks npm run bench Fixed core workloads
CPU profile npm run profile:core V8 profile plus checksummed metadata

Unit tests use Vitest, fake timers for schedules, in-memory stores for pure behavior, temporary files for persistence/restart, and one shared Pi event-bus mock.

Harness review campaigns

See bounded review campaigns for the test-only coverage/closure reference policy and traceable boundary inventory. Failed or skipped mandatory discovery is incomplete, never a clean review. The inventory checks references; the normal suite executes the underlying regressions.

Workflow coverage

The workflow suites prove:

  • creation and embedded initial execution;
  • lease claim, renewal, expiry, and foreign-owner rejection;
  • atomic settlement and unowned destination activation;
  • retry attempts, evidence, terminal behavior, and cadence limits;
  • typed add/revise/redirect definition patches;
  • immutable revision history and current-execution preservation;
  • revision/revision and revision/transition CAS races;
  • legacy normalization and malformed-history rejection;
  • revision-aware LoopList and wake guidance;
  • no workflow-owned TaskStore records.

Primary files:

test/workflow-reducer.test.ts
test/loop-reducer.test.ts
test/store.test.ts
test/loop-tools.test.ts
test/workflow-task-integration.test.ts
test/injection.test.ts
test/property/workflow.property.test.ts

Subagent orchestration coverage

The orchestration suites prove finite-batch bounds, capacity reservation before spawn, reducer/store CAS immutability, lifecycle identity, early-start replay, bounded settlement before consume, retry/uncertainty policy, durable wake acknowledgement, session teardown, cancellation, scheduler exclusion, tool scope gating, compact presentation, and extension wiring.

Primary files:

test/orchestration-reducer.test.ts
test/orchestration-store.test.ts
test/orchestration-runtime.test.ts
test/orchestration-tools.test.ts
test/property/orchestration.property.test.ts

Deterministic integration tests own timeout ambiguity and exact lifecycle ordering because upstream lacks status/list and idempotent dispatch keys. The opt-in live scenario additionally validates the supported protocol-v2 spawn, settlement, consume, and active-cancellation path against a real provider.

Task-backlog coverage

A backlog wake must call TaskList, inspect TaskGet, claim or resume one task, perform work, run observable validation, and settle the task in the same turn. Tests reject status-only progress, missing claims, reasoning-only validation, and deferral to a later wake.

Description-declared prerequisites are followed through TaskGet; TaskStore has no dependency-edge field.

Live E2E

Live tests are opt-in and run in isolated temporary workspaces. Stateful workflow and backlog scenarios use project scope; orchestration uses default file-backed session scope.

PI_LOOP_LIVE_MODEL="openai-codex/gpt-5.6-sol:minimal" npm run test:e2e

Workflow scenarios:

PI_LOOP_LIVE_SCENARIO=retry npm run test:e2e:workflow
PI_LOOP_LIVE_SCENARIO=phases npm run test:e2e:workflow
PI_LOOP_LIVE_SCENARIO=evolution npm run test:e2e:workflow
  • retry: one embedded phase repeats once and completes.
  • phases: at least three embedded phases advance with evidence.
  • evolution: investigation calls WorkflowRevise with add_state, redirect_transition, and revise_state, then follows the revised path.

Every workflow scenario requires explicit phase claims, terminal completion, no standalone task/loop/monitor mutations, and an empty TaskStore.

Backlog scenario:

PI_LOOP_LIVE_MODEL="openai-codex/gpt-5.6-sol:minimal" npm run test:e2e:backlog

The backlog scenario requires this first-run sequence:

TaskList → TaskGet → TaskClaim → write/edit → shell validation → TaskUpdate completed

Controller-routing scenarios:

PI_LOOP_LIVE_ROUTING_MODELS="openai-codex/gpt-5.6-sol:minimal,anthropic/claude-sonnet-4-6" \
npm run test:e2e:routing

Each model/scenario pair runs in a clean temporary Pi process with only WorkflowCreate, TaskCreate, and LoopCreate exposed. Natural-language prompts never name those tools. Critical success requires the exact controller type/count and successful tool validation; payload semantics and first-turn completion contribute to accuracy. Use PI_LOOP_LIVE_ROUTING_SCENARIOS for a comma-separated subset. See controller routing evaluation for the fixed baseline/hold-out matrix and judgment rules.

Subagent orchestration scenario:

PI_LOOP_LIVE_MODEL="openai-codex/gpt-5.6-sol:minimal" \
PI_LOOP_LIVE_SUBAGENTS_EXTENSION="$HOME/.pi/agent/npm/node_modules/@tintinweb/pi-subagents/src/index.ts" \
npm run test:e2e:orchestration

It runs in a temporary workspace with default file-backed session scope, proves two isolated workers settle with durable consumed evidence, then creates and cancels an active worker. The extension path defaults to the standard user package location when PI_LOOP_LIVE_SUBAGENTS_EXTENSION is omitted.

Without PI_LOOP_LIVE_MODEL, live scripts exit with SKIP. Reports are bounded JSON under .artifacts/ and do not record credentials; claim IDs are redacted where applicable.

Property and fuzz replay

The fixed-seed property suite covers cron boundaries, reducer determinism, workflow transition/revision immutability, orchestration CAS/bounds/uncertainty, attempt limits, and file-backed task replay.

Override campaign size:

FC_NUM_RUNS=50000 npm run test:property

Replay a minimized failure with both values printed by fast-check:

FC_SEED=24301 FC_PATH='0:0' npm run test:property

Keep the generalized property and add the minimized example to the nearest ordinary test before fixing production code.

Benchmarks and profiles

npm run bench
npm run bench:baseline
npm run bench:compare
npm run profile:core

npm run bench prints a table of throughput, mean latency, and margin of error. bench:baseline writes those results to .artifacts/benchmarks/baseline.json, and bench:compare reports the throughput change of a fresh run against that baseline. Compare benchmarks only on the same machine, Node version, architecture, timezone, and power state. Profiles are written to .artifacts/profiles/; load the .cpuprofile in a V8-compatible viewer. Shared benchmark workloads live in benchmarks/workloads.ts.

Change evidence

Follow AGENTS.md for the testing contract:

  • Bug: behavioral regression fails on unchanged code, nearby control passes, same regression passes after repair.
  • Feature: acceptance expectations precede implementation; test the affected surface and applicable boundaries.
  • Invariant: independently test authority, ownership, rejection, lifecycle, bounds, and persistence properties that must remain unchanged.
  • Model-facing guidance: validate examples and misuse deterministically; report bounded live model/fixture/attempt evidence separately. A text match is not a model-behavior result.

Name the expectations and invariants in the PR body. Record exact red/green commands and limitations. Compilation, fixture, or provider failures do not establish a product defect; skipped live checks do not pass.

Change-specific minimums

  • Reducer/store mutation: focused reducer + persistence/restart + property tests.
  • Tool schema/copy: tool tests + test/tool-copy-budget.test.ts.
  • Session/wake lifecycle: session, notification, scheduler, and index integration tests.
  • Monitor behavior: manager, tool, onDone runtime, and package smoke tests.
  • RPC contract: copy vendored files to pi-orca, bump VENDOR_REV, and run both repositories' RPC suites.
  • Published surface: build, package smoke, and isolated tarball import.