Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
09b55eb
fix(ce-doc-review): create the run directory on every review and pers…
tmchow Aug 23, 2026
b262f2a
fix(ce-plan): record depth in unified-plan metadata
tmchow Aug 23, 2026
574aabb
fix(ce-doc-review): activate product-lens on an unsettled product pos…
tmchow Aug 23, 2026
896c20f
test(skill-eval-cell): product-lens activation probe rows and a must_…
tmchow Aug 23, 2026
7efdcb3
docs(plans): doc-review cost measurable plan
tmchow Aug 23, 2026
04fa65e
fix(ce-plan): keep the shared html-rendering reference identical acro…
tmchow Aug 23, 2026
bbcebf7
fix(review): scope roster grade terms to the TEAM line; orchestrator …
tmchow Aug 23, 2026
36bc421
test(ce-doc-review): drop the superseded run-directory slot pin
tmchow Aug 23, 2026
2362943
test(skill-eval-cell): run the A/B pair for any row that carries its …
tmchow Aug 23, 2026
7fa3f2d
refactor(skill-eval-cell): reuse lastTrailer for the TEAM roster scope
tmchow Aug 23, 2026
cafa7e2
refactor(skill-eval-cell): one hasBaseline predicate for the A/B deci…
tmchow Aug 23, 2026
d4c1b07
fix(ce-doc-review): persist any local return whose artifact file is a…
tmchow Aug 23, 2026
3d8da53
Revert "fix(ce-plan): keep the shared html-rendering reference identi…
tmchow Aug 23, 2026
882f1af
Revert "fix(ce-plan): record depth in unified-plan metadata"
tmchow Aug 23, 2026
f91d05b
fix(ce-doc-review): manifest records unit and line counts instead of …
tmchow Aug 23, 2026
553ec16
fix(ce-doc-review): keep only the product-lens change; drop review-ar…
tmchow Aug 23, 2026
c55368f
test(skill-eval-cell): a roster probe requires the TEAM trailer
tmchow Aug 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,131 @@
---
title: Product-Lens Activation Condition - Plan
type: fix
date: 2026-08-22
artifact_contract: ce-unified-plan/v1
artifact_readiness: implementation-ready
product_contract_source: ce-plan-bootstrap
execution: code
---

# Product-Lens Activation Condition - Plan

## Goal Capsule

- **Objective:** `ce-doc-review`'s product-lens reviewer runs only when a plan stakes a product position worth a product judgment; routine plans no longer pay for it.
- **Means:** restate the product-lens activation leg in `references/persona-selection.md` as one condition (KTD1), and probe the condition with eval-cell rows that can fail (KTD2).
- **Authority:** this plan; the project's active instructions ("Working on Skills", "Reviewing a skill change", "Right-size new mechanical guards"); the repo-local `ce-skill-work` standard for every edit under `skills/`; `docs/plans/2026-08-22-0934-fix-right-size-skill-ceremony-plan.md` KTD6 (document review stays mandatory on a Lightweight Durable plan).
- **Execution profile:** two units, one branch (`tmchow/doc-review-right-size-a`), one PR to `main`. U2 depends on U1.
- **Stop conditions:** stop and surface if the activation probe shows the restated condition under-fires on a document that states a challengeable product position, or if the routine fixture still announces product-lens after a three-trial rerun — surface it; do not tighten the condition inside this PR.
- **Tail ownership:** the executor owns verification and the PR; the PR carries the probe evidence.

---

## Product Contract

### Summary

One change to one skill. `ce-doc-review`'s product-lens persona activates when the document stakes a product position a stakeholder could challenge that no upstream Product Contract settled, or when the work carries strategic weight. It no longer activates on "solution selection where alternatives plausibly exist", which holds for nearly any fix. The adversarial reviewer, the cross-model pass, the other personas, and the plan template do not change.

### Problem Frame

PR #1514 right-sized ceremony at intake; a small request that still lands as a Durable plan pays for document review in full. Product-lens's premise leg listed "solution selection where alternatives plausibly exist", which is true of almost any plan with a decision in it, so a product-judgment persona — and, through the judgment-trio gate, the cross-model pass — ran on routine bootstrap plans. Persisting review artifacts for a measurement pass and recording plan depth were considered and dropped: the measurement served plugin maintainers, not users, and a recorded depth label invites consumers to route on it instead of reasoning about the document.

### Requirements

- R1. `ce-doc-review`'s product-lens persona activates when the document stakes a product position — what to build, why, or what comes first — that a knowledgeable stakeholder could reasonably challenge and that no upstream Product Contract settled, or when the work carries strategic weight. A choice among mechanisms for building an agreed outcome is not a product position; describing a task or restating known requirements is not either.
- R2. The always-on roster, design-lens, security-lens, scope-guardian, adversarial activation, and the cross-model pass behave exactly as before.
- R3. The condition is probed by eval-cell rows that can fail in both directions: a routine plan must not name product-lens in its declared team, and a plan with a staked position must.

### Key Decisions

- **Adversarial activation and the cross-model pass stay as they are** (session-settled: user-directed — chosen over gating the pass on provenance: a different-family peer catches model-specific consistent errors a same-model reviewer cannot, and the adversarial reviewer has found real findings). Governs R2.
- **Document review stays mandatory on Lightweight Durable plans** — inherits KTD6 of the #1514 plan (session-settled: user-approved). Governs R2.
- **No persisted review artifacts and no depth label** (session-settled: user-directed — chosen over a per-review run directory with manifest and per-persona files, and over `depth:` in plan metadata: the artifacts served maintainers' measurement at every user's expense, and a recorded classification tempts consumers to route on it). Governs the scope below.
- **Verification is proportionate: a pin plus a small activation probe** (session-settled: user-approved — chosen over a three-host matrix or TUI sessions). Governs R3.

### Scope Boundaries

- In: `skills/ce-doc-review/references/persona-selection.md` (product-lens block), `docs/skills/ce-doc-review.md`, the contract test, and eval-cell rows, fixtures, and grader support.
- Out: the plan template and floor; `skills/ce-plan/`; any run-directory or artifact persistence in `ce-doc-review`; `product-lens-reviewer.md`'s own technique suppression (pre-existing, keyed on origin presence; named as follow-up).

---

## Planning Contract

### Key Technical Decisions

- KTD1. **The premise leg is one condition with provenance inside it.** A plan that derives from a validated upstream Product Contract stakes no new position unless it contests what the origin settled; adversarial already carries that provenance rule in the same file. Without it the restated condition fires on any plan with a KTD and pulls the cross-model pass with it on every brainstorm-sourced plan. Governs R1.
- KTD2. **The probe grades the declared team, not narration.** The eval cell's `must_exclude` reads only the ACTIONS trailer, so a roster probe needs a final-answer negative term; that term and `must_include` read the run's trailing `TEAM:` line when present, because narration or an output path can contain a persona's name. A row with its own `baseline_ref` gets an A/B pair. Governs R3.
- KTD3. **Fixtures turn the cross-model pass off.** The probe grades one dimension; the fixture's checkout config sets `cross_model_review_mode: off` so no trial egresses. Governs R3.

### Patterns to Follow

- Eval rows: the `ce-plan` right-size rows in `tests/skill-eval-cell/catalog.ts` (`files_read_post`, `must_include`, `baseline_ref`).
- Trailer parsing: `lastTrailer` in `tests/skill-eval-cell/grade.ts`.

---

## Implementation Units

### U1. Restate the product-lens activation condition

- **Goal:** product-lens activates on a staked, unsettled product position or strategic weight; not on plausible alternatives.
- **Requirements:** R1, R2, KTD1
- **Dependencies:** None
- **Files:**
- Modify: `skills/ce-doc-review/references/persona-selection.md` (product-lens block)
- Modify: `docs/skills/ce-doc-review.md` (product-lens bullet)
- Test: `tests/pipeline-review-contract.test.ts` (pin: the product-lens block contains the unsettled-position condition and the mechanism exclusion, does not contain "alternatives plausibly exist", keeps the strategic-weight leg; the existing security-lens pin keeps passing)
- **Approach:** replace the premise-claims enumeration with the condition in R1, keeping the strategic-weight leg; bring the block to the `ce-skill-work` standard.
- **Test scenarios:**
- The new pin fails on `main` and passes after the edit.
- The security-lens pin is unchanged.
- **Verification:** pins pass; U2 supplies the behavioral evidence.

### U2. Activation probe for product-lens

- **Goal:** Evidence that the restated condition stops the routine case and keeps both positive legs, on Claude and Codex.
- **Requirements:** R3, KTD2, KTD3
- **Dependencies:** U1
- **Files:**
- Create: `tests/skill-eval-cell/fixtures/doc-review-routine-fix/` (a captured real bootstrap fix plan), `doc-review-settled-origin/` (a captured real brainstorm-sourced plan with settled decisions), `doc-review-staked-position/` (explicit prioritization and an outcome prediction), `doc-review-strategic-weight/` (sound premise, strategic weight, no new contested position); each with `.compound-engineering/config.yaml` setting `cross_model_review_mode: off`
- Modify: `tests/skill-eval-cell/catalog.ts` (four `ce-doc-review` rows, `mode:non-interactive`, `git_init: true`, `timeout_secs` sized for a multi-subagent review, `baseline_ref` at the #1514 merge; `must_not_include` on the `Grade` type), `tests/skill-eval-cell/grade.ts` (`must_not_include`; roster terms scoped to the `TEAM:` trailer), `tests/skill-eval-cell/grade.test.ts`, `tests/skill-eval-cell/pack.ts` (A/B for rows with `baseline_ref`), `tests/skill-eval-cell/catalog.test.ts` (required-read allowlist entries)
- **Approach:** each row's task asks the run to end with a `TEAM:` line; routine and settled-origin rows `must_include` the always-on pair and `must_not_include: ["product-lens"]`; staked-position and strategic-weight rows `must_include` `product-lens`. Grade only the product-lens dimension; adversarial activates on every bootstrap fixture by its provenance rule and is not graded.
- **Execution note:** at least one fixture is a captured real plan rather than an authored one (`docs/solutions/skill-design/authored-eval-corpora-contain-the-happy-path.md`).
- **Test scenarios:**
- Pre arm: routine fixture names product-lens; post arm: it does not.
- Staked-position and strategic-weight fixtures name product-lens in both arms.
- Settled-origin fixture: product-lens absent post.
- `catalog.test.ts` and `grade.test.ts` pass with the new rows, allowlist entries, and grader term.
- **Verification:** `bun run test:skill-eval-pack -- --id <row> --arm ab --hosts claude,codex`; results recorded in the PR.

---

## Verification Contract

| Gate | Command | Proves |
|---|---|---|
| Deterministic | `bun run test` | U1 pin, U2 catalog guards and grader tests |
| Release | `bun run release:validate`, `bun run plugin:validate` | plugin and marketplace consistency |
| Behavioral | the four `ce-doc-review` rows on Claude and Codex | activation flips on the routine fixture, stays off on settled-origin, holds on staked-position and strategic-weight |

Conflict call-out on the probe size: `docs/solutions/skill-design/ce-doc-review-calibration-patterns.md` records that single runs of an activation change are noise and sets N=3 per cell as the floor. The session settled one trial per cell as proportionate. If a post-arm result ties with pre on the routine fixture, rerun that cell to three trials before reading it as "no effect".

Not covered, by decision: a roster-regression sweep of the untouched lenses (the security-lens pin stays; the other lenses' text is outside the edited block).

---

## Definition of Done

- R1-R3 hold; every gate in the Verification Contract passes at the PR head.
- `skills/ce-doc-review/` differs from `main` only in the product-lens block; `skills/ce-plan/` is untouched.
- U2 rows and fixtures are committed with their results in the PR body.
- No abandoned experiment files remain under `tests/skill-eval-cell/fixtures/`.

## Sources & Research

- `skills/ce-doc-review/references/persona-selection.md` (adversarial's provenance rule), `references/personas/product-lens-reviewer.md` (technique suppression on origin)
- `tests/skill-eval-cell/grade.ts` (`lastTrailer`), `catalog.ts`, `pack.ts`
- `docs/solutions/skill-design/paired-old-vs-new-injection-skill-evals.md` (grade one dimension), `authored-eval-corpora-contain-the-happy-path.md`, `ce-doc-review-calibration-patterns.md` (variance)
- Cross-model panel and literature this session: self-preference bias and model-specific consistent errors (Panickssery et al. 2024; Self-Correction Bench 2025; "Too Consistent to Detect" 2025); same-model fresh-context review recovers part of the gap (Cross-Context Review 2026)
2 changes: 1 addition & 1 deletion docs/skills/ce-doc-review.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ Document review is harder than code review in specific ways:

Conditional personas activate from what the doc says, not keyword matching:

- **product-lens** when the doc makes challengeable claims about what to build and why, or the work carries strategic weight
- **product-lens** when the doc stakes an unsettled product position — what to build, why, or what comes first — that a stakeholder could challenge, or the work carries strategic weight; a choice among mechanisms is not a product position
- **design-lens** when it contains UI/UX references, user flows, or visual design language
- **security-lens** when it touches auth, public APIs, sensitive data, payments, or third-party trust boundaries
- **scope-guardian** when it has multiple priority tiers, a large requirement count, or scope-boundary language that looks misaligned
Expand Down
4 changes: 2 additions & 2 deletions skills/ce-doc-review/references/persona-selection.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@ The team is `coherence-reviewer` and `feasibility-reviewer` always, plus each co

Activate a conditional persona when the document shows its signals:

**product-lens** — the document makes challengeable claims about what to build and why, or the work carries strategic weight beyond the immediate problem. Users may be end users, developers, operators, maintainers, or any other audience; the criteria are domain-agnostic. Either leg qualifies:
**product-lens** — the document stakes a product position — what to build, why, or what comes first — that a knowledgeable stakeholder could reasonably challenge and that no upstream Product Contract settled, or the work carries strategic weight beyond the immediate problem. Users may be end users, developers, operators, maintainers, or any other audience; the criteria are domain-agnostic. Either leg qualifies:

- *Premise claims* — the document stakes a position a knowledgeable stakeholder could reasonably challenge, not merely describing a task or restating known requirements: non-obvious or debatable problem framing; solution selection where alternatives plausibly exist (implicit or explicit); prioritization that explicitly ranks what gets built vs deferred; goal statements predicting specific user outcomes rather than restating constraints or listing deliverables.
- *Unsettled product position* — a problem framing, a goal predicting a specific user outcome, or a prioritization that ranks what gets built against what is deferred, which the document's origin did not already settle. A choice among mechanisms for an agreed outcome is an implementation decision, not a product position; describing a task or restating known requirements stakes nothing.
- *Strategic weight* — the work could affect trajectory, perception, or positioning even with a sound premise: it shapes what the system becomes known for; it is a complexity or simplicity bet affecting adoption, onboarding, or cognitive load; it opens or closes future directions (path dependencies, architectural commitments); it carries opportunity cost — building this means not building something else.

**design-lens** — UI/UX references, frontend components, or visual design language; user flows, wireframes, screen/page/view mentions; interaction descriptions (forms, buttons, navigation, modals); responsive behavior or accessibility.
Expand Down
14 changes: 14 additions & 0 deletions tests/pipeline-review-contract.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -880,6 +880,20 @@ describe("ce-doc-review contract", () => {
expect(line).not.toMatch(/(^|[;,] )data handling,/i)
})

// product-lens activated on "solution selection where alternatives plausibly exist",
// which holds for nearly any fix, so a judgment persona (and, through the trio gate,
// the cross-model pass) ran on routine bootstrap plans. The leg is now one condition:
// an unsettled product position; a mechanism choice is not one.
test("product-lens activates on an unsettled product position, not on plausible alternatives", async () => {
const content = await readRepoFile("skills/ce-doc-review/references/persona-selection.md")
const block = sliceSection(content, "**product-lens**", "**design-lens**")
expect(block).toMatch(/product position/)
expect(block).toMatch(/no upstream Product Contract settled|origin did not already settle/)
expect(block).toMatch(/choice among mechanisms .* is an implementation decision, not a product position/)
expect(block).not.toMatch(/alternatives plausibly exist/)
expect(block).toMatch(/\*Strategic weight\*/)
})

test("keeps security document review on the parent capability tier", async () => {
const content = await readRepoFile("skills/ce-doc-review/references/dispatch.md")
const modelTierSection = content.slice(content.indexOf("Model tiering lives here"))
Expand Down
4 changes: 4 additions & 0 deletions tests/skill-eval-cell/catalog.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,10 @@ describe("skill-eval-cell catalog", () => {
"ce-commit-push-pr/description-only-no-commit:references/pr-description-writing.md",
"ce-commit-push-pr/babysit-off-preserves-human-decision:references/apply-and-handoff.md",
"ce-debug/pipeline-convergent-fix:references/pipeline-mode.md",
"ce-doc-review/routine-fix-no-product-lens:references/persona-selection.md",
"ce-doc-review/settled-origin-no-product-lens:references/persona-selection.md",
"ce-doc-review/staked-position-keeps-product-lens:references/persona-selection.md",
"ce-doc-review/strategic-weight-keeps-product-lens:references/persona-selection.md",
"ce-debug/pipeline-divergent-defer:references/pipeline-mode.md",
"ce-handoff/resume-asks-does-not-act:references/resume.md",
"ce-ideate/unidentified-subject-reads-scope-gates:references/scope-gates.md",
Expand Down
Loading