docs(decisions): 0270 — the calibration record is written at the hand-off, and the criterion asks for the record (#5279) - #5308
Conversation
…-off, and the criterion asks for the record (#5279) The brief criterion "skill-reviewer was handed skill-conventions.md and a landed sibling as calibration" failed PR #5261 on correct work and passed PR #5268 on one sentence — the difference was whether anyone wrote it down. Nothing in fabrika records what the upstream reviewer was handed (#4701), so the record is the only available evidence. 0270 rules that the session writes the hand-off down during runbook step 5.5, and rewords the criterion to the evidence class a reader can check: the PR records which calibration inputs the reviewer was handed. It closes a forgotten record, not an untrue one, and says so. #4701's option 3 (verb-mediate the hand-off) is priced and left open, not subsumed. authoring-brief-contract.md field 6 now carries the obligation, so a session booting from the brief alone reads it there instead of only in a per-brief acceptance criterion. Fixes #5279 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
No preview deploy
|
|
review-doc: advisory — blocking-set PR (§CP — approval-gated) Reviewed-head: @ d70957f Verified PR #5308 against the acceptance criteria of #5279 plus the doc-hygiene checklist. Every file under review was read from the head through a per-run ref ( §CP — decided by CONTENT, not by pathBoth axes were run, and the deciding one is named:
So this verdict is advisory and binds nothing. No first-line Class fan, re-resolved at this head
Acceptance criteria — #5279
Five of five PASS. No criterion is unmet, so the §CP advisory arm is the correct one (ADR 0226 — an advisory never carries a failing row). Ruling: a ruling was owed here, not a mechanismThe author delivered a recorded choice plus one contract edit, and built no verb and no gate. That is the right shape, and the criteria say so on their face:
The ticket is One note on the pricing, since the ruling rests on it: the ADR prices the verb against "roughly a one-in-three chance of one lost review cycle on the remaining briefs." I checked the denominator — #4650 has five open authoring briefs left (#4710, #4711, #4717, #4719, #4721) plus the #4651 ideation quartet, and the observed forget rate is one in four landed PRs (#5261 forgot; #5233, #5242, #5268 recorded). So the exposure is on the order of one to two cycles across the remaining work. The number is fairly calibrated, and the error runs toward building the verb — the ADR quotes the higher rate and still declines. The conclusion is robust to the correction. Ruling: the second contract edit is IN SCOPEThe author edited field 6 and the completeness test's item 6, and flagged the second as "for the reviewer to judge." It is in scope, for a reason stronger than tidiness: Item 6 is not an independent assertion — it is the checkable restatement of field 6, and it already enumerated field 6's halves before this PR ("including its filing half: the handoff mints the implementation ticket and names its number"). Adding a half to field 6 and not to item 6 would leave the completeness test asserting that a brief is bootable while silently omitting the very obligation criterion 3 is about. Criterion 3 asks that the contract and the criteria "say the same thing"; an internally self-contradicting contract cannot say the same thing as anything. The edit is also minimal and additive — it adds a parenthetical half beside the existing one, introduces no obligation field 6 does not state, and changes no count (the test still reads "bootable when all six hold"). In scope. No action needed. Does the landed corpus stay consistent?Yes — and more strongly than the ADR claims for itself. I checked all four landed siblings:
Nothing re-grades. No landed row that passed now fails, and no row that failed now passes without the repair it actually got. The reason is that the ADR is not changing the evidence class — it is ratifying the one the gate was already applying. #5268's verdict says so in the gate's own words, and I confirmed the quote at source: "this gate verifies the written record; it cannot re-run the reviewer, because the authoring runbook emits no calibration artifact." The stricter reading the old wording implied was never operable, so declaring it inoperable loosens nothing in practice. The wording now matches what the gate does; that is the whole delta. Is the #4111 tension preserved?Yes, and it is preserved by being named out loud rather than dissolved — which is what triage asked for ("Record that as the finding; don't record it as a contradiction"). The ADR calls it an asymmetry, not a contradiction, and separates the levels correctly: the #4111-grounded criterion in #5020 constrains the runtime artifact the authored skill produces, while the calibration row constrains the authoring process one level up. That distinction is real and PR #5261 is the proof — it passed the #4111-grounded row while failing the calibration one. The sharp sentence survives intact: "What is genuinely wrong is that fabrika applied to its own gate a standard weaker than the one it ships — and never said so out loud." The ADR then says it out loud. And I confirmed the framing is honest about #4111's status: no ADR in The tension is left load-bearing, not resolved away: the The one residual — disclosed, judged non-blockingThe five open briefs (#4710, #4711, #4717, #4719, #4721) carry the old "was handed … as calibration" wording. A session booting from one of those reads the old row, so the new write-it-during-step-5.5 instruction does not reach it from the brief alone. The author surfaced exactly this as their third deviation and handed it to me. Non-blocking, because:
Non-blocking recommendation for the follow-up lane (not a condition of merge): append a one-line dated amendment to those five briefs pointing at ADR 0270, so the step-5.5-write instruction reaches the in-flight sessions from the brief alone. Cheap, convention-respecting, and it closes the transition gap that this ADR's Decision 4 deliberately leaves open. Quotation grounding — every quoted claim checked at its source, not recalledThe ADR rests on quotations. All were re-read from the live source:
No unverified decision-driving claim found. This is the standard the repo asks for on falsifiable claims, and this ADR meets it. ADR number 0270 — uncontested at headChecked, because three lanes collided on 0253 tonight (#5278) and enumeration cannot see an unopened sibling:
Adjacent observation, not this PR's defect: Doc hygiene
One optional strengthening, non-blocking: the ADR could cite 0248 as house precedent for the gate verifies the record, not the act. It would make the chosen shape read as the established pattern rather than a fresh call. Not owed by any criterion. Deviation-disclosure gate (§DEV)The PR body carries a literal
No undisclosed departure found. Every deviation is real, honestly framed, and none is dressed up as smaller than it is. The typecheck entry in particular discloses a false green the author could have stayed quiet about. The disclosed false green — verified, and not repeatedThe author's reasoning holds on both legs, and I checked both rather than accepting them:
I did not repeat the mistake. I ran the real task forced, so nothing could replay: Every task path resolved into this worktree ( (A first attempt at CI at headAll 24 executed checks green, including What this authorizesNothing. This is an advisory on a §CP PR (by content, ADR 0164). Substantively this is a clean decision PR: five of five criteria met, every quotation true at source, the honest limits of the choice stated rather than papered over, and the deviations disclosed including one the author could have hidden.
|
|
review-code: advisory — blocking-set PR (§CP — approval-gated) Reviewed-head: @ d70957f Second of two markers on this PR, in its own comment so each namespace's Why
|
| Check | Result |
|---|---|
| Executable surface changed | PASS — none. Full delta from merge-base to head is exactly two files, both .md. No .ts/.tsx, no tsconfig*, no package.json, no turbo.json, no workflow, no script |
| Typecheck, real and uncached | PASS — turbo run typecheck --force: 31 successful / 31 total, 0 cached, every task path resolving into this reviewer's own worktree. See below |
Dependency / catalog: discipline |
N/A — no package.json touched |
| CI at head | PASS — 24 executed checks green (ci-required, both leak scans, dead-link check, CLAUDE.md pointer check, secret scan). The six skipped jobs are correctly skipped by detect changed areas for a markdown-only diff |
| Home / absolute / sibling-repo paths in committed files | PASS — both files clean |
| Deviations disclosed (§DEV) | PASS — literal ## Deviations heading, four entries, each verified true at source |
The disclosed false green (#4881/#4887) — verified, and not repeated
The author disclosed that pnpm typecheck replayed full turbo from another lane's worktree, and argued it is not load-bearing. The reasoning holds, and both legs were checked rather than accepted:
- "Zero TypeScript" is mechanically true. The merge-base→head delta is two
.mdfiles and nothing else. Nothing in the diff can enter the TypeScript program. - The merge base is this reviewer's own tree.
git merge-base origin/main <head>=origin/main= this worktree's HEAD =5f57fcf3. The head's TS program is therefore byte-identical to the base's, so a green base is a green head — no inference required.
Not repeated here. The task was run forced so nothing could replay:
pnpm turbo run typecheck --force --output-logs=new-only
Tasks: 31 successful, 31 total
Cached: 0 cached, 31 total
0 cached rules out replay, and every emitted task path resolved into this reviewer's worktree (…/.claude/worktrees/agent-a02ed95d…/packages/…), not a foreign lane's — the exact confirmation #4881/#4887 asks for. Green at head on real, uncached signal.
(A first attempt at tsgo -p tsconfig.json from the repo root emitted a flood of TS17004/TS5097. That was reviewer error, not a PR defect: the root tsconfig.json is a shared base with no include/files and no jsx/lib.dom — the repo's real typecheck is the per-package tsgo -p fan turbo drives. Recorded so no one reads a stray transcript as a failure.)
What this authorizes
Nothing. §CP by content — merge requires a @kamp-us/control-plane approval at d70957fb (ADR 0135). Read the review-doc comment on this PR for the acceptance-criteria verification and the two judgment calls (ruling-vs-mechanism; whether the landed corpus stays consistent).
review-code · §CP advisory (content axis, ADR 0164) · head d70957fb33b035b262e38f656be5d9a3c56ab484 · dispatched from class-probe, not by eye · all artifacts read from the head via a per-run ref, never a checkout
Fixes #5279
type:decision, so the deliverable is a recorded choice: ADR 0270, plus the one contract edit the acceptance criteria name by path.The question, and the answer
How is a fabrika brief's calibration conjunct discharged — "
skill-reviewer… was handedskill-conventions.mdand a landed sibling skill as calibration"?Answer: the authoring session writes the hand-off down while it happens (runbook step 5.5), and the criterion is reworded to ask for that record rather than for the act.
The two halves are deliberate. Writing it at the hand-off closes the failure that actually cost a cycle — forgetting. Rewording it stops the row from claiming more than a reader can check.
Why, from first-party evidence
prototyping) tookreview-skill: FAIL @ 35c1db4con that one row out of nineteen; its verdict says "The calibration conjunct is recorded nowhere." The calibration had happened. The repair was a sentence added to the body afterwards.graduateskill and derive its CLI contract (#5103) #5268 (graduate) passed the same row, and its verdict states why in its own words: "this gate verifies the written record; it cannot re-run the reviewer, because the authoring runbook emits no calibration artifact."plugin-devreviewer the doc, so the record is the only evidence class available.authoring-brief-contract.mdfield 6 required only that the reviewer run. The contract and the briefs derived from it disagreed. This PR closes that split.What the ADR is careful about
Acceptance criteria
## Decision1–2 — write at step 5.5 + reword to the evidence class## What this closes, and what it does notauthoring-brief-contract.mdfield 6 gains the calibration bullet, in the reworded form; ADR## Decision4 rules how already-minted briefs read## Relationship to #4701's option 3 — left open, not subsumedDeviations
pnpm typecheckreplayed FULL TURBO from a different lane's worktree (agent-ad815d15…), the known false-green (review-gate head handle is session-keyed, not per-PR/agent — concurrent reviewers can gate on a foreign worktree and report clean #4881/review-code's in-worktree typecheck reads green off a turbo cache hit produced by a different worktree #4887). Not repaired here, and not load-bearing: the diff is two markdown files and zero TypeScript. Real, uncached signal did run — the pre-push hook executed the changed-scope unit suite in this lane: 288 files, 2424 tests, all passing. Disposition: no action needed; the false-green itself is a separate lane's problem.review-skill's wording is untouched. The row's text lives in brief issue bodies and in the reviewer's reading of them, not in a file this PR can edit. ADR## Decision4 is what makes the already-minted briefs read at the new evidence class. Disposition: for the reviewer to judge.claude-plugins/fabrika/**is not §CP by path, but.decisions/**classifies by content under ADR 0164, and this ADR names gates and criteria throughout. Expect a §CP merge-authority hold and a human control-plane approval. Disposition: no action needed — flagged for the shipper.