Skip to content

feat: run bounded coding sessions against a pull request - #109

Open
pulkitxm wants to merge 1 commit into
claude/quinjet-context-bundlefrom
claude/quinjet-work-sessions
Open

pulkitxm wants to merge 1 commit into
claude/quinjet-context-bundlefrom
claude/quinjet-work-sessions

Conversation

@pulkitxm

Copy link
Copy Markdown
Owner

Seventh in the review control plane stack. Stacked on #108 (claude/quinjet-context-bundle); review the last commit only.

What this adds

quinjet work runs a bounded coding session against one pull request: an isolated checkout at an exact commit, a task list drawn from something Quinjet already computes, verification commands whose results are recorded, and one local commit at the end. Quinjet does not run a model and does not know which coding tool does the writing. The part that is the same either way is the boundary and the record, and that is what this is.

The boundary is stated, not implied

Every session stores what it may and may not do, and prints both:

May — read and write files inside its own worktree; commit to its own branch; run the verification commands recorded on it.

May not — push the branch or any other ref; comment on the pull request or reply to a thread; resolve, unresolve or otherwise change a review thread; merge, close, reopen or edit the pull request.

Those four are not gaps to fill in later. Quinjet stays the source of truth for everything that reaches GitHub. work publish writes one local commit and then names the commands that would take it further — git push origin quinjet/work/w42-1, quinjet pr gate 42 — without running them.

The exact starting commit

work start --pr 42 --worktree records the head object name at that moment and checks it out on quinjet/work/w42-1 via git worktree add -b, as a sibling of the repository. Nothing a session does appears as an untracked file in the checkout the reviewer is using, and a push landing on the pull request while the session is open changes neither the base the session is measured from nor what it appears to have changed.

Without --worktree or --into a session records its task list and nothing else, which is what you want when the coding tool brings its own sandbox. diff, verify and publish then say plainly that there is no worktree.

Task lists are exactly what the source names

--from feedback takes the blocking rows of pr feedback; --from failed-checks takes the failing gate checks and the failure-severity annotations; --from whole takes nothing. A session started from failing checks does not quietly also carry the review threads, because then nobody could say afterwards what it was for. Each task carries the Quinjet command that resolves it, as information the session cannot act on.

Task summaries and bodies are text written by whoever can reach the pull request, and the text face prints them under a heading that says so.

Verification, honestly recorded

work verify w42-1 -- cargo test spawns the command as argv with the worktree as its working directory. No shell: no pipe, no glob, no expansion. It is stored as an argv array rather than a string for the same reason, because a record that flattened it would imply a shell parsed it.

Re-running a recorded command replaces its row rather than stacking a second one, so the list is always the latest result per distinct command. work verify w42-1 with no command replays exactly what was recorded. A session that has run nothing refuses to report a pass it did not earn, and says so with exit 1.

--exit-code turns the session's verdict into the process's.

Publish and abort

publish previews the exact set it would stage — tracked changes and untracked additions both — names a failing verification rather than hiding it or refusing over it, and commits nothing without --yes. abort counts the checkpoint commits that would go with the branch before removing anything, and leaves the pull request untouched.

Tests

  • src/git/work/tests.rs (13): the exact start commit and branch naming; the forbidden list being present and covering all four operations; feedback sessions taking blocking rows and leaving advisories; failed-check sessions naming checks and never threads; a check name with a space quoted in the command it suggests; whole carrying no tasks even when a queue and gate are supplied; a session with no runs not counting as verified; re-running replacing rather than stacking; the failing verification being the one reported; no-worktree and deleted-worktree both erroring rather than reporting empty.
  • src/state/work_session/tests.rs (7): round trip, replace-and-move-to-front, identifier allocation skipping every stored name, forgetting one leaving its neighbours, the cap dropping the least recently touched, and an unreadable state document reading as no sessions.
  • tests/cli/work.rs (20): the worktree really being checked out at the pull request head on the session branch; the boundary in both faces; both task sources end to end; list; a diff measured from the start commit; passing and failing verifications with their exit codes and --exit-code; replay not stacking; the refusal to replay nothing; publish previewing without committing; publish committing tracked and untracked files, using the message, and making no gh call at all; publish on an empty session; the unverified and failing-verification notices; abort previewing and then removing worktree, branch and record; every session-taking verb returning exit 3 with a hint for an unknown name; two sessions on one pull request getting distinct names and branches.
  • Argument-parsing cases for the new verbs and their rejections; capability count 128 → 136; work added to the root subcommand list.
  • src/cli/tests/arguments.rs was over the 500-line limit after the new cases, so the two case-table tests moved to src/cli/tests/relationships.rs.

Checks

cargo fmt (stable and nightly), cargo clippy --all-targets (stable 1.98 and beta), cargo test, scripts/check_rust_sizes.py, scripts/check_comments.py, scripts/check_secrets.py all clean.

Same unrelated pre-existing failure as the earlier PRs in this stack: git::github::tests::operations::discovers_distinct_fetch_and_push_repositories_for_each_remote fails on main too in this sandbox, because the global git config here rewrites git@github.com: to HTTPS and that unit test does not isolate against it.

Docs

New docs/cli/work/ group with a README covering the boundary, the exact starting commit and the untrusted task text, plus a page per verb. Linked from the CLI index.

🤖 Generated with Claude Code

https://claude.ai/code/session_011n3Xp6dLNibg3gCBN9y5ij


Generated by Claude Code

@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 13:47 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 13:47 — with GitHub Actions Active
@pukbot pukbot Bot added size/XXL documentation Improvements or additions to documentation rust git cli labels Aug 29, 2026
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:04 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:04 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:08 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:08 — with GitHub Actions Active
@pulkitxm
pulkitxm force-pushed the claude/quinjet-work-sessions branch from 11ab2da to 1e2b334 Compare August 29, 2026 14:11
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:11 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:11 — with GitHub Actions Active
@pulkitxm
pulkitxm force-pushed the claude/quinjet-work-sessions branch from 1e2b334 to 85a2c75 Compare August 29, 2026 14:15
@pulkitxm
pulkitxm force-pushed the claude/quinjet-work-sessions branch from 85a2c75 to d0d2ad4 Compare August 29, 2026 14:15
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:15 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:15 — with GitHub Actions Active
@pulkitxm pulkitxm changed the title Add provider-neutral work sessions feat: run bounded coding sessions against a pull request Aug 29, 2026
@pulkitxm
pulkitxm force-pushed the claude/quinjet-work-sessions branch from d0d2ad4 to b541a03 Compare August 29, 2026 14:22
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:22 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:22 — with GitHub Actions Active
Comment thread src/git/github/context/tests.rs Fixed
@pulkitxm
pulkitxm force-pushed the claude/quinjet-work-sessions branch from b541a03 to e6c954c Compare August 29, 2026 14:34
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:34 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:34 — with GitHub Actions Active
@pulkitxm
pulkitxm marked this pull request as ready for review August 29, 2026 14:43
A work session is a bounded place for a coding process to work on one pull
request, plus a record of what it was asked and what it did. Quinjet does
not run a model and does not know which tool does the writing; the part
that is the same either way is the boundary and the record.

The boundary is stated on the session rather than implied. A session may
read and write files inside its own worktree, commit to its own branch,
and run the verification commands recorded on it. It may not push, comment
on the pull request, resolve a thread, or merge, close or edit it. Those
four stay explicit Quinjet operations somebody asks for, so publishing a
session writes one local commit and then names the commands that would
take it further rather than running them.

`work start` records the pull request's head object name exactly and, with
`--worktree`, checks it out on `quinjet/work/<id>` beside the repository,
so nothing a session does shows up in the checkout the reviewer is using
and a later push cannot change what the session appears to have changed.
The task list is exactly what the source names: the blocking rows of the
feedback queue, or the failing checks and their failure annotations, or
nothing at all. Each task carries the Quinjet command that resolves it, as
information the session cannot act on.

`work verify` spawns its command as argv with the worktree as the working
directory, never through a shell, and stores it as argv for the same
reason: a record that flattened it to a string would imply a shell parsed
it. Re-running a recorded command replaces its result rather than stacking
a second one, and a session that has run nothing refuses to report a pass
it did not earn.

`work publish` previews the exact file set it would stage, tracked and
untracked, and names a failing verification rather than either hiding it
or refusing. `work abort` counts the checkpoint commits that would go with
the branch before removing anything, and leaves the pull request untouched
because GitHub never knew the session existed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3Xp6dLNibg3gCBN9y5ij
@pulkitxm
pulkitxm force-pushed the claude/quinjet-work-sessions branch from e6c954c to 4073c18 Compare August 29, 2026 14:57
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:57 — with GitHub Actions Active
@pulkitxm
pulkitxm deployed to pukbot-production August 29, 2026 14:57 — with GitHub Actions Active

Copy link
Copy Markdown
Owner Author

Test (windows-latest) hung on 4073c18 and I have re-run it once. This is not a failure of this pull request, and the underlying issue is pre-existing.

What happened. The job passed the Tests step (cargo-nextest) in 1m30s, then entered Tests (release) at 15:00:02Z and was still in that step 76 minutes later, having produced nothing. On #108 and #110 — same runner label, same minutes of the day — that step took 5m04s and 4m40s. So this is a hang, not a slow release build. #106 hung the same way at the same time.

Why it can hang indefinitely. The two test steps use different runners:

- name: Tests
  run: cargo nextest run --all-features --locked --no-fail-fast
- name: Tests (release)
  run: cargo test --all-features --locked --release

cargo-nextest applies a per-test timeout (it terminates at 120s); plain cargo test has none, and the job sets no timeout-minutes, so the same hang runs to the 6-hour workflow limit.

The likely test. On an earlier commit of #110 the nextest step reported TIMEOUT [120.026s] workspace::tests::reachability_results_update_apps_without_replacing_open_tabs, with the other 817 tests passing. It is in src/workspace/tests.rs, exercises SSH reachability probing, is intermittent (it has passed on #104, #105, #107, #108 and #110), and runs in 0.18s locally.

Why it is not this pull request's. Nothing in this stack touches src/workspace/; git diff --name-only main..claude/quinjet-stack-review contains no path under it. This branch adds quinjet work and its own tests, all of which are Unix-gated and do not run on Windows at all.

What I have not done. I have not modified that test — fixing it is worth doing but belongs in its own change rather than widened into this pull request. The durable fixes are a timeout-minutes on the job and/or making the probe deterministic under test.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

Root cause for the Test (windows-latest) stall, correcting my earlier comment.

Attempt 2 of run 33258952100 ended the same way as attempt 1: Tests (release) ran 16:21:55Z to 22:19:45Z and was cancelled at the six-hour job ceiling. Nothing failed; the step never finished. Every other job on the run passed.

The log names the test, and it is not the one I guessed before:

16:26:37.9  test workspace::tests::replacement_keeps_tab_identity_and_new_tabs_follow_the_close_lifecycle ... ok
16:27:35.2  test workspace::tests::machine_picker_context_follows_replaced_and_appended_projects has been running for over 60 seconds
22:19:45.2  ##[error]The operation was canceled.

It is workspace::tests::machine_picker_context_follows_replaced_and_appended_projects. reachability_results_update_apps_without_replacing_open_tabs, which I named earlier, passed in this same run in 0.27s.

Why it only bites the release step: Tests runs under nextest, which gives every test its own process and a 120s cap, and that step has passed on Windows on every run of this stack. Tests (release) runs plain cargo test, so all tests share one process with no per-test timeout, and the test job has no timeout-minutes either. A single stuck test therefore holds the runner until the six-hour ceiling instead of failing in seconds.

This is not this pull request's failure. src/workspace/ is untouched by every commit in the stack, the test is equally present on main, and the same hang appeared on #106, whose re-run then passed. No fix for it exists anywhere I can port in, so rather than widen this pull request I am leaving it and proposing the patch here. It is two lines in .github/workflows/ci.yml:

   test:
     name: Test (${{ matrix.os }})
+    timeout-minutes: 30
       - name: Tests (release)
-        run: cargo test --all-features --locked --release
+        run: cargo nextest run --all-features --locked --no-fail-fast --cargo-profile release

The first bounds any future hang to 30 minutes instead of six hours. The second gives the release step the same process-per-test isolation and per-test cap that already makes the debug step reliable on Windows, so a stuck test fails as a test rather than as a cancelled job. Whoever picks that up may also want to look at why that test can block on Windows at all, since the timeout only bounds the symptom.

My one permitted re-run is spent, so I am not re-running a third time, and I have not touched the test.


Generated by Claude Code

This branch was successfully deployed

1 active deployment
pukbot-production — 4073c18a Deployed Aug 29, 2026 by pulkitxm via Size label #346
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cli documentation Improvements or additions to documentation git rust size/XXL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants