fix(scaffold): a cli adapter's in-flight children never outlive it (io.pilot.otto) - #114
Merged
Merged
Conversation
… during staging leaves no partial download Stacked on #111 (the adapter exits with its daemon; caller hang-up cancels its call). Generated cli/hybrid adapters still leaked in three ways (io.pilot.otto 0.20.0, measured on darwin-arm64, darwin-amd64 under Rosetta, linux-arm64 and linux-amd64; the same template backs aegis, docker, duckdb, miren, mysql, postgres, redis, smol, sqlite, tldr and upfile): - A CLI child still running when the adapter was SIGKILLed (the supervisor's stop path) or killed by its own Pdeathsig (daemon death) was reparented to init and kept running: `otto.exec ["mcp","serve-http",...]` kept its port open 8 s later on every platform. A SIGTERM could also let the adapter exit before the child it had just signalled was gone. - A stop during the first-spawn asset download (3.5-27.6 s for otto) left the partial pilot-asset-* in the daemon's shared TMPDIR: SIGTERM hit Go's default action because signal.NotifyContext was installed after staging. - A call that hit its per-method timeout came back as a normal reply ({"exit":-1}) instead of an error. Changes: - client_cli: every CLI child stays in the adapter's process group and, on Linux, gets Pdeathsig=SIGKILL from a thread held until Wait returns; a cancelled call SIGTERMs the child before WaitDelay escalates; a deadline or shutdown is an IPC error, not a result. - main: signal handling and the parent watch are installed before staging; shutdown waits (up to 8 s) for in-flight calls, and so their children. - stage: downloads go to $APP/.staged/tmp and are cancelled with the adapter; whatever a SIGKILLed start left there is swept by the next start. This consolidates the template fixes prototyped in the apps sweep for miren (child Pdeathsig, drain, staging) and aegis (SIGTERM on cancel), with their end-to-end tests, verbatim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ith it on macOS too Linux kills an in-flight CLI child when the adapter dies (Pdeathsig, previous commit). macOS has no parent-death signal, so when the adapter itself is SIGKILLed (the app-store supervisor's stop path through v1.0.3 signals only the adapter's pid; also the OOM killer or a crash) the child is reparented to launchd and runs on. io.pilot.otto 0.20.0 on darwin-arm64 and darwin-amd64: an in-flight `otto.exec ["mcp","serve-http",...]` still listened on its port 8 s after the adapter was SIGKILLed (ppid=1), and neither the next spawn's reaper (it matches the adapter's argv) nor anything else would stop it. A supervisor-side group stop (app-store #39) covers the stop path once a daemon release ships it; this covers every daemon version and the other ways an adapter dies, on both OSes. While a CLI child is in flight the adapter keeps a guard: the same binary re-executed with --pilot-child-guard, in the adapter's process group, reading a pipe only the adapter can write (close-on-exec). The kernel closes the pipe however the adapter goes away; on EOF the guard waits until it has been reparented (the adapter really is gone) and SIGKILLs the process group, itself included. A group id is not reused while a member lives, and the guard is a member, so the signal cannot reach an unrelated group. It runs only when the adapter leads its own group (as the supervisor spawns it), retires 1 s after the last child is reaped, and is released and reaped on a clean shutdown so the group is left empty. TestCLIInflightServerE2E (the otto shape: a passthrough call running a server) covers supervisor death with the supervisor's real spawn attributes, SIGKILL of the adapter, a clean stop leaving an empty group, guard idle retirement, the adapter surviving its guard's death, and a group stop. Against origin/main it fails on darwin-arm64, linux-arm64 and linux-amd64; with the guard disabled the SIGKILL subtest fails on darwin ("server child ... outlived the SIGKILLed adapter"); all pass with this change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…download is retried Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it still downloads the app's CLI before the adapter creates its socket. The supervisor gives a new app 3s to create its socket, and waitReady never looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a cli app whose first-start download takes longer than 3s is never marked ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every call on that spawn fails with "appstore: app not ready: io.pilot.miren" until the app respawns. With the real supervisor (spawn + Call), this reproduced on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on the first start, the adapter exits 1 before it listens, so even <ns>.help is unavailable. - stage.go: a Stager runs StageAssetsContext in the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within the call's own deadline. When an attempt fails, each call reports the failure, and a later call retries it (backoff starts at 2s and doubles up to 5 min) instead of the adapter exiting. The idle registry connection is closed after staging. tar and install steps get the same Pdeathsig and locked thread as CLI calls. - main.go: the socket is bound before staging, for cli and hybrid apps that ship assets. After serve returns, main stops staging and waits (up to 5s) for an interrupted download to remove its partial file. - The docs describe the new behaviour. Tests: - TestStageLifecycleE2E gains two subtests: - "socket and help are up while the binary is still downloading": the download stalls, the socket must accept within 2.5s, help must answer, and the call must wait and then run the binary. - "a failed download keeps serving and a later call retries". Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on linux/arm64 ("socket ... did not accept connections within 2.5s"). - TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its SIGKILL subtest on every OS (#114's child guard), and gives the timeout subtest a 4s deadline, because under heavy load a 2s deadline sometimes expired before the child had started. - The lifecycle tests (these two, #111's and #114's) pass 10x on darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64 (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and go vet ./... plus go test ./internal/... pass on linux/arm64. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
Author
|
Aegis measurement against this head (808243f). I built io.pilot.aegis 0.1.5 from it: natives from pilot-protocol/aegis#4, a TEST-ONLY key, and the natives staged over HTTP from a local mirror. I ran the lifecycle harness on darwin/arm64, darwin/amd64 (Rosetta), linux/arm64 and linux/amd64 ( Clean on all four:
Two differences from #116 plus the SIGTERM line (#117):
🤖 Generated with Claude Code |
This was referenced Sep 24, 2026
Merged
fix(scaffold): a cli app is ready before its binary is, and a failed download is retried
TeoSlayer
added a commit
that referenced
this pull request
Oct 1, 2026
…tlive the adapter, cli apps are ready before their binary (#111, #114, #116) * fix(scaffold): generated adapters exit with their daemon and drop abandoned calls Every app built from main.go.tmpl (sixtyfour and the other http, hybrid and cli apps) had four lifecycle faults. The harness sweep found all of them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64, darwin/amd64, linux/amd64 and linux/arm64: - Orphans. The only exit path was SIGINT/SIGTERM. When the daemon dies hard on macOS (no Pdeathsig), or on Linux daemons built before app-store v1.0.3, the adapter is reparented to pid 1 and keeps serving. watchParent records the spawning pid first thing in main and shuts down cleanly once the parent pid changes. - Abandoned calls. ipc.Serve ran handlers on the process ctx, so a caller that hung up (callFrom at its deadline, pilotctl at 120 s) left the backend request and its 2 fds open until the method timeout (60-280 s). serveConn is now the connection's only reader. It feeds ipc.Serve through a pipe and cancels the call's ctx on EOF. - Crash on EMFILE. Under the Linux supervisor's RLIMIT_NOFILE=256, about 120 abandoned or concurrent calls made Accept fail, and log.Fatalf killed the app along with every in-flight call. Accept now backs off from 5 ms to 1 s and keeps going on EMFILE, ENFILE and ECONNABORTED. - Unlinking another instance's socket. Closing the listener unlinked app.sock by name, so an orphan shutting down removed the socket of the instance the new daemon had just started. The listener no longer unlinks on close, and the adapter removes app.sock only when it is still the inode it bound. TestGeneratedAdapterLifecycleE2E builds a real adapter and checks all four. Every subtest fails on main and passes here: macOS -race x10, linux/arm64 x5 and linux/amd64 x5. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scaffold): a cli adapter's in-flight children die with it; a stop during staging leaves no partial download Stacked on #111 (the adapter exits with its daemon; caller hang-up cancels its call). Generated cli/hybrid adapters still leaked in three ways (io.pilot.otto 0.20.0, measured on darwin-arm64, darwin-amd64 under Rosetta, linux-arm64 and linux-amd64; the same template backs aegis, docker, duckdb, miren, mysql, postgres, redis, smol, sqlite, tldr and upfile): - A CLI child still running when the adapter was SIGKILLed (the supervisor's stop path) or killed by its own Pdeathsig (daemon death) was reparented to init and kept running: `otto.exec ["mcp","serve-http",...]` kept its port open 8 s later on every platform. A SIGTERM could also let the adapter exit before the child it had just signalled was gone. - A stop during the first-spawn asset download (3.5-27.6 s for otto) left the partial pilot-asset-* in the daemon's shared TMPDIR: SIGTERM hit Go's default action because signal.NotifyContext was installed after staging. - A call that hit its per-method timeout came back as a normal reply ({"exit":-1}) instead of an error. Changes: - client_cli: every CLI child stays in the adapter's process group and, on Linux, gets Pdeathsig=SIGKILL from a thread held until Wait returns; a cancelled call SIGTERMs the child before WaitDelay escalates; a deadline or shutdown is an IPC error, not a result. - main: signal handling and the parent watch are installed before staging; shutdown waits (up to 8 s) for in-flight calls, and so their children. - stage: downloads go to $APP/.staged/tmp and are cancelled with the adapter; whatever a SIGKILLed start left there is swept by the next start. This consolidates the template fixes prototyped in the apps sweep for miren (child Pdeathsig, drain, staging) and aegis (SIGTERM on cancel), with their end-to-end tests, verbatim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scaffold): a SIGKILLed cli adapter takes its in-flight children with it on macOS too Linux kills an in-flight CLI child when the adapter dies (Pdeathsig, previous commit). macOS has no parent-death signal, so when the adapter itself is SIGKILLed (the app-store supervisor's stop path through v1.0.3 signals only the adapter's pid; also the OOM killer or a crash) the child is reparented to launchd and runs on. io.pilot.otto 0.20.0 on darwin-arm64 and darwin-amd64: an in-flight `otto.exec ["mcp","serve-http",...]` still listened on its port 8 s after the adapter was SIGKILLed (ppid=1), and neither the next spawn's reaper (it matches the adapter's argv) nor anything else would stop it. A supervisor-side group stop (app-store #39) covers the stop path once a daemon release ships it; this covers every daemon version and the other ways an adapter dies, on both OSes. While a CLI child is in flight the adapter keeps a guard: the same binary re-executed with --pilot-child-guard, in the adapter's process group, reading a pipe only the adapter can write (close-on-exec). The kernel closes the pipe however the adapter goes away; on EOF the guard waits until it has been reparented (the adapter really is gone) and SIGKILLs the process group, itself included. A group id is not reused while a member lives, and the guard is a member, so the signal cannot reach an unrelated group. It runs only when the adapter leads its own group (as the supervisor spawns it), retires 1 s after the last child is reaped, and is released and reaped on a clean shutdown so the group is left empty. TestCLIInflightServerE2E (the otto shape: a passthrough call running a server) covers supervisor death with the supervisor's real spawn attributes, SIGKILL of the adapter, a clean stop leaving an empty group, guard idle retirement, the adapter surviving its guard's death, and a group stop. Against origin/main it fails on darwin-arm64, linux-arm64 and linux-amd64; with the guard disabled the SIGKILL subtest fails on darwin ("server child ... outlived the SIGKILLed adapter"); all pass with this change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scaffold): a cli app is ready before its binary is, and a failed download is retried Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it still downloads the app's CLI before the adapter creates its socket. The supervisor gives a new app 3s to create its socket, and waitReady never looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a cli app whose first-start download takes longer than 3s is never marked ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every call on that spawn fails with "appstore: app not ready: io.pilot.miren" until the app respawns. With the real supervisor (spawn + Call), this reproduced on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on the first start, the adapter exits 1 before it listens, so even <ns>.help is unavailable. - stage.go: a Stager runs StageAssetsContext in the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within the call's own deadline. When an attempt fails, each call reports the failure, and a later call retries it (backoff starts at 2s and doubles up to 5 min) instead of the adapter exiting. The idle registry connection is closed after staging. tar and install steps get the same Pdeathsig and locked thread as CLI calls. - main.go: the socket is bound before staging, for cli and hybrid apps that ship assets. After serve returns, main stops staging and waits (up to 5s) for an interrupted download to remove its partial file. - The docs describe the new behaviour. Tests: - TestStageLifecycleE2E gains two subtests: - "socket and help are up while the binary is still downloading": the download stalls, the socket must accept within 2.5s, help must answer, and the call must wait and then run the binary. - "a failed download keeps serving and a later call retries". Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on linux/arm64 ("socket ... did not accept connections within 2.5s"). - TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its SIGKILL subtest on every OS (#114's child guard), and gives the timeout subtest a 4s deadline, because under heavy load a 2s deadline sometimes expired before the child had started. - The lifecycle tests (these two, #111's and #114's) pass 10x on darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64 (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and go vet ./... plus go test ./internal/... pass on linux/arm64. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Teodor Calin <teodor@vulturelabs.io> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #111, which covers the adapter exiting with its daemon and a caller hang-up cancelling its call. This PR covers the CLI child processes, the first-spawn download, and timeouts. Review #111 first. Once #111 merges, this PR's base becomes
main.Problem
These leaks were measured in the published io.pilot.otto 0.20.0 bundles on darwin/arm64, darwin/amd64 (Rosetta), linux/arm64 and linux/amd64 (
docker --init). The same adapter template backs every cli/hybrid app: aegis, docker, duckdb, miren, mysql, otto, postgres, redis, smol, sqlite, tldr and upfile.otto.exec ["mcp","serve-http","--port",P]survives and keeps listening on P 8 s later, with ppid=1, on all 4 platforms.signal.NotifyContextis only installed after staging, so a SIGTERM takes Go's default action (exit 143) and leaves$TMPDIR/pilot-asset-*behind in the daemon's shared TMPDIR. SIGKILL does the same.{"exit":-1,"stdout":"Otto MCP HTTP server listening…"}after 60 s, rather than as an error.Fix (template only)
Commit 1 carries the fixes prototyped in the apps sweep for miren and aegis, copied verbatim (tests included), so those branches can rebase onto this one with no conflict.
client_cli: every CLI child stays in the adapter's process group. On Linux it also getsPdeathsig=SIGKILL, and the forking thread stays locked untilWaitreturns (Pdeathsig fires when that thread exits). A cancelled call sends SIGTERM first; exec'sWaitDelayescalates to SIGKILL after 5 s. A deadline or shutdown is returned as an IPC error.main: signal handling and fix(scaffold): generated adapters exit with their daemon and drop abandoned calls #111's parent watch are installed before staging. Shutdown waits up to 8 s for in-flight calls, and therefore for their children.stage: downloads go to$APP/.staged/tmpand are cancelled along with the adapter. Anything a SIGKILLed start leaves there is swept by the next start.Commit 2 closes the one case that commit 1 leaves open on macOS: a SIGKILLed adapter. macOS has no parent-death signal, so after the adapter dies nothing else would stop the child.
childguard.go(new). While a CLI child is running, the adapter keeps a guard: the same binary re-executed with--pilot-child-guard, in the adapter's process group, reading a pipe that only the adapter can write to (close-on-exec). The kernel closes that pipe however the adapter dies. On EOF, the guard waits until it has been reparented, which proves the adapter is really gone, and then SIGKILLs the process group, itself included. A process group id is not reused while any member is alive, and the guard is a member, so the signal cannot reach an unrelated group.kill -9, on both OSes. app-store http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39 (group SIGKILL on stop) is still worth having, but it only takes effect once a daemon release ships it.docs/CLI-ADAPTER.mdnow documents the child-lifetime and staging guarantees.Tests
TestCLIInflightServerE2E(new) reproduces otto's case: a passthrough call runs a server. It has six subtests: supervisor death using the supervisor's real spawn attributes (Setpgid, plus Pdeathsig on Linux from a locked thread, via a helper process); SIGKILL of the adapter; a clean stop leaving an empty group; guard idle retirement; the adapter surviving its guard's death; and a group stop. The tests from commit 1 areTestCLIAdapterChildLifecycleE2E(timeout returns an error; SIGTERM and SIGKILL with a child in flight) andTestStageInterruptedLeavesNothingE2E.TestCLIInflightServerE2Eonmainserver child … outlived the SIGKILLed adapter-count=5)-count=5)go test ./...build-teston this head (ubuntu-latest,-race)Linux ran in
golang:1.25.13-bookworm, with amd64 under emulation. The amd64 run is where the one bug found in review showed up. The guard originally read the adapter's pid withgetppid()at startup. Under emulation it could start after the adapter had already died, read pid 1, and then never act (adapter's process group … not empty after it was SIGKILLed). The adapter now passes its pid in argv. After that change, darwin and linux-arm64 pass with-count=5, and CI (linux-amd64,-race) passes.Verified on real otto bundles (rebuilt from this branch)
I built
submissions/io.pilot.ottoat version 0.20.1 withpublish.BuildBundle, the same path aspilot-app verify-submission, and signed it with a TEST-ONLY key. Nothing was published.pilot-app verifyand the sweep harness'sverify(manifest, signature, binary sha256, Mach-O/ELF os/arch) pass on all 4 bundles. I then ran the real staged otto CLI, withotto.exec ["mcp","serve-http"]in flight, against both the rebuilt 0.20.1 bundles and the published 0.20.0 ones. All rows come from bundles built from this head. For the Linux lifecycle runs I used copies of the bundles with the otto CLI pre-staged from its sha-pinned asset: R2 downloads inside Docker were taking more than 90 s at the time. The cold-start rows used the real download.{"exit":-1}reply…/otto: context deadline exceeded(darwin-arm64, linux-arm64), child gonepilot-asset-*left in TMPDIR$APP/.staged/tmpempty (darwin-arm64, linux-arm64)$APP/.staged/tmp, swept by the next start, which staged OK on darwin-arm64 in 6.5 s (on linux-arm64 the old partial was swept; the slow download did not finish within the 120 s window)Harness, lifecycle + stopkill + orphan phases:
otto.help×200: 200/200 on all 4.otto.status×200, which execs otto on every call: 200/200 on darwin-arm64 and linux-arm64. Under emulation the 180 s budget allowed 85 calls on darwin-amd64 and 87 on linux-amd64, all ok.Not in this PR
app.sockon disk. The supervisor deletes it before the next spawn.otto startdetaches its relay with setsid, as doredis-server --daemonizeandpg_ctl start) are outside the process group, so neither the group signal nor Pdeathsig reaches them. For otto this is latent: the staged Bun binary can't start its relay at all today (Unable to resolve relay runtime entrypoint, an upstream telepat-io/otto issue). The sweep's postgres branch prototypes acli.teardownhook for this case..shellcase): not covered here.ed25519:mTyrd5ZG/tl76CLpdUEaaGvCjrnE6QLHPVm7XrduH/w=), an R2 upload of the four adapter bundles, and a re-signed catalogue.🤖 Generated with Claude Code