fix(scaffold): generated adapters exit with their daemon and drop abandoned calls - #111
Merged
Merged
Conversation
…ndoned calls Every app built from main.go.tmpl (sixtyfour and the other http, hybrid and cli apps) had four lifecycle faults. The harness sweep found all of them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64, darwin/amd64, linux/amd64 and linux/arm64: - Orphans. The only exit path was SIGINT/SIGTERM. When the daemon dies hard on macOS (no Pdeathsig), or on Linux daemons built before app-store v1.0.3, the adapter is reparented to pid 1 and keeps serving. watchParent records the spawning pid first thing in main and shuts down cleanly once the parent pid changes. - Abandoned calls. ipc.Serve ran handlers on the process ctx, so a caller that hung up (callFrom at its deadline, pilotctl at 120 s) left the backend request and its 2 fds open until the method timeout (60-280 s). serveConn is now the connection's only reader. It feeds ipc.Serve through a pipe and cancels the call's ctx on EOF. - Crash on EMFILE. Under the Linux supervisor's RLIMIT_NOFILE=256, about 120 abandoned or concurrent calls made Accept fail, and log.Fatalf killed the app along with every in-flight call. Accept now backs off from 5 ms to 1 s and keeps going on EMFILE, ENFILE and ECONNABORTED. - Unlinking another instance's socket. Closing the listener unlinked app.sock by name, so an orphan shutting down removed the socket of the instance the new daemon had just started. The listener no longer unlinks on close, and the adapter removes app.sock only when it is still the inode it bound. TestGeneratedAdapterLifecycleE2E builds a real adapter and checks all four. Every subtest fails on main and passes here: macOS -race x10, linux/arm64 x5 and linux/amd64 x5. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 24, 2026
… during staging leaves no partial download Stacked on #111 (the adapter exits with its daemon; caller hang-up cancels its call). Generated cli/hybrid adapters still leaked in three ways (io.pilot.otto 0.20.0, measured on darwin-arm64, darwin-amd64 under Rosetta, linux-arm64 and linux-amd64; the same template backs aegis, docker, duckdb, miren, mysql, postgres, redis, smol, sqlite, tldr and upfile): - A CLI child still running when the adapter was SIGKILLed (the supervisor's stop path) or killed by its own Pdeathsig (daemon death) was reparented to init and kept running: `otto.exec ["mcp","serve-http",...]` kept its port open 8 s later on every platform. A SIGTERM could also let the adapter exit before the child it had just signalled was gone. - A stop during the first-spawn asset download (3.5-27.6 s for otto) left the partial pilot-asset-* in the daemon's shared TMPDIR: SIGTERM hit Go's default action because signal.NotifyContext was installed after staging. - A call that hit its per-method timeout came back as a normal reply ({"exit":-1}) instead of an error. Changes: - client_cli: every CLI child stays in the adapter's process group and, on Linux, gets Pdeathsig=SIGKILL from a thread held until Wait returns; a cancelled call SIGTERMs the child before WaitDelay escalates; a deadline or shutdown is an IPC error, not a result. - main: signal handling and the parent watch are installed before staging; shutdown waits (up to 8 s) for in-flight calls, and so their children. - stage: downloads go to $APP/.staged/tmp and are cancelled with the adapter; whatever a SIGKILLed start left there is swept by the next start. This consolidates the template fixes prototyped in the apps sweep for miren (child Pdeathsig, drain, staging) and aegis (SIGTERM on cancel), with their end-to-end tests, verbatim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ith it on macOS too Linux kills an in-flight CLI child when the adapter dies (Pdeathsig, previous commit). macOS has no parent-death signal, so when the adapter itself is SIGKILLed (the app-store supervisor's stop path through v1.0.3 signals only the adapter's pid; also the OOM killer or a crash) the child is reparented to launchd and runs on. io.pilot.otto 0.20.0 on darwin-arm64 and darwin-amd64: an in-flight `otto.exec ["mcp","serve-http",...]` still listened on its port 8 s after the adapter was SIGKILLed (ppid=1), and neither the next spawn's reaper (it matches the adapter's argv) nor anything else would stop it. A supervisor-side group stop (app-store #39) covers the stop path once a daemon release ships it; this covers every daemon version and the other ways an adapter dies, on both OSes. While a CLI child is in flight the adapter keeps a guard: the same binary re-executed with --pilot-child-guard, in the adapter's process group, reading a pipe only the adapter can write (close-on-exec). The kernel closes the pipe however the adapter goes away; on EOF the guard waits until it has been reparented (the adapter really is gone) and SIGKILLs the process group, itself included. A group id is not reused while a member lives, and the guard is a member, so the signal cannot reach an unrelated group. It runs only when the adapter leads its own group (as the supervisor spawns it), retires 1 s after the last child is reaped, and is released and reaped on a clean shutdown so the group is left empty. TestCLIInflightServerE2E (the otto shape: a passthrough call running a server) covers supervisor death with the supervisor's real spawn attributes, SIGKILL of the adapter, a clean stop leaving an empty group, guard idle retirement, the adapter surviving its guard's death, and a group stop. Against origin/main it fails on darwin-arm64, linux-arm64 and linux-amd64; with the guard disabled the SIGKILL subtest fails on darwin ("server child ... outlived the SIGKILLed adapter"); all pass with this change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 24, 2026
Merged
…download is retried Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it still downloads the app's CLI before the adapter creates its socket. The supervisor gives a new app 3s to create its socket, and waitReady never looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a cli app whose first-start download takes longer than 3s is never marked ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every call on that spawn fails with "appstore: app not ready: io.pilot.miren" until the app respawns. With the real supervisor (spawn + Call), this reproduced on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on the first start, the adapter exits 1 before it listens, so even <ns>.help is unavailable. - stage.go: a Stager runs StageAssetsContext in the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within the call's own deadline. When an attempt fails, each call reports the failure, and a later call retries it (backoff starts at 2s and doubles up to 5 min) instead of the adapter exiting. The idle registry connection is closed after staging. tar and install steps get the same Pdeathsig and locked thread as CLI calls. - main.go: the socket is bound before staging, for cli and hybrid apps that ship assets. After serve returns, main stops staging and waits (up to 5s) for an interrupted download to remove its partial file. - The docs describe the new behaviour. Tests: - TestStageLifecycleE2E gains two subtests: - "socket and help are up while the binary is still downloading": the download stalls, the socket must accept within 2.5s, help must answer, and the call must wait and then run the binary. - "a failed download keeps serving and a later call retries". Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on linux/arm64 ("socket ... did not accept connections within 2.5s"). - TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its SIGKILL subtest on every OS (#114's child guard), and gives the timeout subtest a 4s deadline, because under heavy load a 2s deadline sometimes expired before the child had started. - The lifecycle tests (these two, #111's and #114's) pass 10x on darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64 (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and go vet ./... plus go test ./internal/... pass on linux/arm64. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 24, 2026
fix(scaffold): a cli app is ready before its binary is, and a failed download is retried
fix(scaffold): a cli adapter's in-flight children never outlive it (io.pilot.otto)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Every adapter generated from
internal/scaffold/templates/main.go.tmplhas four lifecycle faults. The harness sweep found all of them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64, darwin/amd64, linux/amd64 and linux/arm64. Rebuilding sixtyfour frommain(fcd0b12,submissions/io.pilot.sixtyfour/case.json) reproduces them, so they come from the template, not from sixtyfour.ipc.Serveruns handlers on the process ctx, so a caller that hangs up (supervisor.callFromat its deadline,pilotctl appstore callat 120 s) doesn't cancel anything. The backend request and its 2 fds stay open until the method timeout (60 s for find_email, 280 s for the intelligence methods). darwin-arm64: 30 calls abandoned after 0.3 s took fds from 7 to 67, and the backend still served every request.RLIMIT_NOFILE=256, 121 (amd64) or 124 (arm64) abandoned calls, or 150 legitimately concurrent ones, crash the app:serve: accept: accept unix app.sock: accept4: too many open files→log.Fatalf→ exit 1. Every in-flight call is lost (0/150 replies).app.sockby name. After a hard daemon death, stopping the old orphan deleted the socket the new instance had just bound, so calls failed withFileNotFoundErrorwhile the new instance was still running.Fix (template only)
watchParent:mainrecordsos.Getppid()before anything slow. A 250 ms ticker cancels the serve ctx once the parent pid changes, so the adapter shuts down the normal way (listener closed, socket removed). An adapter started with ppid ≤ 1 (for example a container's PID 1) doesn't watch.serveConn: a goroutine is the connection's only reader and feedsipc.Servethrough anio.Pipe. When the caller sends EOF, it cancels the call's ctx, so the backend request, or a CLI child throughexec.CommandContext, is aborted immediately. None of the in-tree callers (ipc.Call,supervisor.callFrom) half-close before reading the reply.net/httphandles them, instead oflog.Fatalf.SetUnlinkOnClose(false), andapp.sockis removed only if it's still the inode this process bound (os.SameFile).Tests
internal/scaffold/zz_adapter_lifecycle_e2e_test.go(TestGeneratedAdapterLifecycleE2E) generates and builds a real http adapter, points it at anhttptestbackend, and checks:ulimit -n 32with 40 waiting callers, the adapter hits EMFILE (the test asserts the log line) and keeps running. It serves again once the callers hang up. On macOS the kernel drops a pending connection whenaccept(2)fails with EMFILE, where Linux leaves it queued, so the test tolerates a caller being turned away.Results:
mainall 4 subtests fail:adapter pid … still running 3s after its parent was killed (orphaned),old instance's exit removed the new instance's socket,backend request still open 3s after its caller hung up,adapter exited while out of descriptors: exit status 1.-race -count=1040/40; linux/arm64 and linux/amd64 (golang:1.25.13-bookworm)-count=520/20 each.go test ./...passes on macOS.internal/scaffoldandinternal/publishalso pass on linux/arm64.Verified on real sixtyfour bundles (rebuilt from this branch)
I built
case.jsonwith the version bumped to 0.1.1 throughpublish.BuildBundle(go1.26.8, a TEST-ONLY key, nothing published) and ran the harness on every platform. darwin/amd64 ran under Rosetta, Linux indocker --init, both online and--network none:sixtyfour.help200/200 on all 4 platforms, online and offline. 10,000 calls on darwin-arm64: fd 7→7, RSS levels off at 18.4 MB.find_email×1000 through a local mock broker (SIXTYFOUR_BACKEND_URL): 1000/1000, fd 8→8.X-Pilot-Signature,X-Pilot-CallerandX-Pilot-Timestampare present.RLIMIT_NOFILE=256on Linux: app fds stay at 7–9 (16–18 on amd64 under Rosetta). 150/150 upstream requests were cancelled by the app. No crash, and SIGTERM exits cleanly.NOFILE=256: the app survives. amd64: 150/150 ok. arm64: 124 ok, and 26 got an error reply (socket: too many open filesdialing the backend), so every caller got an answer. Published 0.1.0: exit 1, 0 replies. On darwin (no cap): 150/150.app.socksurvives both A's own exit and a SIGTERM to A, and B keeps serving. Published 0.1.0:app.sock exists=False … FileNotFoundError.Not in this PR
app.sockstays on disk after a stop. The adapter can't clean up after a SIGKILL; the supervisor has to send SIGTERM first. Measured with the real supervisorspawn+ ctx cancel on darwin-arm64, 3 rounds each: app-storemain79e9944 and fix(supervisor): stopping an app stops the processes it started app-store#39 (SIGKILL to the group) giveexit=-1, app.sock left behind=true. A SIGTERM-first stop with a SIGKILL fallback, stacked on http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39, givesexit=0 after ~20 ms, app.sock left behind=falsefor both this build and 0.1.0. That's app-store code and is handled separately. Until it lands, the adapter and the supervisor both delete a stale socket beforeListen, so nothing breaks.ed25519:VoVCiQKPr73di2MlUd091a2Y6TCj/edSbCwRDtnYquI=), which isn't available here. Rebuilding from this commit with go1.26.8 is reproducible: the binary sha256 was identical across two builds on all 4 platforms. Note that a rebuild from currentmainalso addssixtyfour.balancetoexposes, which comes from earlier template changes.🤖 Generated with Claude Code