Repository navigation
fix(scaffold): a cli app is ready before its binary is, and a failed download is retried - #116
Merged
Merged
Conversation
…download is retried Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it still downloads the app's CLI before the adapter creates its socket. The supervisor gives a new app 3s to create its socket, and waitReady never looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a cli app whose first-start download takes longer than 3s is never marked ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every call on that spawn fails with "appstore: app not ready: io.pilot.miren" until the app respawns. With the real supervisor (spawn + Call), this reproduced on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on the first start, the adapter exits 1 before it listens, so even <ns>.help is unavailable. - stage.go: a Stager runs StageAssetsContext in the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within the call's own deadline. When an attempt fails, each call reports the failure, and a later call retries it (backoff starts at 2s and doubles up to 5 min) instead of the adapter exiting. The idle registry connection is closed after staging. tar and install steps get the same Pdeathsig and locked thread as CLI calls. - main.go: the socket is bound before staging, for cli and hybrid apps that ship assets. After serve returns, main stops staging and waits (up to 5s) for an interrupted download to remove its partial file. - The docs describe the new behaviour. Tests: - TestStageLifecycleE2E gains two subtests: - "socket and help are up while the binary is still downloading": the download stalls, the socket must accept within 2.5s, help must answer, and the call must wait and then run the binary. - "a failed download keeps serving and a later call retries". Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on linux/arm64 ("socket ... did not accept connections within 2.5s"). - TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its SIGKILL subtest on every OS (#114's child guard), and gives the timeout subtest a 4s deadline, because under heavy load a 2s deadline sometimes expired before the child had started. - The lifecycle tests (these two, #111's and #114's) pass 10x on darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64 (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and go vet ./... plus go test ./internal/... pass on linux/arm64. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
Author
|
Aegis measurement against this head (0467324). I built io.pilot.aegis 0.1.5 from it: natives from pilot-protocol/aegis#4, a TEST-ONLY key, and the natives staged over HTTP. Then I SIGTERMed the app while
🤖 Generated with Claude Code |
TeoSlayer
force-pushed
the
fix/miren-lifecycle
branch
from
September 24, 2026 10:40
0467324 to
a0d7665
Compare
TeoSlayer
changed the base branch from
fix/sixtyfour-lifecycle
to
fix/otto-lifecycle
September 24, 2026 10:40
TeoSlayer
added a commit
that referenced
this pull request
Oct 1, 2026
…tlive the adapter, cli apps are ready before their binary (#111, #114, #116) * fix(scaffold): generated adapters exit with their daemon and drop abandoned calls Every app built from main.go.tmpl (sixtyfour and the other http, hybrid and cli apps) had four lifecycle faults. The harness sweep found all of them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64, darwin/amd64, linux/amd64 and linux/arm64: - Orphans. The only exit path was SIGINT/SIGTERM. When the daemon dies hard on macOS (no Pdeathsig), or on Linux daemons built before app-store v1.0.3, the adapter is reparented to pid 1 and keeps serving. watchParent records the spawning pid first thing in main and shuts down cleanly once the parent pid changes. - Abandoned calls. ipc.Serve ran handlers on the process ctx, so a caller that hung up (callFrom at its deadline, pilotctl at 120 s) left the backend request and its 2 fds open until the method timeout (60-280 s). serveConn is now the connection's only reader. It feeds ipc.Serve through a pipe and cancels the call's ctx on EOF. - Crash on EMFILE. Under the Linux supervisor's RLIMIT_NOFILE=256, about 120 abandoned or concurrent calls made Accept fail, and log.Fatalf killed the app along with every in-flight call. Accept now backs off from 5 ms to 1 s and keeps going on EMFILE, ENFILE and ECONNABORTED. - Unlinking another instance's socket. Closing the listener unlinked app.sock by name, so an orphan shutting down removed the socket of the instance the new daemon had just started. The listener no longer unlinks on close, and the adapter removes app.sock only when it is still the inode it bound. TestGeneratedAdapterLifecycleE2E builds a real adapter and checks all four. Every subtest fails on main and passes here: macOS -race x10, linux/arm64 x5 and linux/amd64 x5. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scaffold): a cli adapter's in-flight children die with it; a stop during staging leaves no partial download Stacked on #111 (the adapter exits with its daemon; caller hang-up cancels its call). Generated cli/hybrid adapters still leaked in three ways (io.pilot.otto 0.20.0, measured on darwin-arm64, darwin-amd64 under Rosetta, linux-arm64 and linux-amd64; the same template backs aegis, docker, duckdb, miren, mysql, postgres, redis, smol, sqlite, tldr and upfile): - A CLI child still running when the adapter was SIGKILLed (the supervisor's stop path) or killed by its own Pdeathsig (daemon death) was reparented to init and kept running: `otto.exec ["mcp","serve-http",...]` kept its port open 8 s later on every platform. A SIGTERM could also let the adapter exit before the child it had just signalled was gone. - A stop during the first-spawn asset download (3.5-27.6 s for otto) left the partial pilot-asset-* in the daemon's shared TMPDIR: SIGTERM hit Go's default action because signal.NotifyContext was installed after staging. - A call that hit its per-method timeout came back as a normal reply ({"exit":-1}) instead of an error. Changes: - client_cli: every CLI child stays in the adapter's process group and, on Linux, gets Pdeathsig=SIGKILL from a thread held until Wait returns; a cancelled call SIGTERMs the child before WaitDelay escalates; a deadline or shutdown is an IPC error, not a result. - main: signal handling and the parent watch are installed before staging; shutdown waits (up to 8 s) for in-flight calls, and so their children. - stage: downloads go to $APP/.staged/tmp and are cancelled with the adapter; whatever a SIGKILLed start left there is swept by the next start. This consolidates the template fixes prototyped in the apps sweep for miren (child Pdeathsig, drain, staging) and aegis (SIGTERM on cancel), with their end-to-end tests, verbatim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scaffold): a SIGKILLed cli adapter takes its in-flight children with it on macOS too Linux kills an in-flight CLI child when the adapter dies (Pdeathsig, previous commit). macOS has no parent-death signal, so when the adapter itself is SIGKILLed (the app-store supervisor's stop path through v1.0.3 signals only the adapter's pid; also the OOM killer or a crash) the child is reparented to launchd and runs on. io.pilot.otto 0.20.0 on darwin-arm64 and darwin-amd64: an in-flight `otto.exec ["mcp","serve-http",...]` still listened on its port 8 s after the adapter was SIGKILLed (ppid=1), and neither the next spawn's reaper (it matches the adapter's argv) nor anything else would stop it. A supervisor-side group stop (app-store #39) covers the stop path once a daemon release ships it; this covers every daemon version and the other ways an adapter dies, on both OSes. While a CLI child is in flight the adapter keeps a guard: the same binary re-executed with --pilot-child-guard, in the adapter's process group, reading a pipe only the adapter can write (close-on-exec). The kernel closes the pipe however the adapter goes away; on EOF the guard waits until it has been reparented (the adapter really is gone) and SIGKILLs the process group, itself included. A group id is not reused while a member lives, and the guard is a member, so the signal cannot reach an unrelated group. It runs only when the adapter leads its own group (as the supervisor spawns it), retires 1 s after the last child is reaped, and is released and reaped on a clean shutdown so the group is left empty. TestCLIInflightServerE2E (the otto shape: a passthrough call running a server) covers supervisor death with the supervisor's real spawn attributes, SIGKILL of the adapter, a clean stop leaving an empty group, guard idle retirement, the adapter surviving its guard's death, and a group stop. Against origin/main it fails on darwin-arm64, linux-arm64 and linux-amd64; with the guard disabled the SIGKILL subtest fails on darwin ("server child ... outlived the SIGKILLed adapter"); all pass with this change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scaffold): a cli app is ready before its binary is, and a failed download is retried Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it still downloads the app's CLI before the adapter creates its socket. The supervisor gives a new app 3s to create its socket, and waitReady never looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a cli app whose first-start download takes longer than 3s is never marked ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every call on that spawn fails with "appstore: app not ready: io.pilot.miren" until the app respawns. With the real supervisor (spawn + Call), this reproduced on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on the first start, the adapter exits 1 before it listens, so even <ns>.help is unavailable. - stage.go: a Stager runs StageAssetsContext in the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within the call's own deadline. When an attempt fails, each call reports the failure, and a later call retries it (backoff starts at 2s and doubles up to 5 min) instead of the adapter exiting. The idle registry connection is closed after staging. tar and install steps get the same Pdeathsig and locked thread as CLI calls. - main.go: the socket is bound before staging, for cli and hybrid apps that ship assets. After serve returns, main stops staging and waits (up to 5s) for an interrupted download to remove its partial file. - The docs describe the new behaviour. Tests: - TestStageLifecycleE2E gains two subtests: - "socket and help are up while the binary is still downloading": the download stalls, the socket must accept within 2.5s, help must answer, and the call must wait and then run the binary. - "a failed download keeps serving and a later call retries". Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on linux/arm64 ("socket ... did not accept connections within 2.5s"). - TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its SIGKILL subtest on every OS (#114's child guard), and gives the timeout subtest a 4s deadline, because under heavy load a 2s deadline sometimes expired before the child had started. - The lifecycle tests (these two, #111's and #114's) pass 10x on darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64 (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and go vet ./... plus go test ./internal/... pass on linux/arm64. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Teodor Calin <teodor@vulturelabs.io> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #111 (
fix/sixtyfour-lifecycle): this PR uses itswatchParent,serveConnand socket-ownership code and only adds to it. Review #111 first. Once #111 merges, this PR retargets tomain.Problem
The harness sweep of the published io.pilot.miren 0.1.0 bundles (all 4 platforms) found these problems. Each one comes from the cli templates, not from miren:
waitReadywaits 3 s (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3) and never retries. With the real supervisor'sspawn+Call, the first spawn answers every call withappstore: app not ready: io.pilot.mirenuntil the next respawn. This happened on darwin-arm64 and linux-arm64 with both v1.0.2 and v1.0.3. The socket appeared after 11 s (darwin-arm64), 15 s (darwin-amd64), 20 s (linux-arm64) and 22 s (linux-amd64).StageAssetsfails, thenlog.Fatalf, exit 1 before the socket exists, so evenmiren.helpis unavailable.Pdeathsig, but a child does not inherit it.miren loginrunning overmiren.execsurvived a SIGKILL of the adapter and a daemon death with Pdeathsig. It was reparented to pid 1 and ran for 4.1 s (chatty) to 29 s (silent).$TMPDIRpilot-asset-*behind (58–84 KB measured, up to the whole tarball). SIGTERM during staging is not handled, becausesignal.NotifyContextis installed after staging: exit 143, and the file is left behind.{"exit":-1,…}as a result instead of an error.Fix (templates + 2 docs)
childproc_{linux,other}.go.tmpl(new, rendered for cli and hybrid): on Linux, every process the adapter starts (CLI calls,tar, install steps) getsPdeathsig=SIGKILL. The forking thread stays locked untilWaitreturns. On other systems nothing is set. Children always stay in the adapter's process group, so the supervisor's group stop (http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39 in app-store) andreapStale'skill(-pgid)still reach them.stage.go.tmpl: aStagerrunsStageAssetsContextin the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within its own deadline. If an attempt fails, each call reports the failure, and a later call retries it (backoff from 2 s, doubling up to 5 min). The adapter no longer exits. Downloads go to$APP/.staged/tmp. That directory is removed when an attempt ends or is cancelled, and swept at the next start. Download,tarand install steps are bound to the ctx. The idle registry connection is closed after staging.main.go.tmpl:servereturns only after every running call has returned. It closes connections and waits up to 8 s.mainwaits for an interrupted download to clean up before it exits.client_cli.go.tmpl: a cancelled call, or a child ended by a signal, is now an IPC error (backend: …: context deadline exceeded) instead of{"exit":-1}.docs/CLI-ADAPTER.mdanddocs/R2-ARTIFACT-REGISTRY.mddescribe the new behaviour.Tests
TestCLIAdapterChildLifecycleE2Ecovers timeout, SIGTERM, SIGKILL on Linux, and parent death with a call in flight.TestStageLifecycleE2Ecovers four cases: ready and help while the download is stalled; SIGTERM mid-download; the leftover from a SIGKILL mid-download is swept; a failed download keeps serving and a later call retries.80d19b2), darwin/arm64 and linux/arm64 both fail the timeout subtest (out={"exit":-1,…} err=<nil>) and all 4 staging subtests (socket … did not accept connections within 2.5s,download never startedunder$APP/.staged/tmp). linux/arm64 also fails the SIGKILL subtest (in-flight child pid … outlived the SIGKILLed adapter).origin/main(fcd0b12), parent death also fails (adapter pid … kept running after its parent died). fix(scaffold): generated adapters exit with their daemon and drop abandoned calls #111 fixes that one.TestGeneratedAdapterLifecycleE2Epass 5 times in a row on darwin/arm64 and on linux/arm64 (golang:1.25.13-bookworm,--init). The two new tests also passed once on emulated linux/amd64, before this branch was rebased onto fix(scaffold): generated adapters exit with their daemon and drop abandoned calls #111.go test ./...passes on darwin/arm64, andgo vet ./...plusgo test ./internal/...pass on linux/arm64.Verified on real miren bundles built from this branch
publish.BuildBundlebuiltsubmissions/io.pilot.mirenwith version 0.1.1, for all 4DefaultPlatforms, signed with a TEST-ONLY key. Nothing was uploaded or published.pilot-app verifypasses for all 4, and the harnessverifypasses too: sha, signature, pin, and Mach-O/ELF arch. Apart from version, binary sha and signature, the manifest andinstall.jsonare identical to the published 0.1.0.Real supervisor, first spawn (
sup.spawn+sup.Call, nothing staged):app not readymiren.helpok;miren.exec versionok after 9.3 s of stagingHarness (help ×200 + stop/orphan phases;
miren.exec version×200):sandbox-execdeny network /--network none): help 20/20 on all 4.miren.execreturnsinstall assets: stage: fetch …as an error, and the app keeps running.In-flight
miren login(fake device-flow endpoint), then the app is stopped. The table shows whether themirenchild is still alive afterwards:Real supervisor stop with that call in flight:
alive ppid=1with v1.0.2 and v1.0.3, and gone with http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39 (fix/stop-app-process-group).Not in this PR
miren serverthroughmiren.execmight, but that was not verified (it failed in an unprivileged container).🤖 Generated with Claude Code