Skip to content

fix(scaffold): a cli app is ready before its binary is, and a failed download is retried - #116

Merged
TeoSlayer merged 1 commit into
fix/otto-lifecyclefrom
fix/miren-lifecycle
Oct 1, 2026
Merged

TeoSlayer merged 1 commit into
fix/otto-lifecyclefrom
fix/miren-lifecycle

Conversation

@TeoSlayer

Copy link
Copy Markdown
Contributor

Stacked on #111 (fix/sixtyfour-lifecycle): this PR uses its watchParent, serveConn and socket-ownership code and only adds to it. Review #111 first. Once #111 merges, this PR retargets to main.

Problem

The harness sweep of the published io.pilot.miren 0.1.0 bundles (all 4 platforms) found these problems. Each one comes from the cli templates, not from miren:

Published 0.1.0 (measured)
Not ready on first install The adapter downloads its CLI (40–60 MB for miren) before it creates its socket. The supervisor's waitReady waits 3 s (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3) and never retries. With the real supervisor's spawn + Call, the first spawn answers every call with appstore: app not ready: io.pilot.miren until the next respawn. This happened on darwin-arm64 and linux-arm64 with both v1.0.2 and v1.0.3. The socket appeared after 11 s (darwin-arm64), 15 s (darwin-amd64), 20 s (linux-arm64) and 22 s (linux-amd64).
Offline first start StageAssets fails, then log.Fatalf, exit 1 before the socket exists, so even miren.help is unavailable.
LEAK: in-flight CLI child outlives the adapter (Linux) The supervisor gives the adapter Pdeathsig, but a child does not inherit it. miren login running over miren.exec survived a SIGKILL of the adapter and a daemon death with Pdeathsig. It was reparented to pid 1 and ran for 4.1 s (chatty) to 29 s (silent).
LEAK: partial download in the shared $TMPDIR A first-start download that is interrupted leaves pilot-asset-* behind (58–84 KB measured, up to the whole tarball). SIGTERM during staging is not handled, because signal.NotifyContext is installed after staging: exit 143, and the file is left behind.
Timeout reported as success A call cut off by its 60 s deadline replies {"exit":-1,…} as a result instead of an error.

Fix (templates + 2 docs)

  • childproc_{linux,other}.go.tmpl (new, rendered for cli and hybrid): on Linux, every process the adapter starts (CLI calls, tar, install steps) gets Pdeathsig=SIGKILL. The forking thread stays locked until Wait returns. On other systems nothing is set. Children always stay in the adapter's process group, so the supervisor's group stop (http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39 in app-store) and reapStale's kill(-pgid) still reach them.
  • stage.go.tmpl: a Stager runs StageAssetsContext in the background. The runner resolves its command through it (Runner.SetCommandResolver), so a call waits for the binary within its own deadline. If an attempt fails, each call reports the failure, and a later call retries it (backoff from 2 s, doubling up to 5 min). The adapter no longer exits. Downloads go to $APP/.staged/tmp. That directory is removed when an attempt ends or is cancelled, and swept at the next start. Download, tar and install steps are bound to the ctx. The idle registry connection is closed after staging.
  • main.go.tmpl:
  • client_cli.go.tmpl: a cancelled call, or a child ended by a signal, is now an IPC error (backend: …: context deadline exceeded) instead of {"exit":-1}.
  • docs/CLI-ADAPTER.md and docs/R2-ARTIFACT-REGISTRY.md describe the new behaviour.

Tests

TestCLIAdapterChildLifecycleE2E covers timeout, SIGTERM, SIGKILL on Linux, and parent death with a call in flight. TestStageLifecycleE2E covers four cases: ready and help while the download is stalled; SIGTERM mid-download; the leftover from a SIGKILL mid-download is swept; a failed download keeps serving and a later call retries.

Verified on real miren bundles built from this branch

publish.BuildBundle built submissions/io.pilot.miren with version 0.1.1, for all 4 DefaultPlatforms, signed with a TEST-ONLY key. Nothing was uploaded or published. pilot-app verify passes for all 4, and the harness verify passes too: sha, signature, pin, and Mach-O/ELF arch. Apart from version, binary sha and signature, the manifest and install.json are identical to the published 0.1.0.

Real supervisor, first spawn (sup.spawn + sup.Call, nothing staged):

supervisor platform 0.1.0 this PR
app-store v1.0.3 darwin-arm64 never ready, app not ready ready after 675 ms; miren.help ok; miren.exec version ok after 9.3 s of staging
app-store v1.0.3 linux-arm64 never ready ready after 143 ms; exec ok after 25.6 s
app-store v1.0.2 darwin-arm64 / linux-arm64 — ready after 653 ms / 222 ms; exec ok

Harness (help ×200 + stop/orphan phases; miren.exec version ×200):

  • Ready in 558 ms on darwin-arm64, 2.2 s on darwin-amd64 (Rosetta first run, 1.5–1.8 s after that), 25 ms on linux-arm64 and 186 ms on linux-amd64. In the audit's harness runs of 0.1.0, the socket appeared after 6.8 s, 13.5 s, 10.8 s and 25.4 s.
  • help: 200/200 on all 4.
  • exec: 200/200 on darwin-arm64, darwin-amd64 and linux-arm64. linux-amd64 (emulated, host load about 78) did 141/141 before the harness's 3-minute budget ran out.
  • The fd count stayed flat. On darwin-arm64, help ×5000 went fd 11→9 once staging finished, and RSS levelled off at about 19.8 MB.
  • SIGTERM: exit code 0 in 1–40 ms, socket removed, process group empty.
  • Orphan without Pdeathsig: the app exits by itself 205–325 ms after its parent dies (fix(scaffold): generated adapters exit with their daemon and drop abandoned calls #111). With Pdeathsig on Linux: 0–10 ms.
  • Offline first spawn (sandbox-exec deny network / --network none): help 20/20 on all 4. miren.exec returns install assets: stage: fetch … as an error, and the app keeps running.

In-flight miren login (fake device-flow endpoint), then the app is stopped. The table shows whether the miren child is still alive afterwards:

stop darwin-arm64 0.1.0 → PR darwin-amd64 PR linux-arm64 0.1.0 → PR linux-amd64 PR
SIGTERM gone → gone gone gone → gone gone
SIGKILL of the pid (supervisor ≤ v1.0.3 stop) alive 4.1 s → alive 4.2 s alive 4.0 s alive 4.1 s → gone gone
SIGKILL of the group (app-store #39) gone → gone gone — → gone gone
daemon death, Linux Pdeathsig — — alive 4.2 s → gone gone
daemon death, no Pdeathsig orphan, child alive >42 s → gone, app exits in 220 ms gone (132 ms) orphan >42 s → gone (214 ms) gone (70 ms)

Real supervisor stop with that call in flight:

Not in this PR

  • macOS + a pid-only SIGKILL stop (every released supervisor): a SIGKILLed adapter cannot clean up, and macOS has no parent-death signal. The child stays in the app's process group, so app-store http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39's group kill ends it (verified above).
  • Grandchildren: a timed-out call kills its direct child only. Anything that child started stays in the app's process group, until the supervisor's group stop. On Linux, a grandchild gets no Pdeathsig. No curated miren method does this. miren server through miren.exec might, but that was not verified (it failed in an unprivileged container).
  • Shipping: miren 0.1.1 still has to be built with the Miren publisher key and uploaded to R2, and then the submission version bumped. None of that is done here.

🤖 Generated with Claude Code

…download is retried

Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it
still downloads the app's CLI before the adapter creates its socket.

The supervisor gives a new app 3s to create its socket, and waitReady never
looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a
cli app whose first-start download takes longer than 3s is never marked
ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every
call on that spawn fails with "appstore: app not ready: io.pilot.miren" until
the app respawns. With the real supervisor (spawn + Call), this reproduced
on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on
the first start, the adapter exits 1 before it listens, so even <ns>.help is
unavailable.

- stage.go: a Stager runs StageAssetsContext in the background. The runner
  resolves its command through it (Runner.SetCommandResolver), so a call
  waits for the binary within the call's own deadline. When an attempt fails,
  each call reports the failure, and a later call retries it (backoff starts
  at 2s and doubles up to 5 min) instead of the adapter exiting. The idle
  registry connection is closed after staging. tar and install steps get the
  same Pdeathsig and locked thread as CLI calls.
- main.go: the socket is bound before staging, for cli and hybrid apps that
  ship assets. After serve returns, main stops staging and waits (up to 5s)
  for an interrupted download to remove its partial file.
- The docs describe the new behaviour.

Tests:
- TestStageLifecycleE2E gains two subtests:
  - "socket and help are up while the binary is still downloading": the
    download stalls, the socket must accept within 2.5s, help must answer,
    and the call must wait and then run the binary.
  - "a failed download keeps serving and a later call retries".
  Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on
  linux/arm64 ("socket ... did not accept connections within 2.5s").
- TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its
  SIGKILL subtest on every OS (#114's child guard), and gives the timeout
  subtest a 4s deadline, because under heavy load a 2s deadline sometimes
  expired before the child had started.
- The lifecycle tests (these two, #111's and #114's) pass 10x on
  darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64
  (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and
  go vet ./... plus go test ./internal/... pass on linux/arm64.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@TeoSlayer

Copy link
Copy Markdown
Contributor Author

Aegis measurement against this head (0467324). I built io.pilot.aegis 0.1.5 from it: natives from pilot-protocol/aegis#4, a TEST-ONLY key, and the natives staged over HTTP. Then I SIGTERMed the app while aegis.exec install-models (a stub curl) was running, 3 rounds per platform.

🤖 Generated with Claude Code

@TeoSlayer TeoSlayer changed the title fix(scaffold): cli adapters are ready before their binary, and nothing they start outlives them fix(scaffold): a cli app is ready before its binary is, and a failed download is retried Sep 24, 2026
@TeoSlayer
TeoSlayer changed the base branch from fix/sixtyfour-lifecycle to fix/otto-lifecycle September 24, 2026 10:40
@TeoSlayer
TeoSlayer merged commit fabcce3 into fix/otto-lifecycle Oct 1, 2026
6 checks passed
TeoSlayer added a commit that referenced this pull request Oct 1, 2026
…tlive the adapter, cli apps are ready before their binary (#111, #114, #116)

* fix(scaffold): generated adapters exit with their daemon and drop abandoned calls

Every app built from main.go.tmpl (sixtyfour and the other http, hybrid
and cli apps) had four lifecycle faults. The harness sweep found all of
them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64,
darwin/amd64, linux/amd64 and linux/arm64:

- Orphans. The only exit path was SIGINT/SIGTERM. When the daemon dies
  hard on macOS (no Pdeathsig), or on Linux daemons built before
  app-store v1.0.3, the adapter is reparented to pid 1 and keeps serving.
  watchParent records the spawning pid first thing in main and shuts
  down cleanly once the parent pid changes.
- Abandoned calls. ipc.Serve ran handlers on the process ctx, so a
  caller that hung up (callFrom at its deadline, pilotctl at 120 s)
  left the backend request and its 2 fds open until the method timeout
  (60-280 s). serveConn is now the connection's only reader. It feeds
  ipc.Serve through a pipe and cancels the call's ctx on EOF.
- Crash on EMFILE. Under the Linux supervisor's RLIMIT_NOFILE=256,
  about 120 abandoned or concurrent calls made Accept fail, and
  log.Fatalf killed the app along with every in-flight call. Accept now
  backs off from 5 ms to 1 s and keeps going on EMFILE, ENFILE and
  ECONNABORTED.
- Unlinking another instance's socket. Closing the listener unlinked
  app.sock by name, so an orphan shutting down removed the socket of the
  instance the new daemon had just started. The listener no longer
  unlinks on close, and the adapter removes app.sock only when it is
  still the inode it bound.

TestGeneratedAdapterLifecycleE2E builds a real adapter and checks all
four. Every subtest fails on main and passes here: macOS -race x10,
linux/arm64 x5 and linux/amd64 x5.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(scaffold): a cli adapter's in-flight children die with it; a stop during staging leaves no partial download

Stacked on #111 (the adapter exits with its daemon; caller hang-up cancels
its call). Generated cli/hybrid adapters still leaked in three ways
(io.pilot.otto 0.20.0, measured on darwin-arm64, darwin-amd64 under Rosetta,
linux-arm64 and linux-amd64; the same template backs aegis, docker, duckdb,
miren, mysql, postgres, redis, smol, sqlite, tldr and upfile):

- A CLI child still running when the adapter was SIGKILLed (the supervisor's
  stop path) or killed by its own Pdeathsig (daemon death) was reparented to
  init and kept running: `otto.exec ["mcp","serve-http",...]` kept its port
  open 8 s later on every platform. A SIGTERM could also let the adapter exit
  before the child it had just signalled was gone.
- A stop during the first-spawn asset download (3.5-27.6 s for otto) left
  the partial pilot-asset-* in the daemon's shared TMPDIR: SIGTERM hit Go's
  default action because signal.NotifyContext was installed after staging.
- A call that hit its per-method timeout came back as a normal reply
  ({"exit":-1}) instead of an error.

Changes:
- client_cli: every CLI child stays in the adapter's process group and, on
  Linux, gets Pdeathsig=SIGKILL from a thread held until Wait returns; a
  cancelled call SIGTERMs the child before WaitDelay escalates; a deadline or
  shutdown is an IPC error, not a result.
- main: signal handling and the parent watch are installed before staging;
  shutdown waits (up to 8 s) for in-flight calls, and so their children.
- stage: downloads go to $APP/.staged/tmp and are cancelled with the adapter;
  whatever a SIGKILLed start left there is swept by the next start.

This consolidates the template fixes prototyped in the apps sweep for miren
(child Pdeathsig, drain, staging) and aegis (SIGTERM on cancel), with their
end-to-end tests, verbatim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(scaffold): a SIGKILLed cli adapter takes its in-flight children with it on macOS too

Linux kills an in-flight CLI child when the adapter dies (Pdeathsig, previous
commit). macOS has no parent-death signal, so when the adapter itself is
SIGKILLed (the app-store supervisor's stop path through v1.0.3 signals only
the adapter's pid; also the OOM killer or a crash) the child is reparented to
launchd and runs on. io.pilot.otto 0.20.0 on darwin-arm64 and darwin-amd64:
an in-flight `otto.exec ["mcp","serve-http",...]` still listened on its port
8 s after the adapter was SIGKILLed (ppid=1), and neither the next spawn's
reaper (it matches the adapter's argv) nor anything else would stop it. A
supervisor-side group stop (app-store #39) covers the stop path once a
daemon release ships it; this covers every daemon version and the other
ways an adapter dies, on both OSes.

While a CLI child is in flight the adapter keeps a guard: the same binary
re-executed with --pilot-child-guard, in the adapter's process group,
reading a pipe only the adapter can write (close-on-exec). The kernel closes
the pipe however the adapter goes away; on EOF the guard waits until it has
been reparented (the adapter really is gone) and SIGKILLs the process group,
itself included. A group id is not reused while a member lives, and the
guard is a member, so the signal cannot reach an unrelated group. It runs
only when the adapter leads its own group (as the supervisor spawns it),
retires 1 s after the last child is reaped, and is released and reaped on a
clean shutdown so the group is left empty.

TestCLIInflightServerE2E (the otto shape: a passthrough call running a
server) covers supervisor death with the supervisor's real spawn attributes,
SIGKILL of the adapter, a clean stop leaving an empty group, guard idle
retirement, the adapter surviving its guard's death, and a group stop. Against
origin/main it fails on darwin-arm64, linux-arm64 and linux-amd64; with the
guard disabled the SIGKILL subtest fails on darwin ("server child ... outlived
the SIGKILLed adapter"); all pass with this change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(scaffold): a cli app is ready before its binary is, and a failed download is retried

Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it
still downloads the app's CLI before the adapter creates its socket.

The supervisor gives a new app 3s to create its socket, and waitReady never
looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a
cli app whose first-start download takes longer than 3s is never marked
ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every
call on that spawn fails with "appstore: app not ready: io.pilot.miren" until
the app respawns. With the real supervisor (spawn + Call), this reproduced
on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on
the first start, the adapter exits 1 before it listens, so even <ns>.help is
unavailable.

- stage.go: a Stager runs StageAssetsContext in the background. The runner
  resolves its command through it (Runner.SetCommandResolver), so a call
  waits for the binary within the call's own deadline. When an attempt fails,
  each call reports the failure, and a later call retries it (backoff starts
  at 2s and doubles up to 5 min) instead of the adapter exiting. The idle
  registry connection is closed after staging. tar and install steps get the
  same Pdeathsig and locked thread as CLI calls.
- main.go: the socket is bound before staging, for cli and hybrid apps that
  ship assets. After serve returns, main stops staging and waits (up to 5s)
  for an interrupted download to remove its partial file.
- The docs describe the new behaviour.

Tests:
- TestStageLifecycleE2E gains two subtests:
  - "socket and help are up while the binary is still downloading": the
    download stalls, the socket must accept within 2.5s, help must answer,
    and the call must wait and then run the binary.
  - "a failed download keeps serving and a later call retries".
  Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on
  linux/arm64 ("socket ... did not accept connections within 2.5s").
- TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its
  SIGKILL subtest on every OS (#114's child guard), and gives the timeout
  subtest a 4s deadline, because under heavy load a 2s deadline sometimes
  expired before the child had started.
- The lifecycle tests (these two, #111's and #114's) pass 10x on
  darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64
  (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and
  go vet ./... plus go test ./internal/... pass on linux/arm64.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Teodor Calin <teodor@vulturelabs.io>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants