Skip to content

fix(scaffold): generated adapters exit with their daemon and drop abandoned calls - #111

Merged
TeoSlayer merged 7 commits into
mainfrom
fix/sixtyfour-lifecycle
Oct 1, 2026
Merged

TeoSlayer merged 7 commits into
mainfrom
fix/sixtyfour-lifecycle

Conversation

@TeoSlayer

@TeoSlayer TeoSlayer commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Every adapter generated from internal/scaffold/templates/main.go.tmpl has four lifecycle faults. The harness sweep found all of them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64, darwin/amd64, linux/amd64 and linux/arm64. Rebuilding sixtyfour from main (fcd0b12, submissions/io.pilot.sixtyfour/case.json) reproduces them, so they come from the template, not from sixtyfour.

Published 0.1.0 (measured)
Orphan The only exit is SIGINT/SIGTERM. After its parent is SIGKILLed, the app is reparented to pid 1 and keeps answering IPC, on all 4 platforms. Linux daemons only catch this when built with app-store ≥ v1.0.3 (Pdeathsig); no released daemon pins that yet, and macOS has no Pdeathsig at all.
Abandoned calls ipc.Serve runs handlers on the process ctx, so a caller that hangs up (supervisor.callFrom at its deadline, pilotctl appstore call at 120 s) doesn't cancel anything. The backend request and its 2 fds stay open until the method timeout (60 s for find_email, 280 s for the intelligence methods). darwin-arm64: 30 calls abandoned after 0.3 s took fds from 7 to 67, and the backend still served every request.
EMFILE crash Under the Linux supervisor's RLIMIT_NOFILE=256, 121 (amd64) or 124 (arm64) abandoned calls, or 150 legitimately concurrent ones, crash the app: serve: accept: accept unix app.sock: accept4: too many open files → log.Fatalf → exit 1. Every in-flight call is lost (0/150 replies).
Socket hijack Closing the listener unlinks app.sock by name. After a hard daemon death, stopping the old orphan deleted the socket the new instance had just bound, so calls failed with FileNotFoundError while the new instance was still running.

Fix (template only)

  • watchParent: main records os.Getppid() before anything slow. A 250 ms ticker cancels the serve ctx once the parent pid changes, so the adapter shuts down the normal way (listener closed, socket removed). An adapter started with ppid ≤ 1 (for example a container's PID 1) doesn't watch.
  • serveConn: a goroutine is the connection's only reader and feeds ipc.Serve through an io.Pipe. When the caller sends EOF, it cancels the call's ctx, so the backend request, or a CLI child through exec.CommandContext, is aborted immediately. None of the in-tree callers (ipc.Call, supervisor.callFrom) half-close before reading the reply.
  • Accept backoff: EMFILE, ENFILE and ECONNABORTED back off from 5 ms to 1 s and keep serving, the same way net/http handles them, instead of log.Fatalf.
  • Socket ownership: SetUnlinkOnClose(false), and app.sock is removed only if it's still the inode this process bound (os.SameFile).

Tests

internal/scaffold/zz_adapter_lifecycle_e2e_test.go (TestGeneratedAdapterLifecycleE2E) generates and builds a real http adapter, points it at an httptest backend, and checks:

  1. It exits within 3 s of its parent being SIGKILLed and removes its socket.
  2. SIGTERM to an old instance leaves a newer instance's socket in place, and the newer instance still answers.
  3. A caller hanging up cancels the backend request, checked 5 times in a row.
  4. Under ulimit -n 32 with 40 waiting callers, the adapter hits EMFILE (the test asserts the log line) and keeps running. It serves again once the callers hang up. On macOS the kernel drops a pending connection when accept(2) fails with EMFILE, where Linux leaves it queued, so the test tolerates a caller being turned away.

Results:

  • On main all 4 subtests fail: adapter pid … still running 3s after its parent was killed (orphaned), old instance's exit removed the new instance's socket, backend request still open 3s after its caller hung up, adapter exited while out of descriptors: exit status 1.
  • On this branch: macOS arm64 -race -count=10 40/40; linux/arm64 and linux/amd64 (golang:1.25.13-bookworm) -count=5 20/20 each.
  • go test ./... passes on macOS. internal/scaffold and internal/publish also pass on linux/arm64.

Verified on real sixtyfour bundles (rebuilt from this branch)

I built case.json with the version bumped to 0.1.1 through publish.BuildBundle (go1.26.8, a TEST-ONLY key, nothing published) and ran the harness on every platform. darwin/amd64 ran under Rosetta, Linux in docker --init, both online and --network none:

  • sixtyfour.help 200/200 on all 4 platforms, online and offline. 10,000 calls on darwin-arm64: fd 7→7, RSS levels off at 18.4 MB.
  • find_email ×1000 through a local mock broker (SIXTYFOUR_BACKEND_URL): 1000/1000, fd 8→8. X-Pilot-Signature, X-Pilot-Caller and X-Pilot-Timestamp are present.
  • Orphan phase without Pdeathsig: SELF-EXIT after 240–325 ms on all 4 platforms (0.1.0: ORPHAN). The Linux Pdeathsig phase still passes (1–8 ms).
  • 150 calls abandoned after 0.3 s, with RLIMIT_NOFILE=256 on Linux: app fds stay at 7–9 (16–18 on amd64 under Rosetta). 150/150 upstream requests were cancelled by the app. No crash, and SIGTERM exits cleanly.
  • 150 concurrent held calls with NOFILE=256: the app survives. amd64: 150/150 ok. arm64: 124 ok, and 26 got an error reply (socket: too many open files dialing the backend), so every caller got an answer. Published 0.1.0: exit 1, 0 replies. On darwin (no cap): 150/150.
  • Socket race on darwin arm64 and amd64: B's app.sock survives both A's own exit and a SIGTERM to A, and B keeps serving. Published 0.1.0: app.sock exists=False … FileNotFoundError.

Not in this PR

  • The supervisor's stop path still SIGKILLs the app, so app.sock stays on disk after a stop. The adapter can't clean up after a SIGKILL; the supervisor has to send SIGTERM first. Measured with the real supervisor spawn + ctx cancel on darwin-arm64, 3 rounds each: app-store main 79e9944 and fix(supervisor): stopping an app stops the processes it started app-store#39 (SIGKILL to the group) give exit=-1, app.sock left behind=true. A SIGTERM-first stop with a SIGKILL fallback, stacked on http adapter: REST path params + PATCH/PUT/DELETE + broker pattern allow-list #39, gives exit=0 after ~20 ms, app.sock left behind=false for both this build and 0.1.0. That's app-store code and is handled separately. Until it lands, the adapter and the supervisor both delete a stale socket before Listen, so nothing breaks.
  • Shipping sixtyfour 0.1.1 needs the app's publisher key (ed25519:VoVCiQKPr73di2MlUd091a2Y6TCj/edSbCwRDtnYquI=), which isn't available here. Rebuilding from this commit with go1.26.8 is reproducible: the binary sha256 was identical across two builds on all 4 platforms. Note that a rebuild from current main also adds sixtyfour.balance to exposes, which comes from earlier template changes.

🤖 Generated with Claude Code

…ndoned calls

Every app built from main.go.tmpl (sixtyfour and the other http, hybrid
and cli apps) had four lifecycle faults. The harness sweep found all of
them in the published io.pilot.sixtyfour 0.1.0 bundles on darwin/arm64,
darwin/amd64, linux/amd64 and linux/arm64:

- Orphans. The only exit path was SIGINT/SIGTERM. When the daemon dies
  hard on macOS (no Pdeathsig), or on Linux daemons built before
  app-store v1.0.3, the adapter is reparented to pid 1 and keeps serving.
  watchParent records the spawning pid first thing in main and shuts
  down cleanly once the parent pid changes.
- Abandoned calls. ipc.Serve ran handlers on the process ctx, so a
  caller that hung up (callFrom at its deadline, pilotctl at 120 s)
  left the backend request and its 2 fds open until the method timeout
  (60-280 s). serveConn is now the connection's only reader. It feeds
  ipc.Serve through a pipe and cancels the call's ctx on EOF.
- Crash on EMFILE. Under the Linux supervisor's RLIMIT_NOFILE=256,
  about 120 abandoned or concurrent calls made Accept fail, and
  log.Fatalf killed the app along with every in-flight call. Accept now
  backs off from 5 ms to 1 s and keeps going on EMFILE, ENFILE and
  ECONNABORTED.
- Unlinking another instance's socket. Closing the listener unlinked
  app.sock by name, so an orphan shutting down removed the socket of the
  instance the new daemon had just started. The listener no longer
  unlinks on close, and the adapter removes app.sock only when it is
  still the inode it bound.

TestGeneratedAdapterLifecycleE2E builds a real adapter and checks all
four. Every subtest fails on main and passes here: macOS -race x10,
linux/arm64 x5 and linux/amd64 x5.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
teovl and others added 2 commits September 24, 2026 12:33
… during staging leaves no partial download

Stacked on #111 (the adapter exits with its daemon; caller hang-up cancels
its call). Generated cli/hybrid adapters still leaked in three ways
(io.pilot.otto 0.20.0, measured on darwin-arm64, darwin-amd64 under Rosetta,
linux-arm64 and linux-amd64; the same template backs aegis, docker, duckdb,
miren, mysql, postgres, redis, smol, sqlite, tldr and upfile):

- A CLI child still running when the adapter was SIGKILLed (the supervisor's
  stop path) or killed by its own Pdeathsig (daemon death) was reparented to
  init and kept running: `otto.exec ["mcp","serve-http",...]` kept its port
  open 8 s later on every platform. A SIGTERM could also let the adapter exit
  before the child it had just signalled was gone.
- A stop during the first-spawn asset download (3.5-27.6 s for otto) left
  the partial pilot-asset-* in the daemon's shared TMPDIR: SIGTERM hit Go's
  default action because signal.NotifyContext was installed after staging.
- A call that hit its per-method timeout came back as a normal reply
  ({"exit":-1}) instead of an error.

Changes:
- client_cli: every CLI child stays in the adapter's process group and, on
  Linux, gets Pdeathsig=SIGKILL from a thread held until Wait returns; a
  cancelled call SIGTERMs the child before WaitDelay escalates; a deadline or
  shutdown is an IPC error, not a result.
- main: signal handling and the parent watch are installed before staging;
  shutdown waits (up to 8 s) for in-flight calls, and so their children.
- stage: downloads go to $APP/.staged/tmp and are cancelled with the adapter;
  whatever a SIGKILLed start left there is swept by the next start.

This consolidates the template fixes prototyped in the apps sweep for miren
(child Pdeathsig, drain, staging) and aegis (SIGTERM on cancel), with their
end-to-end tests, verbatim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ith it on macOS too

Linux kills an in-flight CLI child when the adapter dies (Pdeathsig, previous
commit). macOS has no parent-death signal, so when the adapter itself is
SIGKILLed (the app-store supervisor's stop path through v1.0.3 signals only
the adapter's pid; also the OOM killer or a crash) the child is reparented to
launchd and runs on. io.pilot.otto 0.20.0 on darwin-arm64 and darwin-amd64:
an in-flight `otto.exec ["mcp","serve-http",...]` still listened on its port
8 s after the adapter was SIGKILLed (ppid=1), and neither the next spawn's
reaper (it matches the adapter's argv) nor anything else would stop it. A
supervisor-side group stop (app-store #39) covers the stop path once a
daemon release ships it; this covers every daemon version and the other
ways an adapter dies, on both OSes.

While a CLI child is in flight the adapter keeps a guard: the same binary
re-executed with --pilot-child-guard, in the adapter's process group,
reading a pipe only the adapter can write (close-on-exec). The kernel closes
the pipe however the adapter goes away; on EOF the guard waits until it has
been reparented (the adapter really is gone) and SIGKILLs the process group,
itself included. A group id is not reused while a member lives, and the
guard is a member, so the signal cannot reach an unrelated group. It runs
only when the adapter leads its own group (as the supervisor spawns it),
retires 1 s after the last child is reaped, and is released and reaped on a
clean shutdown so the group is left empty.

TestCLIInflightServerE2E (the otto shape: a passthrough call running a
server) covers supervisor death with the supervisor's real spawn attributes,
SIGKILL of the adapter, a clean stop leaving an empty group, guard idle
retirement, the adapter surviving its guard's death, and a group stop. Against
origin/main it fails on darwin-arm64, linux-arm64 and linux-amd64; with the
guard disabled the SIGKILL subtest fails on darwin ("server child ... outlived
the SIGKILLed adapter"); all pass with this change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…download is retried

Stacked on #114 (fix/otto-lifecycle). #114 makes staging ctx-aware, but it
still downloads the app's CLI before the adapter creates its socket.

The supervisor gives a new app 3s to create its socket, and waitReady never
looks again (app-store v1.0.2, which daemon v1.13.9 pins, and v1.0.3). So a
cli app whose first-start download takes longer than 3s is never marked
ready. io.pilot.miren 0.1.0 is one: 40-60 MB, socket after 11-22 s. Every
call on that spawn fails with "appstore: app not ready: io.pilot.miren" until
the app respawns. With the real supervisor (spawn + Call), this reproduced
on darwin-arm64 and linux-arm64 with v1.0.2 and v1.0.3. With no network on
the first start, the adapter exits 1 before it listens, so even <ns>.help is
unavailable.

- stage.go: a Stager runs StageAssetsContext in the background. The runner
  resolves its command through it (Runner.SetCommandResolver), so a call
  waits for the binary within the call's own deadline. When an attempt fails,
  each call reports the failure, and a later call retries it (backoff starts
  at 2s and doubles up to 5 min) instead of the adapter exiting. The idle
  registry connection is closed after staging. tar and install steps get the
  same Pdeathsig and locked thread as CLI calls.
- main.go: the socket is bound before staging, for cli and hybrid apps that
  ship assets. After serve returns, main stops staging and waits (up to 5s)
  for an interrupted download to remove its partial file.
- The docs describe the new behaviour.

Tests:
- TestStageLifecycleE2E gains two subtests:
  - "socket and help are up while the binary is still downloading": the
    download stalls, the socket must accept within 2.5s, help must answer,
    and the call must wait and then run the binary.
  - "a failed download keeps serving and a later call retries".
  Both fail on the base (80d19b2 + #114, 808243f) on darwin/arm64 and on
  linux/arm64 ("socket ... did not accept connections within 2.5s").
- TestCLIAdapterChildLifecycleE2E gains a parent-death subtest, runs its
  SIGKILL subtest on every OS (#114's child guard), and gives the timeout
  subtest a 4s deadline, because under heavy load a 2s deadline sometimes
  expired before the child had started.
- The lifecycle tests (these two, #111's and #114's) pass 10x on
  darwin/arm64, 5x on linux/arm64 and 1x on emulated linux/amd64
  (golang:1.25.13-bookworm). go test ./... passes on darwin/arm64, and
  go vet ./... plus go test ./internal/... pass on linux/arm64.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
fix(scaffold): a cli app is ready before its binary is, and a failed download is retried
fix(scaffold): a cli adapter's in-flight children never outlive it (io.pilot.otto)
@TeoSlayer
TeoSlayer merged commit fc52a9b into main Oct 1, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants