Skip to content

Staging fails (exit 12) on lean base images that lack rsync #116

Description

@brandon-behring

Problem

runpod_deploy's workspace staging invokes rsync in SSH-transport mode
(rsync -az ... -e 'ssh ...' . root@host:/workspace/repo), which requires rsync
on both the local client and the remote container. Lean RunPod base images — e.g.
runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04 — do not ship rsync, so
staging fails immediately:

bash: line 1: rsync: command not found
rsync error: error in rsync protocol data stream (code 12)

No GPU compute is consumed, but the failure mode is non-obvious and aborts the run.

Repro

  1. Use runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04 as the pod image.
  2. Run run_job with any staging entry.
  3. Observe CalledProcessError (exit 12) at the workspace-push step.

Workaround (currently required)

Add a setup step to the job spec to install rsync before staging:

setup:
  - command: apt-get update -qq && apt-get install -y -qq rsync
    timeout_sec: 120

~10 s of pod time; works, but is boilerplate noise in every spec targeting a lean image.

Suggested fix (pick one)

  • (a) Document the rsync-on-remote requirement prominently in the quickstart/README.
  • (b) Auto-install rsync as part of the SSH-ready validation step before staging fires.
  • (c) Offer a --transfer-mode scp|tar fallback for images without rsync.

Found while running a one-shot LoRA sweep (prompt-injection-portfolio Lane 1).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions