Skip to content

Duplicate of #1849: missing SSH client during sandbox create attachment #4172

Description

@Stormxftw

Correction / Disposition

Withdrawing this report as a duplicate of #1849 after rechecking the local deployment and the complete earlier issue thread. The failure was real, but the original report overstated its novelty and incorrectly described parts of our custom configuration as NVIDIA's documented Compose setup. It also incorrectly treated the nonzero exit as necessarily wrong.

This correction replaces those claims. It does not claim an upstream fix has been shipped.

User Story

This report came from agent-run setup checks of a local OpenShell 0.1.2 deployment on Windows with Docker Desktop. The intended workflow was noninteractive sandbox creation followed by command execution. The assistant had mounted the Linux CLI into the gateway container at /opt/openshell/openshell; that mount and the associated PATH configuration were local customizations.

The upstream Compose example instead lists a workstation-installed CLI as a prerequisite. It does not ship our /opt/openshell mount. See the v0.1.2 Compose example.

Problem Statement

An attached sandbox create -- <command> invocation from our customized, minimal gateway container can successfully provision the sandbox and then fail to launch the local SSH attachment because ssh is absent from that container's PATH.

The observed message is unhelpful (No such file or directory (os error 2)), but a nonzero exit is legitimate when the requested attachment cannot start. Successful provisioning alone does not imply the entire foreground CLI invocation succeeded. --no-tty does not mean --detach.

This same missing-client failure was already described and assessed in #1849, including the triage analysis and the reporter's confirmation that installing an SSH client resolved their case. The original version of this report incorrectly said it had not been confirmed there.

Minimal Reproduction and Verified Results

In the customized local setup described above, with CLI and gateway version 0.1.2:

# Our minimal gateway container does not provide the external SSH client:
docker exec openshell-gateway ssh -V
# exec: "ssh": executable file not found in $PATH

# Requests a foreground attachment and fails locally after provisioning:
docker exec openshell-gateway /opt/openshell/openshell \
  sandbox create --name repro --cpu 1 --memory 2Gi --no-auto-providers \
  -- /bin/sleep infinity
# exit 1: No such file or directory (os error 2)

# The sandbox nevertheless reached Ready and gRPC exec works:
docker exec openshell-gateway /opt/openshell/openshell \
  sandbox get repro --output json
docker exec openshell-gateway /opt/openshell/openshell \
  sandbox exec -n repro --no-tty --no-login-shell -- /bin/echo LOCAL_WORKFLOW_OK
# exit 0; LOCAL_WORKFLOW_OK

# The documented detached workflow succeeds without an SSH client:
docker exec openshell-gateway /opt/openshell/openshell \
  sandbox create --name repro-detached --cpu 1 --memory 2Gi \
  --no-auto-providers --detach -- /bin/sleep infinity
# exit 0; sandbox Ready

# Remove only the test sandboxes:
docker exec openshell-gateway /opt/openshell/openshell \
  sandbox delete repro repro-detached

The recheck used unique test names and captured native process return codes directly, without a trailing echo masking the result. Test sandboxes were deleted and their absence verified.

At v0.1.2, run.rs skips attachment for --detach; the foreground path calls the SSH helper. ssh.rs constructs Command::new("ssh"), and its noninteractive wait path propagates a spawn error through into_diagnostic().

Impact / Local Resolution

We have corrected local automation to use the documented --detach mode, followed by sandbox exec, which uses the gRPC path. The test now passes without adding tools to the gateway image, changing authentication, or ignoring failures.

Additional checks confirmed that a repeated explicit name is rejected without replacing the existing sandbox, and a remote command exiting 7 still returns 7. The original suggestion that a retry with the same explicit name necessarily creates duplicate sandboxes was not supported and is withdrawn.

Acceptance / Remaining Scope

A clearer missing-SSH diagnostic could still be useful, but it is already part of #1849's discussion. No new upstream implementation request is being made here. Closing as a duplicate/local workflow correction, not as proof that the diagnostic has been fixed upstream.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions