Skip to content

feat(supervisor): netlink/syscall network setup so the privileged supervisor needs no workload-image tools #3280

Description

@akram

User Story

As an OpenShell operator, I want a sandbox to start with the default generic Alpine image in proxy mode without the sandbox image having to ship iproute2/nftables, so that the #3116 Alpine default works on unmodified base images.

Problem Statement

In proxy mode the privileged supervisor sets up the sandbox network namespace by shelling out to tools resolved from the workload image: ip (netns/veth/addr/route — 32 call sites in crates/openshell-supervisor-process/src/netns/mod.rs), nsenter (9), nft (6), and dmesg (bypass monitoring). A bare Alpine image only ships busybox ip (no netns subcommand) and no nftables/iptables, so the supervisor fails at startup with Network namespace creation failed ... iproute2 is installed and the sandbox container exits.

The community base image previously provided these tools; #3116 removes that dependency by defaulting to bare Alpine, which surfaces the gap. Only setns (8 call sites) is already a direct syscall today — namespace/veth/route creation and firewall rules still spawn external binaries.

This is the networking subset of #2750 (make the privileged supervisor independent of workload-image code), scoped down so it can land as a focused change that unblocks #3116.

Impact / Why This Matters

Without this, the #3116 default (bare Alpine) cannot run in proxy mode — the enforced-egress path that Secure Agent Workspace and any default-deny deployment rely on. The current workarounds are to keep shipping iproute2/nftables inside every sandbox image (re-introducing the exact image dependency #3116 removes) or to disable proxy mode (losing egress isolation). Neither is acceptable for a generic default.

Proposed Design

Replace the privileged network setup's external-helper calls with in-process kernel interfaces, so the supervisor is self-contained:

  • Namespace + veth + addresses + routes: route netlink (rtnetlink) plus setns/unshare with FD-owned namespaces (removes the /run/netns requirement).
  • Enter namespaces: setns directly (already used for the enter path).
  • Firewall / bypass rules: nf_tables netlink instead of the nft binary.
  • Bypass monitoring: NFLOG instead of dmesg (also drops the CAP_SYSLOG requirement; see feat(supervisor): NFLOG-based bypass detection to drop the CAP_SYSLOG requirement #2382).

The supervisor binary stays musl-static so it carries no dynamic loader or NSS dependency on the workload image. Behavior on base images that already ship iproute2/nftables must be unchanged.

Acceptance Criteria

  • A sandbox created from an unmodified docker.io/library/alpine:* image starts in proxy mode with no ip/nsenter/nft/dmesg executed from the workload image.
  • Network namespace, veth, addressing, and routing are created without spawning ip/nsenter.
  • Bypass-detection firewall rules are programmed without the nft binary.
  • Bypass monitoring works without dmesg / CAP_SYSLOG.
  • Verified on the Docker, rootless Podman, and Kubernetes runtimes.
  • No behavior change for base images that already ship the tools.

Alternatives Considered

Scope

In: the networking subset above.

Out (remains in #2750): the Phase 3 execution boundary (deny execve after startup, unprivileged workload/SSH execution), the Podman health-check and Kubernetes PVC-seeding shells, the hostile-image test suite, and the VM guest path.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions