Roadmap

The Alpha proves the design on one Linux machine, with stand-ins for the agent and the network. What comes next puts real agents under vpnw on more systems, then leads to a first release, and after that to the parts a company would pay for. Durations are estimates, and what real agents turn up could stretch them.

Two lanes. The engine: Alpha, done, a working engine with run, trace, guard and learn on Linux; Beta, next, about 14 weeks for WireGuard, the file-system layer, macOS, packages and agent recipes; toward v1.0, 6 to 9 months of product work (estimate) for workspaces, a frozen event format, an outside review and pilots; v1.0, the first release, with the command line, policy files and events frozen. After v1.0, what companies would pay for: team policies, audit retention, managed exits and fleet configuration.

Stage Status What it covers
Alpha Done run, trace, guard and learn on Linux, with a sealed sandbox, tests and a demo
Beta Next WireGuard, a file-system layer, guard on macOS, packages, and real coding agents under vpnw for four weeks
Toward v1.0 Planned Workspaces, a frozen event format, an outside security review, pilots
v1.0 Planned A first release, with the command line, policy files and events frozen
After v1.0 Planned Team policies, audit retention, managed exits, fleet configuration

The Beta: Real Agents

The Beta has two goals. The first is an engine ready for real use: the pieces real agents and real networks need, on more systems, with every number from the Alpha measured again on each of them. The second is the demo for real: coding agents that developers use every day, working under vpnw guard for four weeks.

The Machines

The Alpha ran on one Linux machine with two CPUs and no IPv6. The Beta spreads out:

Machine System Why this one
A developer laptop Ubuntu 24.04 LTS, x86-64 The most common Linux desktop. AppArmor restricts user namespaces there, so it tests the profile vpnw needs to run without sudo
A second laptop Fedora, current release, x86-64 A newer kernel, SELinux and another package format
Raspberry Pi 5 Raspberry Pi OS, 64-bit The Alpha’s arm64 build compiles but has never run
A cloud virtual machine Debian 13, x86-64, with IPv6 The two IPv6 rows of the bypass matrix, and a real cloud metadata service at 169.254.169.254
CI runners Hosted Linux runners of GitHub Actions and GitLab CI Where many agents run with nobody watching
A Mac macOS on Apple silicon guard on macOS

Work Packages

  1. Packages. A Debian package, an RPM, a Homebrew formula, a signed static archive and an AppArmor profile, so that Ubuntu 23.10 and later allow vpnw’s sandbox without sudo. vpnw doctor checks each of these and says what is missing.
  2. The WireGuard path, and proxies over TLS. A third kind of path next to direct and proxy. vpnw would bring the WireGuard tunnel up inside its own process, in user space, with its own small network stack: no root, no new network interface on the machine, nothing changed for other programs. The likely base is wireguard-go and its user-space network stack. That would be the engine’s first third-party code, so it gets reviewed, pinned and weighed: the goal is a binary still under 10 MB. Proxies reached over TLS (https://), which the Alpha refuses, come in the same package.
  3. The file-system layer. Two limits the Alpha lacks. Which folders a program may read and write, so a sealed agent can’t read keys it has no business with, and which local sockets it may use, so a program that needs ssh-agent gets that one socket instead of --allow-unix-sockets lifting the filter for everything. Landlock, the kernel’s own unprivileged sandbox, is the likely base. The bypass matrix grows rows for files.
  4. guard on macOS. macOS has no network namespaces. The plan is the system sandbox that other agent sandboxes use on macOS, with a profile that lets the program connect to vpnw’s two local ports and nowhere else, and the same bypass matrix run against it. If it can’t give a boundary as tight as on Linux, guard keeps refusing to run on macOS, and the Beta report says why.
  5. Walls that report. The Alpha’s walls are silent: a program that tries to go around vpnw fails, but the record doesn’t show it. Linux can hand a refused system call to a supervising program, which would let vpnw record a refused Unix socket as an event. The Beta tries that, and looks for a way to count direct connection attempts too.
  6. Agent recipes. A reviewed policy and notes for each agent the Beta runs, and ready-made CI steps for GitHub Actions and GitLab CI that wrap a job in vpnw guard and keep its trace as a build artifact. Git over SSH, which ignores proxy settings, gets a documented route. The common tools agents call are tested under guard too: wget, git, pip, Python’s requests and httpx, Go’s net/http and Node.js.
  7. Measurements again. Start-up, cost per connection, throughput, memory, the bypass matrix and the planted bugs, on every Beta machine, plus the WireGuard path’s own figures. New planted bugs go into the new parts.

The Demo, for Real

Three coding agents that developers use every day, from different makers and built on different language runtimes, run under vpnw guard every working day for four weeks: on a laptop and in CI. Their policies start as learn drafts and are reviewed by a person.

Each agent also gets a poisoned task, like the one in the Alpha demo, carrying a canary token: a fake secret that is worthless but easy to spot if it ever arrives anywhere. The test passes only if the canary never arrives and the real work goes through. The Beta counts what matters to someone using vpnw every day: how often the policy got in the way of real work, how long a review of a learn draft takes, and what vpnw costs in time.

Timeline: About 14 Weeks

Weeks Work
1 to 2 Packages and the AppArmor profile. The Beta machines set up, and the Alpha’s tests and measurements run on all of them, arm64 and IPv6 included
3 to 6 The WireGuard path. The first agent under guard on a laptop, with its recipe
5 to 9 The file-system layer and walls that report. The CI steps for GitHub Actions and GitLab CI
7 to 10 guard on macOS, with its own bypass matrix
9 to 12 Four weeks of daily use by three agents, on laptops and in CI, and the fixes they call for
13 to 14 Beta release: packages, documentation, and a Beta report with every figure measured again

The Beta needs two to three people: a lead engineer, a second engineer who knows Linux isolation (namespaces, seccomp, Landlock), and part-time help for macOS and packaging. Apart from laptops and a Mac the team already has, the machines cost little: a Raspberry Pi, a small cloud machine for three months and CI minutes come to well under $1,000.

When the Beta Is Done

  1. The bypass matrix passes on every Linux machine, x86-64 and arm64, with both IPv6 rows run on a machine that has IPv6.
  2. guard on macOS passes its own bypass matrix, or stays switched off with the reasons published.
  3. Three real agents worked under guard for four weeks with reviewed policies, every run recorded, and each agent’s canary token blocked.
  4. The WireGuard path works with a real WireGuard server, with its throughput and cost per connection measured.
  5. The packages install on Ubuntu, Debian, Fedora and macOS, and vpnw doctor passes on each without sudo.
  6. Every Alpha figure measured again on each machine and published, and every planted bug, old and new, caught.
  7. Fuzzing runs every night with no open crash.

Out of scope for the Beta: Windows, a graphical interface, a hosted service, team features, and UDP and QUIC.

After the Beta: Toward v1.0

A first release should follow about 6 to 9 months after the Beta (estimate). Most of the work turns a tested engine into something a team can depend on:

  • Workspaces. A small file kept in a repository that names the path, the policy and the trace settings for its agents, so a developer, a CI job and an agent all run with the same network: vpnw workspace run agent-research -- ./agent. Secrets such as proxy passwords are referenced from elsewhere, not written into the file.
  • A frozen event format, with a written schema, so tools and audit systems can rely on it from one version to the next.
  • An outside security review of the sandbox and the broker, by people who do this for a living, published together with the fixes.
  • Pilots at companies running agents, on their own machines.
  • UDP and QUIC, if the pilots need them. Programs that can fall back to TCP do so today.
  • Windows, once the interfaces have settled.
  • v1.0 itself: from then on the command line, the policy files and the events stay stable, so scripts and tools built on them keep working.

After v1.0: What Companies Would Pay For

The local engine stays complete on its own: run, trace, guard and learn will never need an account or a hosted service. What a company would pay for is coordination across people and machines:

  • Team policies. One policy for every agent a team runs, kept and reviewed in one place.
  • Audit retention. Every trace kept, searchable and sent to the tools a company already uses.
  • Managed exits. Known exit addresses per team or per customer, run for them.
  • Fleet configuration. Paths and policies pushed to every machine that runs agents.

Later: More Uses

Coding agents come first, but the same pattern fits CI jobs and package installs, tool servers, browser agents, crawlers that need a regional exit, platforms that host other people’s agents, audit trails in regulated teams and everyday developer tools. None of them is tested yet. The post on next uses looks at what each would need.

Open Questions

  • The license. It will be chosen before the first public release. The plan is a proprietary engine, with an open-source branch under consideration.
  • The first third-party code. User-space WireGuard means depending on code the project didn’t write. The alternative, the kernel’s own WireGuard, needs root and changes the machine’s network. The Beta picks one and says why.
  • macOS. Whether the system sandbox can give a boundary as tight as Linux namespaces is an open question until the Beta’s bypass matrix has run there.
  • Pilots. The Beta’s agents run on the project’s own machines. Pilots on other people’s machines come after it, and the project would like to hear from teams who want to take part. See Contact.