fullauto.online

Harnesses

Sandboxes and self-writing agents: what isolation Prime Agent and Agent Zero give you

Both agents write and run their own code. One tells you outright that it is not a sandbox; the other wraps everything in a container. What each boundary protects, and what neither does.

Published
10 Sep 2026
Reading
9 min
Class
harnesses

half-life 90dfrom 10 Sep 2026

An agent that writes and executes its own code is powerful precisely because nothing limits what it can try. That is also the whole security problem. Two open-source projects, Prime Agent and Agent Zero, take opposite positions on who is responsible for the boundary.

Prime Agent: the boundary is yours

Prime Agent's only tool is a persistent Python kernel, and its documentation is unusually blunt: worker and kernel processes run with your local user permissions, and it is not a security sandbox. Anything the model can do as you, it can do. The advice is to use clean worktrees or restricted, disposable environments, especially with untrusted repositories.

That honesty is useful, and it puts the design work on you. A disposable container or VM, a throwaway user, a scoped API key and no access to your real credentials are all things you need to arrange before the first run.

Agent Zero: the container is the boundary

Agent Zero ships inside Docker. The agent gets a whole Linux desktop but not your computer, and its projects keep files, secrets and memory apart. That is a real boundary, and it is the reason the project is comfortable letting the agent install software and run arbitrary code.

The boundary has a door, though. The optional CLI connector and any mounted folders extend the agent's reach to your host, and the project's own guidance is to avoid mounting your whole home directory, keep credentials in project secrets and review anything that touches accounts, money or production.

What neither protects you from

  • Prompt injection. Text on a web page or in a document can carry instructions. A container does not stop an agent from obeying them inside it.
  • Data the agent can see. Whatever is inside the boundary — secrets, files, tokens — can be leaked by an agent that is tricked or simply wrong.
  • Network egress. Neither design promises to stop the agent talking to the internet. If it can reach a service, it can send data there.

A workable baseline

A disposable environment per task, a dedicated low-privilege account, credentials scoped to the single job, egress limited to what it needs, and a human approval on anything that spends money, sends messages or deletes data. Both projects work better inside that frame than outside it.

The pattern matters more than the product. Treat any agent that writes its own code as untrusted code you have chosen to run, and the rest follows. The same thinking applies to the tool-calling harnesses on the harness board, where permission prompts and sandboxes are the field to read first.

Sources

(p) => ` `