fullauto.online

← Harnesses

Harness · spec sheet

Codex CLI

OpenAI·coding
Category
Coding agent
Interface
Terminal
License
Open source
Languages
Any
Models
OpenAI (GPT-5.6 line)
Isolation
Sandboxed by default
Best for
Open-source + sandbox as hard requirements

Sandboxed by default and open source. The usual pick when those two are hard requirements.

What it is

Codex CLI is OpenAI's terminal coding agent: a local program that reads your repository, edits files, runs commands and iterates until the task is done or it runs out of rope. It is the same category as Claude Code, with two differences that show up in the spec above and account for most of the reason people pick it. The source is public, and the sandbox is on unless you turn it off.

The models are OpenAI's — the GPT-5.6 line as of this page's last check. The client being open source does not make the harness model-agnostic; it makes the loop auditable while the intelligence stays behind someone's API.

How the loop works

Standard agent cycle: the model asks for a tool, the CLI runs it inside the sandbox, the result goes back into context, repeat. Tools are file reads and edits, search, and shell execution, with MCP servers available for anything external. Project-level instructions live in a plain Markdown file at the repository root — the AGENTS.md convention, which several other agents on this board now read too, making it the closest thing to a portable standing-instructions format.

It runs interactively or non-interactively. The headless mode is what makes it a reasonable thing to put in CI: one command, a task, a diff or a patch out the other end, with the sandbox already doing the job you would otherwise need a disposable runner for.

Isolation and permissions

This is the strongest isolation story among the local agents here, and the reason for the rank. Commands execute under OS-level confinement — Seatbelt on macOS, Landlock and seccomp on Linux — with the writable area scoped to the working directory and network access off by default. Approval is a small number of named modes rather than a per-tool rule set: roughly read-only, edit-with-approval, and full access, with escalation prompts when the agent wants something the current mode forbids.

The practical consequences are worth knowing before you hit them. Network-off breaks package installs and anything else that fetches, so first runs in an uncached repository often fail in confusing ways until you widen the policy or pre-install dependencies. The usual mistake is reaching straight for full access to make the noise stop, which discards the entire advantage. Mode names and flags have changed more than once across releases; check the repository's docs rather than trusting a blog post, including this one.

Who it is for

Anyone for whom "I can read the harness" and "it cannot touch the rest of my machine" are requirements rather than nice-to-haves: regulated environments, shared machines, CI pipelines, and people running agents against code they did not write. Also the obvious choice if your organisation is already standardised on OpenAI models and billing.

Limits

  • OpenAI models only. Open client, closed brain. For genuine model choice see OpenCode or OpenClaw.
  • The sandbox costs you friction. Network restrictions are the single most common source of confusing failures, and the temptation to disable them is constant.
  • Fewer context features. Subagents, memory and long-session compaction are less developed than in the Anthropic line; long tasks hit the window sooner.
  • Terminal only. No first-party desktop or IDE surface here; editor integration means a separate product.
  • Churn. It has been rewritten and re-flagged repeatedly. Pin a version if you are scripting against it.

Alternatives on this board

  • Claude Code — richer context management and surfaces, proprietary, sandbox opt-in.
  • OpenCode — open source too, but provider-agnostic, if model lock is the thing you are avoiding.
  • Gemini CLI — the other big-vendor open-licensed terminal agent, strong on whole-repository context.
  • Editor agents — if the point of the sandbox was really just a small blast radius.

Sources

Hand-maintained editorial spec, not vendor copy — the read on each tool is judgement. Last checked 16 Sep 2026 · back to harnesses.