Harnesses
Prime Agent: a harness with one tool and a habit of rewriting itself
Prime Intellect's open-source harness hands the model a single persistent Python kernel and lets it edit its own skills. What that buys you, what the benchmark numbers do and don't say, and why it is not a sandbox.
- Published
- 29 Aug 2026
- Reading
- 9 min
- Class
- harnesses
half-life 90dfrom 29 Aug 2026
Most coding agents give the model a toolbelt: read a file, write a file, run a command, search the repository. Prime Agent, released by Prime Intellect in the first week of August under the MIT licence, does the opposite. It gives the model one tool — a persistent IPython kernel — and expects it to write Python for everything else.
It is worth understanding even if you never run it, because it is the clearest example so far of a design argument that will reach other harnesses: that the tool schema is the wrong abstraction, and code is the right one.
The design in two ideas
A recursive language model. Instead of stuffing history into the
prompt, context lives as variables inside the kernel and the model works on it
programmatically. Sub-agents are function calls: the agent launches one with an
asynchronous await rlm("task", name="…"), gets a handle straight back,
and collects the result later. Sub-agent sessions keep their own model, kernel and
history, so you can return to one after it has finished.
A continual harness. The agent can create, read, update and delete its
own prompts, memory, skills and sub-agent definitions, and the changes persist to disk
across sessions. The /refine command reviews the current trajectory and
applies small, evidence-backed edits. It never rewrites the immutable base system
prompt, and earlier refinements can be rolled back.
What you actually get to use
- Install and run. One command on Linux or macOS, then
prime-agentin a project directory and/loginfor your provider — subscription or API key, including local endpoints. - Background daemon. Sessions are owned by a daemon, so you can
detach and reattach, browse them with
prime-agent agents, and resume from the JSONL session log. - Autonomous mode. Goals, heartbeats and continuation, bounded by
--autonomous-max-turns,--autonomous-max-tokensand--autonomous-timeout-ms, with an optional gate command that has to pass. - Compaction without amnesia. Context is compacted at thresholds, but the full history stays reachable from code.
The numbers, read carefully
Prime Intellect reports 95.5% on ARC-AGI-3 with Opus 5, against a human expert baseline of 95.4%, and long-context results well ahead of another harness on the same model — for example 0.700 against 0.420 on OOLONG at 128k context, using GLM-5.2. Those are striking, and they are also the vendor's own runs on benchmarks the design was built for. Treat them as evidence the approach is worth trying, not as a ranking.
The authors are refreshingly direct about the biggest caveat: no model has yet been trained around this harness. If the gains hold up under independent replication, the training-in-the-loop version is the interesting one.
Not a sandbox
The project says plainly that Prime Agent runs model-generated Python with your user permissions and is not a security sandbox. Run it in a clean worktree or a disposable environment, and never point it at an untrusted repository on your main machine.
Who should try it
Long-horizon work where context management is the bottleneck — large codebases, data wrangling, research runs — and people comfortable reading the state the agent writes about itself. It is a poor first harness: if you want permissions, approvals and a settled workflow, start with Claude Code or Codex CLI. Our Prime Agent spec sheet has the fields side by side, and Anatomy of an agent harness explains the six pieces this design deliberately collapses into one.
Sources
- Prime Intellect: Prime Agent, a self-improving RLM agent — design, benchmarks and caveats, August 2026.
- PrimeIntellect-ai/prime-agent — install, commands, licence and the sandbox warning.