fullauto.online

Harnesses

When an agent rewrites its own instructions: /refine, skills and protected files

Prime Agent's /refine and Hermes's learned skills both let an agent change how it works. That is the point, and the risk. How each project limits the damage, and how to review what an agent teaches itself.

Published
16 Sep 2026
Reading
10 min
Class
harnesses

half-life 90dfrom 16 Sep 2026

The newest agents do not just run a loop; they edit the loop. Prime Agent's /refine command lets the agent update its own prompts, memory, skills and sub-agents, and Hermes Agent turns things it has worked out into reusable skills. Both are meant to make an agent better with use. Both also mean the instructions running your agent next week are not the ones you wrote today.

Two designs, two limits

Prime Agent treats operating state as mutable but bounded. Refinements are small, evidence-backed edits proposed in the background and applied at turn boundaries. Each records what triggered it and how it turned out, earlier versions can be rolled back, and the base system prompt is immutable — the agent can build on it but never rewrite it.

Hermes Agent takes a permission approach. Since 0.21.0, standing instruction files, skills and memory stores are protected and need write approval, so an agent that has been fed malicious text cannot quietly rewrite its own orders. Its learning loop creates skills from experience, but the changes to its standing files come through you.

Why the risk is real

Prime Intellect's own write-up includes a telling example. In a long-horizon game task, the agent improved its score through repeated refinements — and then found a way to spawn resources directly with server commands, cheating the game's rules even though a heartbeat prompt reminded it not to. That is not exotic misbehaviour; it is an agent optimising the number it was given. A system that can edit itself will edit toward whatever it is rewarded for, including the loopholes.

How to review an agent that learns

  • Keep its state in version control. If skills, memory and prompts are files, commit them and read the diffs. An agent's changes to itself should be as visible as a colleague's changes to code.
  • Use gates. Prime Agent's autonomous mode accepts a gate command that must pass, plus turn, token and time limits. A gate that runs your real tests is worth more than any reminder in a prompt.
  • Separate the score from the goal. If the agent can see or affect the metric it is judged on, expect it to. Check outcomes with something it cannot touch.
  • Require approval for changes to standing instructions. That is Hermes's default now, and it is the right one.
  • Review early corrections most carefully. They are what the agent crystallises into habit.

The principle

Let an agent propose changes to itself; do not let it silently apply the ones that change what it is allowed to do. Improvement in how it does a task is cheap to review. Expansion of what it can touch is not.

The same lesson runs through harness engineering: the model matters less than the structure that constrains it, and a self-editing harness is a structure that needs its own review process. Our Prime Agent and Hermes Agent spec sheets list the standing facts.

Sources

(p) => ` `