Framework · spec sheet
OpenAI Agents SDK
- Category
- Orchestration
- Type
- Library
- License
- Open source (OpenAI)
- Languages
- Python, JS
- Focus
- Agents, handoffs, guardrails, sessions
- Best for
- Lightweight multi-agent apps
Lightweight multi-agent orchestration: agents, handoffs, guardrails and sessions in a small, sturdy API.
What it is
The OpenAI Agents SDK is an open-source Python and JavaScript library for building multi-agent applications out of a deliberately small set of primitives: agents, handoffs, guardrails and sessions. It is the grown-up version of an experimental project OpenAI circulated as Swarm, and the design bet is the same one — that most multi-agent systems need four concepts done properly, not forty done decoratively.
It is open source, so you can read the loop rather than trust it, but it is built OpenAI-model-first. The tight integration is the point and the trade: the abstractions line up with how OpenAI's models take instructions and tools, and switching vendors is a consideration rather than a given.
How it works
An agent is instructions, a model and tools — plain Python functions with typed schemas, or tools reached over protocols like MCP. Handoffs are the multi-agent part: one agent explicitly transfers control to another, passing the conversation along, which is cleaner than two agents pretending to email each other inside one context. Guardrails wrap a run with input and output checks that can stop it before or after the model speaks, and sessions keep conversation state across turns so you are not rebuilding history by hand.
The surrounding machinery is thin on purpose: a runner that executes agent loops, agents-as-tools for delegation that returns a result rather than a handoff, and tracing that records what actually happened across agents and tool calls. That last one matters more than it sounds — multi-agent runs are hard to debug precisely because the interesting work happens between the prompts.
When it earns its place over a plain loop
When you want specialised agents with clean transfers of control — a triage agent that hands off to a specialist, an agent that uses another as a tool — and you want that without adopting a state machine or a role-play framework. The API is small enough to learn in an afternoon, and guardrails plus sessions cover the two things every plain loop grows by hand: safety checks and message history.
If one agent with a good tool set does the job, a plain loop is still smaller. The SDK earns its dependency when handoffs or guardrails are requirements, not when they sound nice.
Limits
- OpenAI gravity. Model-agnosticism is not the design centre; treat non-OpenAI backends as something to verify case by case before committing.
- Guardrails are code, not a boundary. They catch what you anticipated and shaped checks for; they do not contain a determined or unexpected failure.
- Durability is thin. Sessions hold state, but long-running, resumable, checkpointed work is not the primitive here — that is what LangGraph is for.
- It moves quickly. The API is young and evolving; pin versions and read the release notes rather than trusting any write-up, this one included.
Alternatives on this board
- LangGraph — when the flow is a graph and you need checkpointing and interrupts.
- Google ADK — code-first multi-agent with deployment paths on the other cloud.
- Pydantic AI — typed tools and outputs first, provider choice second.
- CrewAI — when the split is into roles with backstories rather than handoffs.
Sources
- OpenAI Agents SDK documentation — the place to verify primitives, tool wiring and current API names.
- OpenAI — openai-agents-python repository — source, examples and release history.
- Anatomy of an agent harness — what the loop is doing underneath, vendor-neutral.
- Framework leaderboard for the full comparison.
Hand-maintained editorial spec, not vendor copy — the read on each tool is judgement. Last checked 16 Sep 2026 · back to frameworks.