Framework · spec sheet
DSPy
- Category
- Prompt optimiser
- Type
- Library
- License
- Open source (Stanford NLP)
- Languages
- Python
- Focus
- Compile prompts against a metric
- Best for
- Optimising instead of hand-writing prompts
Compiles prompts against a metric instead of asking you to write them.
What it is
DSPy is an open-source Python library from Stanford NLP that treats prompting as a compilation problem. Instead of hand-writing instructions and polishing them by eye, you declare what a step should take and return, give the program a metric and some examples, and let an optimiser search for the prompt (and the few-shot demonstrations) that score well against it. The note above is literal: it compiles prompts against a metric.
It is a library, not a service and not an agent harness. You import it, define modules, and run optimisation as an ordinary step in your own codebase. That framing — prompts as build artefacts rather than prose — is the whole pitch, and it changes who on the team is allowed to touch them.
How it works
A program in DSPy is a graph of modules. Each module wraps a step — a plain prediction, a chain-of-thought step, a retrieval call — and each step declares a signature: the typed inputs it consumes and the outputs it must produce. You compose modules into a pipeline the way you would compose functions, and the pipeline runs against whatever model provider you configure.
Then comes the part that separates it from prompt templates. You supply a small set of labelled examples and a scoring function — exact match, a rubric, an LLM judge, whatever you can defend — and an optimiser runs the pipeline, proposes instructions and demonstrations, scores the result, and keeps what worked. The output is a compiled artefact you can version and ship. Change the model or the data, recompile, and diff the results.
When it earns its place over a plain loop
When you have a metric and a handful of examples, and prompt churn is costing you time. The classic case is a pipeline that works at eighty per cent and is stuck there: hand-tuning has plateaued, and searching over instructions and demonstrations systematically beats another evening of adjectives. It also earns its place when several people edit the same prompts — a compiled artefact with a score attached is easier to review than a string someone tweaked at midnight.
Over a plain loop, the trade is setup cost. If you have no metric and no examples, DSPy has nothing to optimise against, and a single well-written prompt in a simple loop will outperform the scaffolding for a long time.
Limits
- The metric is the ceiling. Optimisation will find whatever your scoring function rewards, including its blind spots. A weak metric compiles confidently wrong prompts.
- Optimisation is not free. Searching over instructions and demonstrations costs model calls, and the search itself has settings worth understanding before you trust a result.
- Debugging moves sideways. When the behaviour lives in a compiled artefact, the question “why did the prompt change?” becomes a diff and a score history rather than a memory of what you typed.
- It is not an agent framework. Tool loops, permissions and session state are out of scope; DSPy improves the steps, it does not run them for you in production.
Alternatives on this board
- Pydantic AI — when the problem is structured outputs and typed tools rather than prompt quality.
- LangGraph — when you need the orchestration around the steps, not the steps compiled.
- OpenAI Agents SDK — a lightweight loop with handoffs, if that is what you actually lack.
Sources
- DSPy documentation — the place to verify module names, optimisers and current API shapes.
- Stanford NLP — DSPy repository — source, examples and the licence text.
- Evals that catch regressions — why a metric has to exist before any of this pays.
- Framework leaderboard for the full comparison.
Hand-maintained editorial spec, not vendor copy — the read on each tool is judgement. Last checked 16 Sep 2026 · back to frameworks.