fullauto.online

← Frameworks

Framework · spec sheet

DSPy

Stanford NLP·optimizer
Category
Prompt optimiser
Type
Library
License
Open source (Stanford NLP)
Languages
Python
Focus
Compile prompts against a metric
Best for
Optimising instead of hand-writing prompts

Compiles prompts against a metric instead of asking you to write them.

What it is

DSPy is an open-source Python library from Stanford NLP that treats prompting as a compilation problem. Instead of hand-writing instructions and polishing them by eye, you declare what a step should take and return, give the program a metric and some examples, and let an optimiser search for the prompt (and the few-shot demonstrations) that score well against it. The note above is literal: it compiles prompts against a metric.

It is a library, not a service and not an agent harness. You import it, define modules, and run optimisation as an ordinary step in your own codebase. That framing — prompts as build artefacts rather than prose — is the whole pitch, and it changes who on the team is allowed to touch them.

How it works

A program in DSPy is a graph of modules. Each module wraps a step — a plain prediction, a chain-of-thought step, a retrieval call — and each step declares a signature: the typed inputs it consumes and the outputs it must produce. You compose modules into a pipeline the way you would compose functions, and the pipeline runs against whatever model provider you configure.

Then comes the part that separates it from prompt templates. You supply a small set of labelled examples and a scoring function — exact match, a rubric, an LLM judge, whatever you can defend — and an optimiser runs the pipeline, proposes instructions and demonstrations, scores the result, and keeps what worked. The output is a compiled artefact you can version and ship. Change the model or the data, recompile, and diff the results.

When it earns its place over a plain loop

When you have a metric and a handful of examples, and prompt churn is costing you time. The classic case is a pipeline that works at eighty per cent and is stuck there: hand-tuning has plateaued, and searching over instructions and demonstrations systematically beats another evening of adjectives. It also earns its place when several people edit the same prompts — a compiled artefact with a score attached is easier to review than a string someone tweaked at midnight.

Over a plain loop, the trade is setup cost. If you have no metric and no examples, DSPy has nothing to optimise against, and a single well-written prompt in a simple loop will outperform the scaffolding for a long time.

Limits

  • The metric is the ceiling. Optimisation will find whatever your scoring function rewards, including its blind spots. A weak metric compiles confidently wrong prompts.
  • Optimisation is not free. Searching over instructions and demonstrations costs model calls, and the search itself has settings worth understanding before you trust a result.
  • Debugging moves sideways. When the behaviour lives in a compiled artefact, the question “why did the prompt change?” becomes a diff and a score history rather than a memory of what you typed.
  • It is not an agent framework. Tool loops, permissions and session state are out of scope; DSPy improves the steps, it does not run them for you in production.

Alternatives on this board

  • Pydantic AI — when the problem is structured outputs and typed tools rather than prompt quality.
  • LangGraph — when you need the orchestration around the steps, not the steps compiled.
  • OpenAI Agents SDK — a lightweight loop with handoffs, if that is what you actually lack.

Sources

Hand-maintained editorial spec, not vendor copy — the read on each tool is judgement. Last checked 16 Sep 2026 · back to frameworks.