What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
If your LLM feature works but depends on a sprawling prompt that is hard to test or maintain, DSPy offers a different approach: describe the task with structured inputs and outputs, implement it as reusable Python modules, then evaluate candidate program configurations against a metric you choose. The important shift is from editing prompt text by intuition to improving an LLM program through an explicit evaluation loop. That process can help, but it does not guarantee better results; outcomes depend on the task, examples, metric, model, and evaluation setup.
What DSPy changes about prompt engineering
DSPy is a Python framework for building AI systems. Its central idea is to represent an LLM task and the code around it as a program, rather than treating one hand-written prompt as the whole application. DSPy describes itself as “a declarative way to build with LLMs.” In practice, you define task inputs and outputs, choose modules to perform the work, connect those modules into a program, and use an optimizer to tune selected parts against a metric.
This does not mean that prompts disappear. Instructions and demonstrations may still be part of the resulting program. The difference is that they are components the framework can configure and evaluate, rather than text you must continually edit by hand. DSPy’s overview introduces signatures, modules, and optimizers as the building blocks.
How signatures and modules fit together
Signatures describe the task
A signature names the fields an operation receives and the fields it should return. Instead of embedding every requirement in a long prompt, you state the shape of the task—for example, that a step takes a question and source text and produces an answer. The signature makes the interface explicit; it does not by itself ensure that an answer is correct or safe.
#1 Best Overall
Modules implement reusable steps
A module provides the behavior for a step in the program. The getting-started guide discusses Predict, ChainOfThought, and ReAct as module choices, and explains how to compose custom modules from independent stages. A simple prediction step and a multi-step workflow can therefore use the same programming model. See DSPy’s modules guide for the documented module concepts.
Modules can be called and composed with ordinary Python control flow. That makes it possible to express a larger application as a sequence of stages or other program logic, instead of forcing the entire task into a single prompt.
Rank #2
Build an evaluation-driven optimization loop
Optimization is meaningful only when you can define what “better” means for the application. A DSPy optimizer uses a program, a metric, and training inputs; depending on the method, it may use labeled examples, inputs without labels, or feedback. The metric is consequential: it determines which outputs the search rewards. A metric that checks only superficial similarity, for instance, may fail to penalize an answer that is invalid or harmful.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems- Describe the task. Create a signature with named input and output fields, focused on what the step does.
- Choose the implementation. Select a module for each step and compose the steps into a Python program.
- Establish a baseline. Run the unoptimized program on representative evaluation cases and record its scores and important failure modes.
- Define the metric and examples. Make the metric reflect application success, including relevant validity or safety requirements. Assemble representative training inputs and the labels or feedback required by your chosen method.
- Select what to optimize. Decide whether the likely opportunity is in examples, instructions, model weights, or program composition, then choose an optimizer suited to that change and your available evaluation signal.
- Compare candidates fairly. Evaluate the optimized candidate against the baseline on held-out cases that were not used to guide optimization. Inspect outputs as well as aggregate scores, and retain the metric, data split, model configuration, and program version with the result.
- Save and reload. Treat the selected program as a development artifact. DSPy’s getting-started tutorial covers saving and reloading optimized programs.
The official “Program, don’t prompt” guide walks through configuring a language model, declaring fields, creating a predictor, invoking it, and saving or reloading optimized programs. The official documentation describes the workflow; no benchmark result here establishes that optimization will improve a particular application.
Choose an optimizer by what should change
DSPy’s optimizer guide groups methods by the parts of a program they can tune. These approaches are not interchangeable: they can require different evaluation signals and compute, and the guide does not establish a universal cost ranking or a single best choice for every task.
| Approach | What changes | Examples named in DSPy’s guide | Evaluation signal and practical consideration |
|---|---|---|---|
| Few-shot demonstration construction | Examples supplied to guide the program | LabeledFewShot, BootstrapFewShot |
Demonstrations are central; the exact data requirements vary by method. Check the installed-version guide for the method’s requirements. |
| Instruction and demonstration search | Natural-language instructions and demonstrations | COPRO, MIPROv2, SIMBA, GEPA |
The guide describes MIPROv2 as proposing instructions and demonstrations, and GEPA as using reflection and textual feedback. Appropriate feedback and evaluation quality matter. |
| Fine-tuning | Underlying model weights | BootstrapFinetune |
This changes model weights rather than only prompt components. Consider the data, model support, and compute needed for the specific workflow. |
| Program combination | How programs are combined | Ensemble |
Combines programs rather than simply changing one instruction. Evaluate whether the combined behavior improves the application’s metric. |
The names and descriptions above follow DSPy’s optimizer documentation. Because the cited optimizer and overview pages are versioned differently, verify API names and examples against the version installed in your project.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Make the metric reflect real success
A score is evidence about the criterion you chose, not a general verdict on an LLM application. If a response must be factually grounded, correctly formatted, and safe, a metric that rewards only one of those properties can favor a candidate that performs poorly for users. Design the evaluation around the actual acceptance criteria, and inspect representative outputs for failure modes that a numerical score may conceal.
- Use examples that reflect the inputs and edge cases the application actually receives.
- Keep held-out cases separate from the data or feedback used to steer optimization.
- Compare baseline and candidate with the same model and evaluation setup unless changing those is itself part of the experiment.
- Record the metric definition and program configuration so later changes can be checked against the same baseline.
DSPy documentation explains optimization and scoring workflows, but it does not establish a guaranteed gain, a universally best optimizer, or a universal compute or cost ranking. A measured improvement applies to the metric and evaluation set used; it should not be presented as proof that every user or task will see the same result.
Best Value
Use DSPy as Python software
DSPy is a software framework used in Python, not a physical product. Its programs are configured with a language model, and optimization evaluates model-backed behavior. Installation details and APIs can change, so follow the instructions for the version you intend to use in the DSPy tutorials and verify examples against your installed release.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

