agentsclimarketplace

Dspy production deployment

Skill OmidZamani/dspy-skills/skills/dspy-production-deployment

Collection of Claude Skills for DSPy framework - program language models, optimize prompts, and build RAG pipelines systematically

Install
npx -y skills add OmidZamani/dspy-skills --skill dspy-production-deployment

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Use for deploying DSPy with save/load, configure_cache, restrict_pickle, track_usage, async execution, streaming, and production runtime controls.

SKILL.md

3.3 KB, as published. Nobody here has run it

DSPy Production Deployment

Goal

Prepare a DSPy program for repeatable, observable, scalable, and safer production execution.

Cache Hardening

DSPy enables memory and disk caches by default. Disk cache deserialization uses pickle unless restricted. Enable the allowlist mode in production:

import dspy

dspy.configure_cache(restrict_pickle=True)

Register trusted custom cache types only when needed:

dspy.configure_cache(
    restrict_pickle=True,
    safe_types=[MyResult, Metadata],
)

Disable a cache layer explicitly when a deployment cannot persist data or requires fresh model responses:

dspy.configure_cache(
    enable_disk_cache=False,
    enable_memory_cache=True,
)

Save and Load

Prefer state-only JSON for readable, safer artifacts:

compiled.save("./artifacts/program.json", save_program=False)

loaded = MyProgram()
loaded.load("./artifacts/program.json")

Use whole-program save only for trusted artifacts. It uses cloudpickle:

compiled.save("./artifacts/program/", save_program=True)
loaded = dspy.load("./artifacts/program/")

Keep the DSPy major version compatible when loading saved programs.

Usage Tracking

dspy.configure(
    lm=dspy.LM("openai/gpt-4o-mini"),
    track_usage=True,
)

prediction = program(question="What is DSPy?")
print(prediction.get_lm_usage())

Cached calls return no new token usage.

Async Execution

Most built-in modules support acall():

import asyncio

async def main():
    prediction = await program.acall(question="What is DSPy?")
    print(prediction.answer)

asyncio.run(main())

Implement aforward() for custom async modules. Use dspy.asyncify(program) only when adapting a synchronous callable is the right boundary.

Streaming

import asyncio
import dspy

stream_program = dspy.streamify(
    dspy.Predict("question -> answer"),
    stream_listeners=[
        dspy.streaming.StreamListener(signature_field_name="answer"),
    ],
)

async def main():
    async for chunk in stream_program(question="Explain DSPy briefly."):
        print(chunk)

asyncio.run(main())

For looped modules such as ReAct, set allow_reuse=True on listeners for repeated fields. Cache hits yield the final Prediction without replaying token chunks.

Production Checklist

  1. Pin the stable DSPy series.
  2. Use state-only JSON unless whole-program pickle is necessary and trusted.
  3. Enable restrict_pickle=True.
  4. Record usage, latency, errors, and traces.
  5. Load-test async and streaming paths separately.
  6. Use dspy-debugging-observability for MLflow and callbacks.

Official Documentation

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.