agentsclimarketplace

Run deep swe

Skill adriannoes/awesome-agentic-ai/cursor-claude-codex/skills/david-ondrej/agent-orchestration/run-deep-swe

329 agent skills (Cursor, Claude Code & Codex), 5,380 OpenClaw skills, 201 ML notebooks, 7 textbooks, 52 research papers, 17 industry reports for PMs, Designers & Developers.

Install
npx -y skills add adriannoes/awesome-agentic-ai --skill run-deep-swe

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Score any AI model on the DeepSWE coding-agent benchmark via the OpenRouter API. Use when the user wants an independent, reproducible coding-agent eval — "run DeepSWE", "benchmark this model on DeepSWE", "score model X on the coding benchmark", "test a model via OpenRouter on DeepSWE", or to verify vendor-reported coding scores. Covers setup, the OpenRouter wiring for mini-swe-agent, single-task / subset / full 113-task runs, and leaderboard submission.

SKILL.md

4.4 KB, as published. Nobody here has run it

Run DeepSWE via OpenRouter

DeepSWE (deepswe.datacurve.ai) is a 113-task Harbor-compatible coding-agent benchmark. It runs via Pier (Harbor fork) driving mini-swe-agent (model-agnostic). Any model reachable through OpenRouter can be scored.

Prerequisites — state-check first

which uv git docker || echo "MISSING: install uv, git, docker"
docker info >/dev/null 2>&1 || echo "MISSING: Docker daemon not running (Pier's default sandbox)"
echo "OPENROUTER_API_KEY set? ${OPENROUTER_API_KEY:+YES}"

Docker must be running — Pier sandboxes each task in Docker by default (--env modal for cloud instead).

David has a dedicated OpenRouter key for this benchmark exported globally in ~/.zshrc (weekly hard spend limit set as a safeguard). A fresh shell already has OPENROUTER_API_KEY available. If it's somehow not set, re-source the shell:

source ~/.zshrc && echo "key loaded? ${OPENROUTER_API_KEY:+YES}"

If still unset, ask David — never invent a key.

Setup

git clone https://github.com/datacurve-ai/deep-swe && cd deep-swe
uv tool install datacurve-pier            # PyPI (preferred)
# or: uv tool install git+https://github.com/datacurve-ai/pier
# pier bundles mini-swe-agent as the --agent driver

Run all pier commands from inside deep-swe/, using relative -p tasks/....

OpenRouter wiring (the part the docs don't spell out)

mini-swe-agent has a native OpenRouter model class. Both routes below use OPENROUTER_API_KEY and the OpenRouter slug (vendor/model, e.g. minimax/minimax-m3):

Route A — native OpenRouter class (preferred, hits openrouter.ai/api/v1 directly):

pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter

Route B — LiteLLM provider prefix (fallback; same key):

pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model openrouter/minimax/minimax-m3

Notes:

  • Slug = the exact OpenRouter slug. Verify it at openrouter.ai/models before running.
  • Free/zero-cost models: OpenRouter cost tracking can error. Set export MSWEA_COST_TRACKING=ignore_errors.
  • Flag spelling can vary by version — confirm with pier run --help and mini --help.

Smoke test FIRST (1 task — do this before any full run)

Always validate end-to-end wiring on a single task before spending tokens on the corpus:

pier run -p deep-swe/tasks/<task-id> --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter
# list available task ids:
ls deep-swe/tasks

Pass criteria: run completes, model returns actions (not auth/format errors), a score/trajectory is emitted. If it 401s → key wrong. If "provider not provided"/"model not mapped" → fix slug or switch route.

Subset run (deterministic sample)

pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter \
  --n-tasks 10 --sample-seed 0

Full 113-task corpus (costs tokens + time — confirm with user first)

pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter
# add `--env modal` to run in parallel Modal sandboxes (needs Modal configured)

Output & leaderboard

  • Trials land in jobs/<run>/<trial_id>/. Inspect with pier view jobs/<run>, pier analyze jobs/<run>, or pier critique run jobs/<run>.
  • Report: the exact command used, pass/fail, score, and any blockers.
  • Submit results for the official leaderboard to: <email-address>

Failure modes

SymptomCauseFix
HTTP 401bad/missing keyre-export OPENROUTER_API_KEY
"LLM Provider NOT provided"missing slug prefixuse Route B openrouter/... or Route A with --model-class openrouter
"model isn't mapped"/cost errorunknown cost for modelexport MSWEA_COST_TRACKING=ignore_errors
unknown flagversion driftcheck pier run --help

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.