agentsclimarketplace

Molmo

Skill graph-robots/open-robot-skills/tools/molmo

Skill and tool bundles for gap (graph as policy) — Anthropic Agent Skills format, discovered by path

Install
npx -y skills add graph-robots/open-robot-skills --skill molmo

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Visual pointing and Q&A via the Molmo VLM served from a self-hosted vLLM endpoint (OpenAI-compatible API). Use when a workflow needs a single pixel coordinate for a named object (point_prompt) or Molmo-grade visual question answering and a vLLM server is available; for an API-only zero-GPU alternative use the gemini-er bundle.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

3.1 KB, as published. Nobody here has run it

molmo

Molmo visual pointing + Q&A. The bundle itself is zero-GPU (httpx client only), but Molmo has no hosted API — you must serve it yourself with vLLM and point GAP_MOLMO_BASE_URL at it. If you can't self-host, use gemini-er.detect (hosted Gemini Robotics-ER) instead.

Hosting recipe (vLLM)

Lifted from the dev tree's run book (training/README.md):

# Serve Molmo2-8B on an OpenAI-compatible endpoint
CUDA_VISIBLE_DEVICES=0 PYTHONNOUSERSITE=1 \
  python -m vllm.entrypoints.openai.api_server \
  --model allenai/Molmo2-8B \
  --trust-remote-code \
  --dtype bfloat16 \
  --port 8122 \
  --gpu-memory-utilization 0.5 \
  --max-model-len 4096 \
  --max-num-batched-tokens 4096

# Smoke-test
curl -s http://127.0.0.1:8122/v1/models | jq '.data[0].id'

# Point the bundle at it
export GAP_MOLMO_BASE_URL=http://127.0.0.1:8122/v1

Operational notes from the dev tree: pin Molmo to its own GPU when running alongside other perception services — under heavy parallel evaluation it becomes the throughput bottleneck if co-located; on a dedicated GPU you can push --gpu-memory-utilization 0.85 --max-num-batched-tokens 8192 for ~2× perception throughput. The server can also run on a remote machine and be port-forwarded in (the 4090 real-robot profile did exactly this).

Config

EnvMeaningDefault
GAP_MOLMO_BASE_URLvLLM OpenAI-compatible base URL— (required)
GAP_MOLMO_MODELModel name served by vLLMallenai/Molmo2-8B

Notes

  • molmo.point_prompt sends the canonical "Point at <query>" prompt and parses all four Molmo point output formats (Molmo2 <points coords=...>, Molmo1 <point x= y=>, legacy <points x1= y1= ...>, plain x, y fallback), converting normalized coordinates to pixels. found=False means the model emitted no parseable point.
  • molmo.query_yes_no coerces with the source-verbatim rule: answer is true iff "yes" appears in the lowercased reply.
  • Backend unreachable after 3 retries raises ToolError (route on_error).

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.