Gemini er
Skill and tool bundles for gap (graph as policy) — Anthropic Agent Skills format, discovered by path
npx -y skills add graph-robots/open-robot-skills --skill gemini-erAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Open-vocabulary 2D object detection via the Gemini Robotics-ER API — one call returns pixel-space bounding boxes with labels and scores for a text query. Use when a workflow needs a detection box to seed segmentation (e.g. sam3.segment_box) or coarse localization without any local GPU model.
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
2.4 KB, as published. Nobody here has run it
gemini-er
API-backed 2D detection on Gemini Robotics-ER. Zero GPU. The canonical
perception recipe (from the dev tree's perceive_gemini_er workflow script)
is: gemini-er.detect → best box by score → sam3.segment_box on the full
frame at that box → depth projection → OBB fit.
Install
uv sync --extra gemini-er # google-genai (pip: pip install -e ".[gemini-er]")
export GOOGLE_API_KEY=... # or GEMINI_API_KEY
Config
| Env | Meaning | Default |
|---|---|---|
GAP_GEMINI_ER_MODEL | Gemini model name | gemini-robotics-er-1.5-preview |
GOOGLE_API_KEY / GEMINI_API_KEY | API key (SDK default resolution) | — |
Contract
gemini-er.detect(image, query) returns
{"detections": [{"box": BoundingBox2D, "label": str, "score": float}]}:
boxis pixel-space{x1, y1, x2, y2}(top-left → bottom-right), clamped to the image bounds. The model emits the Geminibox_2dconvention ([ymin, xmin, ymax, xmax]normalized 0–1000); conversion happens here.scoredefaults to 1.0 when the model reports none — callers select the best detection withmax(..., key=score).- No match (or unparseable model output) → empty
detections, never an error. Treat empty as "object not visible".
When to use
- Detection boxes for open-vocabulary prompts, no local weights.
- Prefer
molmo.point_promptwhen a single click point is enough, andvlm.query_yes_nofor semantic checks without localization.