Sam3
Skill and tool bundles for gap (graph as policy) — Anthropic Agent Skills format, discovered by path
npx -y skills add graph-robots/open-robot-skills --skill sam3Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Segment Anything 3 — text-, point-, and box-prompted instance segmentation, plus a stateful streaming video tracker that carries object identity through SAM3's memory bank. Use when a workflow needs open-vocabulary masks from an RGB image or needs to follow one object across frames.
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
3.8 KB, as published. Nobody here has run it
sam3
The SAM3 image servicer + video-tracker servicer as in-process tools. Images
are RGB uint8 [H, W, 3] numpy arrays; masks come back as gap Mask
(uint8 [H, W], 0 background / 255 foreground), score-sorted best-first.
When to use
segment_textfor open-vocabulary "find the X" masks (one mask per instance; checkscores[0]— callers typically reject below ~0.3).segment_boxafter a detector (e.g.grounding-dino.detect) for a pixel-accurate mask inside the detection box; add the point prompt (use_point=True) when a pointing model supplies one.tracker_init/tracker_update/tracker_closeto follow a single target across an observation stream (e.g. for visual servoing).
Install
uv sync --extra sam3 # torch + torchvision + the upstream sam3 package
# (pip: pip install -e ".[sam3]")
Model weights download on first model build. Device is taken from
GAP_SAM3_DEVICE (default cuda); the image model also runs on cpu
(slow), the video tracker is CUDA-only in practice.
Gotchas (carried over from the servicers)
- Lazy singletons: the image model and the video predictor each load on first call and stay resident; importing the bundle never imports torch.
segment_textcaps results atmax_results=5by default — cluttered scenes emit 100+ instances (~1 MB/mask at 720p) and downstream consumes only the top mask. Passmax_results<=0for everything.- The video tracker JIT-compiles Triton NMS kernels via the
CCenv var; a staleCC(e.g. a Ray env pointing at a non-existent gcc-13) surfaces asFileNotFoundErrorinsidetracker_init. The bundle forcesCCto a real compiler before tracker use (_ensure_cc_compiler). - Tracker prompt precedence is box > point > text; a point prompt is converted to a small (10% of image) box because the predictor's box path is more reliable for init than a single point.
- The tracker is built with
apply_temporal_disambiguation=False— the default hotstart heuristics silently delete the masklet around frame 3 in streaming mode (no fresh text re-detection per frame). - Drift handling in
tracker_update: a mask-area jump >1.5x the running median or confidence <0.30 keeps the LAST GOOD mask and reportsconfidence=0.0withobject_present=True(skip this frame); after 5 consecutive drift hitsobject_present=False— re-init the tracker. - Sessions idle longer than 120 s are evicted lazily on the next tracker
call; an evicted/unknown
tracker_idraisesToolError. tracker_initreturnsobject_present=Falsewith an emptytracker_id(no exception) when the initial detection finds nothing.