Gemini image
A curated collection of Claude Code skills — code-review loop, local chat-history search, session end/resume, Gemini image generation, transcript retrospectives — that run machine-wide with just python and node.
npx -y skills add BryceEWatson/claude-global-skills --skill gemini-imageAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Generate and edit images with Google's Gemini models (Nano Banana / Gemini 3 Pro Image / 3.1 Flash Image) from a zero-dependency Python CLI. Use whenever the user wants to create, generate, edit, restyle, or compose an image with Gemini / Google AI; produce listing photos, mockups, banners, icons, or marketing visuals via Gemini; or attach reference images for image-to-image editing. Also handles text/conversation/model-listing — it is the maintained successor to the older `gemini-client` plugin. Triggers: "generate an image with Gemini", "use Google AI to make a picture", "Gemini image", "nano banana", "edit this image with Gemini", "Imagen".
SKILL.md
5.8 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it
gemini-image
The maintained, machine-wide core for image generation and editing with
Google Gemini. One zero-dependency Python file — runs on any machine with
python, no per-project install. It is the canonical successor to the
gemini-client plugin (strict superset; that plugin is retired in its favour).
Quick start
GEM=~/.claude/skills/gemini-image/scripts/gemini_image.py
# Generate
python "$GEM" --image -o castle.png "A watercolor painting of a medieval castle"
# Pick a model / aspect ratio / resolution
python "$GEM" --image -o hero.png --model gemini-3-pro-image \
--aspect-ratio 2:3 --image-size 2K "Botanical oil painting of peonies"
# Edit or compose from reference image(s) — repeat -i (3-pro up to 8, 3.1-flash up to 14)
python "$GEM" --image -o staged.png -i room.png -i sofa.png \
"Place this sofa in this room with warm afternoon light"
# See which image models your key can use
python "$GEM" --list-models --images-only
Auth resolves in order: --api-key → GEMINI_API_KEY → GOOGLE_API_KEY →
.env (cwd, then next to the script). The key is sent as x-goog-api-key.
Models (verified against the live API, 2026-06)
| Model | Status | Use for |
|---|---|---|
gemini-3-pro-image / -preview | GA / preview | Auto-selected default — highest fidelity, accurate text rendering, up to 8 reference images. Mockups, design, data-viz. |
gemini-3.1-flash-image / -preview | GA / preview | Balanced (~half the cost, close to Pro on most prompts). Extreme aspect ratios (1:4…8:1), the 512 size, up to 14 reference images, video-to-image. |
gemini-2.5-flash-image | GA | Cheapest/fastest, lowest fidelity. The offline/error fallback floor. |
imagen-4.0-* | GA | Pure text-to-image via the separate :predict endpoint — not covered by this generateContent client. |
Aspect ratios: 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9 21:9 (3.1-flash adds the
extreme ratios). Sizes: 1K 2K 4K universally; 512 only on gemini-3.1-flash-image.
Default model: best available, auto-selected
Omit --model (CLI) or model= (library) and the skill resolves the best
image model your key can access — currently gemini-3-pro-image — instead of
a fixed id. It stays current automatically: a future GA Pro model (e.g.
gemini-4-pro-image) is adopted the day your key gains access, no code change.
A newer flash model never displaces a GA pro one (newer ≠ higher fidelity).
- Explicit always wins:
--model <id>/model="<id>"skips resolution entirely (and the/modelslookup). - Pin a default:
GEMINI_IMAGE_DEFAULT=<id>hard-pins (e.g. a cost-sensitive project pinsgemini-2.5-flash-image); explicit--modelstill overrides it. - Cost/speed: the best model costs more and is slower than the old default (approx. ~$0.13 vs ~$0.04/image, ~3× latency — Google list pricing, 2026-06). CLI use prints a one-line stderr notice when a non-default model is auto-selected; library use stays silent.
- Never fails: the resolution is cached per key for 24h (a temp-file keyed by
a hash of the API key — the key itself is never written). If
/modelsis unreachable it falls back togemini-2.5-flash-image, so a network blip can never break generation or cause a surprise bill.
What it does beyond the old plugin
- Reference-image input (
-i/--input-image, repeatable) — image editing and multi-image composition, which the plugin could not do. - Multi-image output — when the model returns several images they are all
saved (
out.png,out-2.png, …) instead of overwriting one file. - Safety-block diagnostics — a blocked/empty result reports the
blockReason/finishReason/ safety ratings instead of a bare "no image". - Verified model menu + exponential-backoff retry (429/5xx) + a
GeminiErrorexception contract so projects can vendor the file and build thin wrappers (CLI exit 2 = image blocked, 1 = error, 0 = ok).
Importable
generate, generate_image, resolve_best_image_model, extract_text,
extract_images, extract_usage, list_models, and GeminiError are public.
generate_image() returns
{output_paths, output_path, mime_type, text, blocked, diagnostic, raw}.
Notes
- Every generated image carries an invisible SynthID watermark (Google, non-optional).
- The saved file's extension matches the model's actual output format (the
image models choose it —
gemini-3-pro-imagereturns JPEG). A.pngrequest that comes back as JPEG is saved as.jpgwith a stderr note, so a file never lies about its contents. - The request uses
generationConfig.imageConfig— the shape proven across all current production callers and accepted by the live API. The newerresponseFormat.imageshape is tracked as a future migration inSPEC.md, which is the canonical contract this skill and the project-specific TypeScript callers all conform to. - Per-project business logic (brand denylists, print-spec validation, Sharp compositing, pack rules) stays in those projects — this core only does correct generation/editing.