Image comprehension ollama
Skill aosama/image-comprehension-ollama/skills/image-comprehension-ollama
If you are a vision enabled model then you do not need this skill. Use this skill to analyze image files on disk via a local vision model (default: moondream:1.8b via Ollama). Invoke it whenever you encounter an image provided as a file path — such as a screenshot, photo, diagram, chart, or scan — and need a text description of its contents. The local model processes the image and returns a description to stdout, supplementing your workflow when images are not directly viewable in chat or when a local analysis is preferred. Supports PNG, JPEG, GIF, WebP, BMP.From its SKILL.md
npx -y skills add aosama/image-comprehension-ollama --skill image-comprehension-ollamaAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its file declares
Copied from the file, not written here
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
10.2 KB, ~2.3k tokens by cl100k_base, as published. Nobody here has run it
Image Comprehension with Ollama
Use this skill to analyze image files via a local vision model. When an image is provided only as a file path, or when you want a local description of a screenshot, photo, diagram, or other visual artifact, invoke this skill. It runs a local vision model that returns a text description, allowing you to act on visual information even when images are not displayed directly in the conversation.
When you encounter an image file — whether from a screenshot the user shared, a photo on disk, a rendered webpage capture, a chart, a diagram, or any other visual artifact — this skill lets you understand its contents. It runs a local vision model that analyzes the image and returns a text description to stdout. The model loads on demand, describes the image, and unloads immediately to free memory.
Use this skill proactively whenever visual information is relevant to your task and you cannot directly perceive it.
When to use this skill
- A user shares a screenshot and asks you to debug, review, or explain it
- You take a browser screenshot for QA and need to understand what rendered on screen (Playwright snapshots show structure; this skill shows you what it actually looks like)
- You encounter image files in the repo (diagrams, charts, photos, scans)
- You need to extract text from an image (OCR-like use case)
- You need to verify visual output (e.g., does the UI look correct?)
- Any situation where a human would look at an image and you need the same information
Prerequisites
- Ollama installed and running on your machine.
python3available on your PATH. No virtual environment is required.- A vision model pulled locally. By default the skill uses
moondream:1.8b(tiny, ~1.6 GB, works on CPU). Run once manually:
To use a different model, set theollama pull moondream:1.8bOLLAMA_VISION_MODELenvironment variable or pass--model.
Recommended models
| Model | Size | Quality | Best for |
|---|---|---|---|
moondream:1.8b | ~1.6 GB | Good | Default. Tiny, works on CPU, fast downloads |
minicpm-v:2.6 | ~2.5 GB | Very good | Better accuracy, still lightweight |
llava:7b | ~4.7 GB | Strong | High-quality descriptions, needs GPU for speed |
gemma4:e2b | ~7.2 GB | Excellent | Best quality for those with disk space and GPU |
Resolve the skill path first
Always call the script by absolute path. On this machine, and as a default convention, use:
$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh
Agent workflow
- Obtain or confirm the image path from the user, or from a screenshot you took.
- Decide what question to ask about the image:
- For general descriptions: use the default prompt.
- For specific information: provide a focused question via
--prompt. - For text extraction: ask about text content explicitly.
- Run the comprehension command.
- The description is printed to stdout (progress logs go to stderr).
- Use the returned description to complete your task — debug, verify, analyze, answer the user's question.
Default prompt
If you do not provide --prompt, the skill uses:
Describe this image in detail
Override with --prompt for specific questions about the image. Better prompts produce better descriptions — be specific about what you need to know.
Quick start
# Basic usage with default prompt and model
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image /path/to/image.png
# Custom question about the image
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image /path/to/screenshot.png --prompt "What text is visible in this screenshot?"
# Use a different model
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image /path/to/image.png --model llava:7b
# Use a different model via environment variable
OLLAMA_VISION_MODEL=llava:7b "$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image /path/to/image.png
# Analyze a chart or diagram
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image /path/to/diagram.png --prompt "Explain the flow and relationships in this diagram."
# Run the built-in smoke test
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --test
# Show help
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --help
Supported image formats
- PNG (
.png) - JPEG (
.jpg,.jpeg) - GIF (
.gif) - WebP (
.webp) - BMP (
.bmp)
Configuration
| Setting | CLI flag | Environment variable | Default |
|---|---|---|---|
| Vision model | --model | OLLAMA_VISION_MODEL | moondream:1.8b |
| Timeout | — | COMPREHEND_IMAGE_TIMEOUT_SECONDS | 180 |
- Model: Pass
--model <name>on the command line, or setOLLAMA_VISION_MODELin your environment. The CLI flag takes precedence. If neither is set,moondream:1.8bis used. - Timeout: Set
COMPREHEND_IMAGE_TIMEOUT_SECONDSto change the shell wrapper timeout (default 180 seconds). The inner Python script also enforces this as its maximum wait time for API calls.
Timing
Image comprehension typically takes 5 to 60 seconds depending on:
- Image complexity and resolution
- Question complexity
- Hardware capabilities (CPU/GPU)
Progress logs are printed to stderr during comprehension so coding agents can see that work is still progressing.
Options
--image <path>— Path to the image file to analyze (required, unless using--test).--prompt <text>— Custom prompt/question for the image (default: "Describe this image in detail").--model <name>— Ollama model to use (default:moondream:1.8b, override withOLLAMA_VISION_MODELenv var).--test— Run the built-in smoke test.--help— Show the help text.
Output
The image description is printed to stdout. Progress logs and error messages go to stderr.
To capture only the description:
description=$("$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image /path/to/image.png 2>/dev/null)
Concurrency guidance
Do not parallelize image comprehension requests. Each comprehension may monopolize GPU/CPU resources. Run comprehension one at a time and wait for each to complete before starting the next.
How it works
comprehend_image.shforwards all arguments tocomprehend_image.pyusing your existingpython3.- The Python script validates the image path exists.
- If Ollama isn't already running, it attempts to auto-start
ollama servefrom PATH. - It checks that the requested model is installed.
- It encodes the image as base64 and sends a POST request to the Ollama API with
keep_alive: "0". - The description is printed to stdout.
- The
keep_alive: "0"parameter unloads the model immediately after processing, freeing GPU/CPU memory.
If Ollama is not running or the model is not installed, the script prints a clear error to stderr and exits with a non-zero code. The agent should report the error to the user and suggest running ollama serve or ollama pull <model>.
Example prompts
| Use case | Prompt example |
|---|---|
| General description | "Describe this image in detail" (default) |
| Object detection | "What objects are present in this image?" |
| Text extraction | "Extract and transcribe all visible text in this image." |
| Chart analysis | "What does this chart show? Describe the trends and key data points." |
| UI/UX review | "Describe the user interface elements and their layout." |
| Document reading | "What is the content of this document? Summarize the key points." |
| Error diagnosis | "What error or issue is shown in this screenshot?" |
| Browser QA | "Does this webpage render correctly? Describe the layout, any visual errors, and whether the content matches what you'd expect." |
| Dark/light theme | "Is this page in dark mode or light mode? Describe the color scheme and any theme-related issues." |
How to test the skill
# Run the built-in smoke test (creates a minimal test image)
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --test
# Test with a real image
"$HOME/.agents/skills/image-comprehension-ollama/scripts/comprehend_image.sh" --image ~/Downloads/some-image.png
Caveats
- The vision model can hallucinate details, misread text, or miss elements. Treat descriptions as strong evidence, not ground truth. When precision matters (e.g., reading error messages, checking exact UI text), cross-reference with accessibility snapshots, page source, or other tools.
- Small text, thin fonts, and low-contrast regions are the most common sources of misreadings. For OCR-critical tasks, ask the user to confirm key values.
- The description is only as good as your prompt. A vague prompt like "what is this" will produce a vague answer. A specific prompt like "read the error message in the red banner at the top of this screenshot" will produce a targeted answer.
- The model unloads after each call. If you need to analyze multiple images, expect a short load time on each call.
Notes
- No API key is required — everything runs locally via Ollama's HTTP API.
- The Python code uses only the standard library.
- Progress logs go to stderr, description goes to stdout — easy to capture programmatically.
- If Ollama isn't already running, the script can auto-start
ollama serveifollamais on your PATH.
What ships with it: 2 files
15.7 KB alongside SKILL.md, 2 of them executable
scripts/
- comprehend_image.pyruns14.8 KB
- comprehend_image.shruns938 B