Image
Extract text from images using a vision LLMFrom its SKILL.md
npx -y skills add axoviq-ai/synthadoc --skill imageAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its file declares
Copied from the file, not written here
The file declares its own license as AGPL-3.0-or-later. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
2.1 KB, 419 tokens by cl100k_base, as published. Nobody here has run it
Image Skill
Base64-encodes the image and passes it to a vision-capable LLM that extracts
all text and key information. Returns the LLM's response as result.text.
Setup
No pip dependency — the skill uses only the Python standard library plus a
LLM provider you supply at construction time. The provider can be any object
that implements the complete() interface (see below).
Standalone usage
import asyncio
from synthadoc.skills.image.scripts.main import ImageSkill
# ImageSkill REQUIRES a vision-capable provider — calling extract() without
# one raises ValueError immediately.
skill = ImageSkill(provider=my_provider)
async def main():
result = await skill.extract("/path/to/screenshot.png")
print(result.text) # extracted text from the image
print(result.metadata) # {"tokens_input": N, "tokens_output": N}
asyncio.run(main())
Provider interface — any object with this async method:
async def complete(
messages: list, # list of Message objects from synthadoc.skills.base
system: str | None = None,
temperature: float = 0.0,
max_tokens: int = 4096,
) -> object # must have .text (str), .input_tokens (int), .output_tokens (int)
Build the provider with any vision-capable model. Message is importable
from synthadoc.skills.base — no dependency on synthadoc.providers:
from synthadoc.skills.base import Message
Supported image formats: .png, .jpg/.jpeg, .webp, .gif, .tiff
When this skill is used
- Source path ends with
.png,.jpg,.jpeg,.webp,.gif, or.tiff - User intent contains:
image,screenshot,diagram,photo
What ships with it: 2 files
2.2 KB alongside SKILL.md, 2 of them executable
scripts/
- __init__.pyruns0 B
- main.pyruns2.2 KB