Vision Reader Skill
Enables pure-text LLMs (DeepSeek, etc.) to read binary files by calling Python tools.
Install Path
All tools are located at ~/opencode-vision-skill/tools/. The ~ means your home directory:
- Windows:
%USERPROFILE%\opencode-vision-skill\tools\ - macOS/Linux:
~/opencode-vision-skill/tools/
If you installed elsewhere, adjust the paths below accordingly.
Available Tools
| Tool | Command |
|---|---|
pdf-reader | python ~/opencode-vision-skill/tools/pdf-reader.py <filepath> |
image-reader | python ~/opencode-vision-skill/tools/image-reader.py <filepath> |
ppt-reader | python ~/opencode-vision-skill/tools/ppt-reader.py <filepath> |
screenshot | python ~/opencode-vision-skill/tools/screenshot.py [x,y,w,h] |
Always use absolute file paths for the filepath argument.
Tool Details
pdf-reader — Read PDF files
python ~/opencode-vision-skill/tools/pdf-reader.py "/absolute/path/to/file.pdf"
Returns Markdown with page headings, text content, and tables.
image-reader — Read and describe images
python ~/opencode-vision-skill/tools/image-reader.py "/absolute/path/to/image.png"
Returns three sections:
- Metadata — Format, mode, size, DPI
- Visual Description — AI-generated caption describing objects, layout, arrows, diagrams
- OCR Text — Recognized text from the image
ppt-reader — Read PowerPoint files
python ~/opencode-vision-skill/tools/ppt-reader.py "/absolute/path/to/slides.pptx"
Returns slide-by-slide text with titles, bullet points, and tables.
screenshot — Capture the screen
python ~/opencode-vision-skill/tools/screenshot.py
Captures full screen. For a region: python ~/opencode-vision-skill/tools/screenshot.py 0,0,800,600.
Returns the saved PNG file path.
Notes
- The first
image-readercall downloads Florence-2 model (~400MB) if not pre-downloaded - Tesseract OCR is optional — image-reader works without it using easyocr
- Screenshot requires a display server (X11/Wayland on Linux, native on Windows/macOS)