Communitygithub.com

Liu-Zhangzhu/opencode-vision-skill

Give pure-text LLMs the ability to see — in one click.一键让纯文本大模型获得视觉能力。

opencode-vision-skill とは?

opencode-vision-skill is a OpenCode agent skill that give pure-text LLMs the ability to see — in one click.一键让纯文本大模型获得视觉能力。.

対応~Claude Code~Codex CLI~CursorOpenCode
npx skills add Liu-Zhangzhu/opencode-vision-skill

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

Vision Reader Skill

Enables pure-text LLMs (DeepSeek, etc.) to read binary files by calling Python tools.

Install Path

All tools are located at ~/opencode-vision-skill/tools/. The ~ means your home directory:

  • Windows: %USERPROFILE%\opencode-vision-skill\tools\
  • macOS/Linux: ~/opencode-vision-skill/tools/

If you installed elsewhere, adjust the paths below accordingly.

Available Tools

ToolCommand
pdf-readerpython ~/opencode-vision-skill/tools/pdf-reader.py <filepath>
image-readerpython ~/opencode-vision-skill/tools/image-reader.py <filepath>
ppt-readerpython ~/opencode-vision-skill/tools/ppt-reader.py <filepath>
screenshotpython ~/opencode-vision-skill/tools/screenshot.py [x,y,w,h]

Always use absolute file paths for the filepath argument.

Tool Details

pdf-reader — Read PDF files

python ~/opencode-vision-skill/tools/pdf-reader.py "/absolute/path/to/file.pdf"

Returns Markdown with page headings, text content, and tables.

image-reader — Read and describe images

python ~/opencode-vision-skill/tools/image-reader.py "/absolute/path/to/image.png"

Returns three sections:

  • Metadata — Format, mode, size, DPI
  • Visual Description — AI-generated caption describing objects, layout, arrows, diagrams
  • OCR Text — Recognized text from the image

ppt-reader — Read PowerPoint files

python ~/opencode-vision-skill/tools/ppt-reader.py "/absolute/path/to/slides.pptx"

Returns slide-by-slide text with titles, bullet points, and tables.

screenshot — Capture the screen

python ~/opencode-vision-skill/tools/screenshot.py

Captures full screen. For a region: python ~/opencode-vision-skill/tools/screenshot.py 0,0,800,600. Returns the saved PNG file path.

Notes

  • The first image-reader call downloads Florence-2 model (~400MB) if not pre-downloaded
  • Tesseract OCR is optional — image-reader works without it using easyocr
  • Screenshot requires a display server (X11/Wayland on Linux, native on Windows/macOS)

関連スキル