Community影像github.com

ainotate

Use when an answer depends on where something is or how it looks on a screen - "where is X", "how do I do Y in this app/site", explaining a UI or a screenshot the user sends, reporting a bug or writing a ticket (Jira, GitHub, Freshdesk, Slack) about something visible, showing before/after of a change, writing a step-by-step guide, or any reply that would otherwise say "top right, under the..." in words. Also when the user says "ainotate", "show me", "take a screenshot", "annotate".

ainotate 是什麼?

ainotate is a Claude Code agent skill that use when an answer depends on where something is or how it looks on a screen - "where is X", "how do I do Y in this app/site", explaining a UI or a screenshot the user sends, reporting a bug or writing a ticket (Jira, GitHub, Freshdesk, Slack) about something visible, showing before/after of a change, writing a step-by-step guide, or any reply that would otherwise say "top right, under the..." in words. Also when the user says "ainotate", "show me", "take a screenshot", "annotate".

相容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/dontpayfull/AInotate/tree/main/skills/ainotate

Installed? Explore more 影像 skills: steipete/songsee, affaan-m/frontend-design, mattpocock/improve-codebase-architecture · View all 6 →

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

ainotate 是做什麼的?

Show, don't describe: an annotated screenshot plus short text that refers to the marks by number. Plain-text or code answers: skip it.

When to reach for it unprompted

  • Explaining where a control or setting is on a website or app.
  • Answering "where is / how do I" questions about a UI.
  • Reporting a bug found during browsing, testing, or scraping.
  • Whenever you are about to use spatial words ("top right", "under the second menu", "left sidebar").
  • Verifying UI changes: before and after.
  • A step only the user may take: a permission switch, a 2FA code, a payment or consent screen, a CAPTCHA. Do not click it for them: show exactly what to press, and say what the click does.

CLI ainotate (fallback python -m ainotate); ainotate doctor checks the machine. Every spec option: references/spec.md. Exit codes: 2 bad spec, 3 cannot draw, 4 capture failed or a dependency is missing (the message says what to install), 5 text matches several places (ask which, or pass nth), 6 target not found (closest text listed). Fix the spec, never work around.

1. Pick a source

SourceHow
Public web pageainotate shoot spec.json: capture, find targets, redact, annotate, save
Logged-in pagethe same spec with "cdp_url" + "page_url_contains" (user's running browser). Read references/browsers.md for every browser surface
User's screenshottheir file; clipboard: ainotate clipboard
Desktop appainotate windows <app>, then ainotate window <id> (macOS: add --target for exact control boxes, below); ainotate screen

Never move or resize the user's windows; capture and crop.

2. Locate - never guess coordinates from a preview

The image you see when you Read a file is downscaled; positions read off it are wrong.

  • Web: "target" on each mark (shoot resolves it), or ainotate capture URL --target save='{"text": "Save"}' for rects in CSS px plus the scale. Targets: {"text"}, {"selector"}, {"role", "name"}, optional within, nth, exact.
  • Desktop apps on macOS: ainotate window <id> --target allow='{"element": "Allow", "role": "button"}' returns rects in window points plus the scale (Accessibility; exact even for icons). ainotate elements --app "System Settings" lists the labels it can match.
  • Images with text: ainotate locate img.png "Save" (OCR), or "target": {"text": "Save"}.
  • Icons: ainotate grid img.png, then ainotate zoom img.png x1 y1 x2 y2 (fine grid in source px; one zoom per element, under ~500px wide). ainotate info img.png gives the real size.

Units, one rule: every coordinate (rect, crop, label_at, at, pad) is in one unit; scale converts it to image pixels. DOM rects: dicts {x,y,w,h} with the returned scale. OCR/grid/zoom pixels: lists [x1,y1,x2,y2] with scale: 1. Never mix.

3. Annotate

{"url": "https://news.ycombinator.com", "name": "hn header",
 "marks": [
  {"type": "step", "n": 1, "target": {"text": "new", "exact": true}, "label": "Newest"},
  {"type": "step", "n": 2, "target": {"text": "login"}, "label": "Sign in"}]}

ainotate shoot spec.json --draft. "actions" (click, fill, wait, scroll_to) run before measuring, e.g. to open a menu. Existing image: same marks with "rect", plus "input" and "scale", then ainotate annotate spec.json --draft. Presets: laptop, wide, phone.

typeuse
stepnumbered badge + outline; n matches your text
box"this thing", with a label
arrowlabel pointing at a target, no outline
click / keysclick ripple / keycaps like ["⌘", "K"]
magnifyloupe that enlarges a small control
highlight / spotlight / texttint / dim the rest / free note
redactsolid block for anything sensitive
blur / pixelatedecorative only; secrets always redact

Placement: leave labels to the engine. It puts each label in free space near its target, plans all labels together and curves arrows around text; it covers text only when the frame has no room, and then warns. Fix that warning by widening the crop toward empty space (margins, gutters) or shortening the label; use label_at only as a last resort, and never onto text. Crop: findable in 2 seconds. "auto" (default) frames the marks; an explicit crop when a landmark matters (header, sidebar) and always when redacting. Far-apart marks: one image each. Frame (guides, docs): "frame": {"chrome": "browser", "bg": "ocean"}; tickets stay plain. Privacy: shoot redacts emails, phones, cards, keys, tokens and password fields by default; on images only with "privacy": "auto" (OCR). Best-effort: you still verify. Colors: look orange (default), bad red (only wrong/bug), good green, info blue. Max 6 marks, labels 1-4 words in the audience's language, UI text quoted verbatim. User's own marks: keep them, use a color they did not, say which are yours.

4. Verify (required)

Read the output image. Each mark sits on its element, labels and arrows do not cover text that matters, and no sensitive data shows anywhere in the frame, marked or not. --debug draws the exact rects (cyan) and label boxes (magenta). Fix and rerun; then save without --draft.

5. Deliver

Saves go to the output folder (default ~/Pictures/AInotate), never overwriting; --copy also puts the image on the clipboard.

  • Chat: image path on its own line, then 1-3 sentences: "At ① open the cart; ② ...".
  • Ticket: Summary, Steps (numbers match badges), Expected vs Actual, URL/env, image, in the ticket's language.
  • Guide: one image per step, same preset and crop width, then ainotate guide steps.json. Before/after: ainotate compare a.png b.png. Flows: ainotate animate (APNG or GIF).

Patterns and mistakes

  • Tables: box the whole row, label beside it via label_at, "arrow": false.
  • Dense pages: a "covers page text" warning means no free room. Widen the crop first; then label_at into empty space next to its own target. Crossing arrows confuse more than they help.
  • Measure the visible control (the bordered box), not an inner <input> or icon.
  • Hidden, duplicate or offscreen element: the outline lands on nothing. Hidden at this preset (sidebar on phone): say so and show where it is instead. Never mark empty space.
  • Red means wrong; for emphasis use look.

相關技能