Communitygithub.com

KaiyiHu/ResearchFigureSkill

Evidence-locked scientific visual compiler for auditable research figures

What is ResearchFigureSkill?

ResearchFigureSkill is a Claude Code agent skill that evidence-locked scientific visual compiler for auditable research figures.

Works with~Claude Code~Codex CLI~Cursor
npx skills add KaiyiHu/ResearchFigureSkill

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Research Figure Compiler

Public workflow version: 1.0.

Use the shortest safe workflow:

allowed paper content
  → resolve whether existing paper figures may be inspected
  → one useful paper summary
  → motivation and/or pipeline prompt template
  → editable figure + preview
  → one fast critical check

Do not create an evidence ledger, role-analysis file, FigureSpec, provenance bundle, audit JSON, or multi-round review package unless the user explicitly asks for one.

Resolve SKILL_ROOT as the directory containing this file.

1. Establish the user brief

Target the communication quality expected of leading AI conference papers: clear five-second contribution, source-bounded claims, reviewer-readable text, precise arrows, compact information density, disciplined alignment, accessible color, and an editable master. Examples include NeurIPS, ICML, ICLR, AAAI, CVPR, ACL, and KDD. This is a quality target, not an automatic claim of venue compliance. If a specific venue's current technical requirements matter, verify them from official sources.

Ask the user to provide these three decisions before analysis:

Figure type: Motivation, Pipeline, or both?
Existing corresponding figure in the paper: yes, no, or unknown?
If yes: completely replace it, use it as a reference and improve it, or
preserve it and repair selected parts?

If the request already answers a decision, do not ask it again. If one or more decisions are missing, ask one concise combined question before generating. The user may answer unknown for existing figures; resolve that case in step 3 without visually inspecting the figure first.

2. Select outputs without role classification

This Skill has two fixed figure types:

  • motivation: why the problem matters and why the research is needed;
  • pipeline: how the proposed method transforms input into output.

Follow the user's command:

  • If the user requests motivation, Figure 1, problem/gap, or 动机图, generate only motivation.
  • If the user requests pipeline, method, workflow, architecture, 方法图, or 流程图, generate only pipeline.
  • If the user requests both or a complete figure set, generate both.
  • If the user does not specify a type, ask whether they want Motivation, Pipeline, or both before continuing.

Do not infer one “winning role” and suppress the other.

3. Resolve existing paper figures before inspecting them

Respect user exclusions first. Never use excluded pages, figures, captions, or supplements even to decide whether they should become references.

For the remaining allowed source, detect the likely presence of an existing motivation or pipeline figure without visually inspecting it first. Use non-visual signals when practical, such as PDF image-object counts, section placement, or a visible Figure marker in already allowed extracted text.

If the paper appears to contain an existing figure of a selected type and the user has not already given an instruction, pause before opening that figure or reading its caption and ask one concise question:

The paper appears to contain an existing Motivation/Pipeline figure. Should I
(1) regenerate independently and ignore it, (2) use it only as a visual/layout
reference, or (3) preserve it and repair it?
  • For independent regeneration, do not inspect or use the old figure or its caption. Build a new composition from the allowed paper text and this Skill.
  • For reference-led regeneration, inspect it only after permission and treat it as visual structure/style, never as scientific evidence.
  • For repair, preserve verified elements and change only requested or failed parts.

If the user already says to discard, ignore, reference, or repair an existing figure, follow that instruction without asking again. Do not ask about incidental plots or figures unrelated to the selected output type.

4. Read and summarize the allowed source

Read the full allowed paper or brief once. Respect every user exclusion before text extraction or visual inspection. Never use excluded captions, figures, pages, or supplements as evidence.

Create one paper-summary.md, normally 500–900 words or equivalent, containing:

  1. research problem and importance;
  2. current approach and concrete gap;
  3. bounded thesis and contributions;
  4. method input, 3–7 main stages, handoffs, and output;
  5. strongest exact results that help understand the paper;
  6. limitations and interpretations the figure must not imply;
  7. exact terminology, numbers, and labels likely to appear in a figure;
  8. inspected scope, exclusions, and missing material.

Use page/section anchors where practical, but do not create a separate evidence ledger. Preserve units, signs, qualifiers, and uncertainty. Do not invent missing modules, values, relations, or causal claims.

5. Fill the fixed prompt template

Read references/prompt-templates.md every time. Fill only the template or templates selected in step 2.

For each prompt:

  • replace every {{PLACEHOLDER}};
  • copy exact scientific labels from the summary;
  • keep one clear five-second message;
  • list every visible entity and every arrow as source → target | label/payload;
  • state what must not be shown;
  • request live editable text and vector/native shapes;
  • include the short negative prompt and critical QA checklist.

Save motivation-prompt.md, pipeline-prompt.md, or both. Do not paste the entire paper into a drawing prompt.

Every filled prompt must include a layout safety contract:

  • reserve a clean outer safe area of 4–6% on every canvas edge;
  • keep titles, labels, arrowheads, borders, and meaningful icons fully inside that safe area; only nonessential background texture may enter it;
  • distribute visual mass across the useful canvas, normally occupying 75–88% without a large accidental blank region;
  • align related panel edges, title insets, stage centers, row baselines, and repeated objects to explicit shared guides;
  • keep gaps consistent and make whitespace intentional;
  • avoid both edge crowding and a composition with most content compressed into one corner.

Reference-led style contract

When the user supplies one or more desired examples, treat each one as a primary style and composition reference, not as optional inspiration and not as scientific evidence.

Before rendering, translate the examples into observable instructions for:

  • panel topology and approximate region ratios;
  • border style, stroke character, and corner treatment;
  • title lettering, body lettering, and typography hierarchy;
  • icon family, arrow rhythm, information density, and whitespace;
  • semantic accent colors and where fills are or are not used.

The filled prompt must say explicitly that the renderer must preserve this visual language and must not redesign it into a generic house style. Replace all reference-specific scientific content, text, numbers, logos, and unique icons with the validated inventory from the current paper.

If no visual reference is supplied, default to a hand-drawn academic infographic: white paper background, slightly irregular black linework, colored dashed rounded panel borders, handwritten-looking headings, simple scientific doodle icons, sparse pale highlights, and compact but readable information density. Do not default to a modern corporate card grid.

Motivation boundary

Use:

status quo → observed limitation or blind spot → bounded research need

Show at most three primary messages. Do not reveal the full architecture, training procedure, or result leaderboard.

Pipeline boundary

Use:

typed input → 3–7 verb-led stages → typed output

Name every handoff. Show branches or feedback only when the source explicitly supports them. Do not add benchmark victory badges or decorative modules.

6. Generate the figure

Use the filled prompt as the renderer input. Do not silently reinterpret it into a different visual system.

6.1 Choose the rendering order

Use image-first rendering when any of these is true:

  • the user supplied a visual reference;
  • the user asks for GPT/image-model generation;
  • matching a hand-drawn or illustrative visual language is more important than perfect vector purity.

In image-first mode, pass the filled prompt and the supplied reference image(s) directly to the image generator. The first artifact is the style-faithful PNG. Do not replace this step with a manually designed corporate SVG.

Use vector-first rendering only when the user explicitly prioritizes a fully editable deterministic master, or when exact plots/equations dominate. Vector-first output must still implement the declared style contract; editable does not mean clean sans-serif cards.

When both style fidelity and editability are requested:

  1. generate and approve the style-faithful image first;
  2. create an editable SVG companion that preserves the same panel topology, dashed borders, hand-drawn line character, lettering hierarchy, icons, and accent palette;
  3. keep exact scientific labels, values, and arrows live and correct;
  4. disclose if any illustrative layer remains raster rather than claiming the whole figure is fully editable.

Default deliverables remain:

motivation.svg + motivation.png
pipeline.svg + pipeline.png

Use another editable format only when the user asks. Exact quantitative plots, axes, equations, and sensitive numbers remain deterministic. Short diagram labels may be generated in the image-first style pass, but must be checked character by character and corrected in the editable companion when needed.

Before rendering, convert the layout safety contract into concrete coordinates or proportions. For a provisional 1800 × 1000 canvas, keep critical content about 72–108 px from the left/right edges and 40–60 px from the top/bottom. Align repeated regions and objects to named guides. Do not use apparent full-bleed borders that risk clipping after export.

When a reference image is supplied, preserve its observable visual grammar closely enough that the result belongs to the requested style family. Do not copy its scientific content, text, values, logos, or source-specific symbols.

7. Run one fast critical check

Inspect the actual preview at 100% and one 200% view. Check only these gates:

  1. Science — no invented or stronger-than-source claim, value, component, or legal/clinical conclusion.
  2. Structure — required entities are present; stage order and every arrow endpoint/direction are correct.
  3. Text — exact spelling, symbols, and numbers; no pseudo-text, missing glyphs, or unreadable labels.
  4. Optics and layout — no blur, fuzzy/melted shapes, overlap, clipping, off-canvas content, or obvious low-resolution enlargement. Confirm a visible 4–6% perimeter safe area, balanced use of the canvas, no large accidental blank region, consistent gaps, and alignment of related titles, panels, rows, stages, and repeated objects.
  5. Editability — the master retains live text and separate vector/native objects rather than one flattened bitmap.

When a reference was supplied, the optics gate also fails if panel topology, border treatment, lettering character, icon language, or information density has drifted into a visibly different style family.

For SVG, optionally run the lightweight check:

python3 "${SKILL_ROOT}/scripts/quick_qa.py" figure.svg figure.png

For an honestly disclosed image-first hybrid whose illustration layer remains raster, use --allow-hybrid. Never use this flag to describe a flattened image as fully editable.

If any gate fails, make one targeted repair and inspect again. Stop at the first passing result. Use at most two render attempts unless the user asks for more iteration.

8. Keep delivery small

Default files:

paper-summary.md
motivation-prompt.md        # only when selected
motivation.svg
motivation.png
pipeline-prompt.md          # only when selected
pipeline.svg
pipeline.png

Report the five QA gates in the final response; do not create a QA file. Temporary renders belong in an OS temporary directory and should be removed after the final files pass.

Boundaries

  • Do not expose private or unpublished material to an external provider without authorization.
  • Do not use image-generated axes, benchmark values, equations, or final high-risk scientific labels without deterministic verification.
  • Do not claim venue compliance unless current official requirements were checked.
  • Do not imitate a living artist or closely copy a reference figure.
  • Do not replace expert scientific, statistical, clinical, or legal review.

Related Skills