Communitygithub.com

Peytontalismanic424/paper-share-skills

Convert a PDF (especially academic papers) to clean Markdown using MinerU — preserves reading order, LaTeX equations, tables, and figures, and runs on the GPU. Use when the user asks to convert/extract a PDF to markdown/text. Invoke directly — no need for the user to type the skill name.

paper-share-skills とは?

paper-share-skills is a Claude Code agent skill that convert a PDF (especially academic papers) to clean Markdown using MinerU — preserves reading order, LaTeX equations, tables, and figures, and runs on the GPU. Use when the user asks to convert/extract a PDF to markdown/text. Invoke directly — no need for the user to type the skill name.

対応~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/Peytontalismanic424/paper-share-skills/tree/HEAD/pdf-to-markdown

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

PDF → Markdown (MinerU)

Converts PDFs to clean, structure-preserving Markdown using MinerU (opendatalab/MinerU), the highest-fidelity open-source engine for academic papers. Runs the pipeline backend on the GPU (CUDA). Output keeps heading hierarchy, inline LaTeX math ($...$), figures, and tables (as HTML).

When to use

  • The user asks to "convert / extract / turn this PDF into markdown (or text)."
  • Best fit: academic papers — multi-column layouts, LaTeX math, tables, references.
  • Also handles images, .docx, .pptx, .xlsx (pass the file directly).

How to run

MinerU lives in a dedicated isolated env and is driven through the bundled convert.py wrapper, which calls MinerU's Python API in-process.

Do not use the mineru CLI here — in 3.4.0 it spins up a local API server that returns 502 Bad Gateway on Windows. The wrapper bypasses that.

"$MINERU_PYTHON" \
  "<SKILLS_DIR>/pdf-to-markdown/convert.py" \
  "<input.pdf>" "<output_dir>" --lang en

- Arg 1 — input file (pdf/image/docx/pptx/xlsx).
- Arg 2 — output directory.
- `--lang` — OCR language hint: `en` for English, `ch` for Chinese/mixed (default `ch`,
  which also handles English). Only affects OCR'd (scanned) pages.
- The script prints `DONE -> <path to .md>` on success.

### Output layout

<output_dir>//auto/.md # the markdown <output_dir>//auto/images/ # extracted figures (referenced from the md) <output_dir>//auto/_*.json # layout/content metadata (can be ignored)


## After converting

1. Read the printed `DONE -> ...md` path and report it to the user.
2. Optionally show a short preview (title + first headings).
3. Note where extracted images were written.
4. Quality is normally high: headings as `#`, inline math as `$...$`, tables as
   `<table>` HTML (with LaTeX inside cells). If a scanned PDF comes out garbled,
   re-run with `--method ocr`.

## Preflight / setup (only if something is missing)

- Interpreter with MinerU (env `MINERU_PYTHON`, default `python`).
- Verify CUDA: `$MINERU_PYTHON -c "import torch; print(torch.cuda.is_available())"` → `True`.
- Models are pre-downloaded to `~/.cache/huggingface/hub/` (one-time).

If the env is missing, recreate it (install `uv` first if needed):

uv venv --python 3.12 "" uv pip install --python "/Scripts/python.exe" torch torchvision --index-url https://download.pytorch.org/whl/cu124 uv pip install --python "/Scripts/python.exe" -U "mineru[core]" "/Scripts/mineru-models-download.exe" -s huggingface -m pipeline

Then set `MINERU_PYTHON` to `<your-env>/Scripts/python.exe` (Windows) or
`<your-env>/bin/python` (macOS/Linux). On macOS/Linux the models-download binary is
`mineru-models-download` (no `.exe` suffix).

## Troubleshooting
- **`cuda.is_available()` is False** → reinstall torch from the `cu124` index (see setup).
  The pipeline still runs on CPU if needed, just slower.
- **Model download fails / slow** → re-run `mineru-models-download` with `-s modelscope`.
- **Scanned PDF garbled** → add `--method ocr` to the `convert.py` call.
- **`vlm-engine` / `hybrid-engine` backends** → these need vLLM/SGLang (Linux); they 502
  on Windows. Stick with the default `pipeline` backend used by the wrapper.

## Lightweight fallback (simple / Office docs)

For non-paper files where layout fidelity doesn't matter, Microsoft's **MarkItDown**
is faster and simpler — but it mangles academic two-column PDFs and tables, so prefer
MinerU for papers:

"$MINERU_PYTHON" -m pip install markitdown "$MINERU_PYTHON" -m markitdown "" > out.md

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

Peytontalismanic424/paper-share-skills

Validate and upload narrated paper videos to Bilibili with biliup. Handles required landscape-first/portrait-second submissions, series uploads, CST scheduling, diagnostics, dry-run validation, and durable upload receipts. Trigger on: "upload to bilibili", "bilibili upload", "发布到B站", "上传到B站".

Peytontalismanic424/paper-share-skills

Download the TeX source (e-print tar.gz) of an arXiv paper from an arXiv URL or bare ID, and unpack it into paper_src/. Use FIRST whenever a pipeline input is an arXiv link and the TeX source is preferred over the PDF (paper-to-beamer, paper-to-bilibili, paper-venue-discovery). Invoke with an arXiv URL (abs/pdf/e-print, arxiv.org or export.arxiv.org) or a bare ID like 2509.07996v4.

Peytontalismanic424/paper-share-skills

Convert Beamer PDF slides into a fail-closed narrated MP4 with page-aligned narration, TTS audio, cover, and upload metadata. Trigger on narrated slides, slide video, or paper video requests.

Peytontalismanic424/paper-share-skills

Automated pipeline: paper (PDF or TeX source) → SUSTech Beamer slides (11-section 论文分享 structure) → compiled PDF. When TeX source is available, skip MinerU and extract content directly from LaTeX. Use when the user asks to create Beamer slides from an academic paper, or says "论文分享", "paper to slides", "make beamer from paper". Invoke with an absolute or relative path to a PDF, .tar.gz TeX archive, .tex file, or an arXiv URL/ID (source TeX is downloaded first via skill://paper-download-arxiv-paper-source).

Peytontalismanic424/paper-share-skills

Orchestrate a paper PDF through MinerU, SUSTech Beamer slides, complete Chinese narration, landscape and portrait video, metadata gates, and Bilibili upload. Supports single-paper and batch slides, video, or bilibili phases. Trigger on full paper-to-Bilibili requests, PDF-to-video publication, and batch paper publishing.

Peytontalismanic424/paper-share-skills

Generate a deterministic 16:10 video-cover poster from a paper directory. Extracts paper metadata and figures and atomically produces poster.tex, poster.pdf, and poster.png. Use for paper-video covers, not conference posters.

Peytontalismanic424/paper-share-skills

Fix for undefined \setsource/\setdomains/\setpresenter/\setvenue commands when compiling SUSTech Beamer slides from the standard template. Copy the extended beamerthemesustech.sty from the sustech-slides-template repository.

関連スキル