Communitygithub.com

baronguyen001/ai-automation-skills

Extract text and simple table-like rows from a PDF for downstream AI without OCR binaries. Use when the user asks to read a PDF, turn a PDF into text, pull simple tables from statements/reports, or feed PDF content into an LLM pipeline.

ai-automation-skills란 무엇인가요?

ai-automation-skills is a Gemini CLI agent skill that extract text and simple table-like rows from a PDF for downstream AI without OCR binaries. Use when the user asks to read a PDF, turn a PDF into text, pull simple tables from statements/reports, or feed PDF content into an LLM pipeline.

지원 대상~Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/baronguyen001/ai-automation-skills/tree/HEAD/skills/pdf-text-extract

즐겨 사용하는 AI에게 물어보기

이 에이전트 스킬이 미리 로드된 새 채팅을 엽니다.

문서

PDF Text Extract

Use this skill when a workflow needs machine-readable text from a digital PDF before summarization, classification, or structured extraction. The helper uses pure-Python PDF libraries when available and includes a simple line-based table detector for downstream cleanup.

When to invoke

  • User says: "extract text from this PDF", "feed a PDF to AI", "pull tables from this report", "read this statement".
  • The PDF already contains selectable text and does not require OCR.

When NOT to invoke

  • The PDF is a scanned image; use an OCR workflow instead.
  • The user needs pixel-perfect table reconstruction with merged cells and layout fidelity.

Concrete example

User input:

Extract the text and rough tables from this PDF so Gemini can summarize it.

Output:

# Copy assets/extract.py into your project, then:
from extract import extract_pdf

doc = extract_pdf("downloads/report.pdf")
print(doc["text"][:2000])
for table in doc["tables"]:
    print(table["page"], table["rows"][:3])

Install either pypdf or pdfminer.six in the target project. No OCR binary is required, and the helper fails clearly when the PDF has no extractable text.

Pattern to apply

  1. Prefer digital text extraction first; do not add OCR unless the source is scanned.
  2. Preserve page boundaries so downstream prompts can cite page numbers.
  3. Keep table extraction simple: split rows on tabs or repeated spaces, then let a later schema pass normalize columns.
  4. Cap text sent to an LLM by page/range when the PDF is long.
  5. Fail loudly when no text is found instead of sending an empty prompt downstream.

Reference: assets/extract.py.

Source

Distilled from production use across the author's automation projects. v1.0.0. See also: [[gemini-structured-output]], [[csv-report-writer]], [[s3-uploader]].

→ Build the full runnable bot with Trawlkit.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

baronguyen001/ai-automation-skills

Schedule any script to run on a recurring schedule on Windows (Task Scheduler) or Linux (cron) - register, list, and remove jobs from one command, with logging to a file and a guard against overlapping runs. Use for schedule a script, run nightly, set up a cron job, windows task scheduler, or run on a timer.

baronguyen001/ai-automation-skills

Turn a run's list of result dicts into a schema'd CSV and a Markdown table from one column spec - declare columns once, emit both, with stable ordering and safe escaping, stdlib only, no pandas. Use when the user asks to write results to CSV, export a report, make a markdown summary table, or save a run's output as a spreadsheet.

baronguyen001/ai-automation-skills

Strip HTML to clean plain text with the standard library only (no BeautifulSoup/lxml): drops script/style, turns block tags into line breaks, and collapses whitespace. Use when the user wants readable text from an HTML page/email, to clean scraped HTML before sending it to an LLM, or to build a text index from web content.

baronguyen001/ai-automation-skills

Persist records between scheduled runs as append-only JSON Lines (one object per line) with streaming reads and optional key-based dedup, stdlib only. Use when the user wants to log run results to JSONL, append events to a file, dedup records by id across runs, or keep a simple durable history without a database.

baronguyen001/ai-automation-skills

Rotate a pool of HTTP/SOCKS proxies with round-robin selection, failure tracking, and a cooldown that benches dead proxies before retrying. BYO proxy list via env - none are shipped. Use for rotate proxies, spread requests across proxies, avoid IP bans, retry through a different proxy, or proxy health check.

baronguyen001/ai-automation-skills

Parse an RSS 2.0 or Atom feed into normalized item dicts (title, link, id, published, summary) with the standard library only - no feedparser. Use when the user asks to read an RSS/Atom feed, poll a blog/news feed, extract feed entries, or watch a site that publishes a feed.

baronguyen001/ai-automation-skills

Upload and download run artifacts from S3-compatible storage with BYO bucket and credentials from env. Use when the user asks to store scraper output, archive reports, publish artifacts to object storage, fetch a prior run file, or use MinIO/R2/Spaces/S3 without hardcoding credentials.

baronguyen001/ai-automation-skills

Draft a short-form TikTok/Reels/Shorts or long-form video script from a topic and creator persona file - hook variants, beat-structured body, CTA, b-roll cues, caption overlays, and [STORY]/[NUMBER] slots so the model never fabricates personal facts. Use for write a video script, tiktok script, youtube script, or /shortform-script.

baronguyen001/ai-automation-skills

Give scheduled scripts memory between runs with one SQLite file - a seen-set for dedup, a key/value cursor to resume where you left off, and order-preserving new-item filtering. Use for dedup across runs, don't re-alert the same item, remember the last id, resume a scraper, or persist state between cron runs.

관련 스킬