Communitygithub.com

baronguyen001/ai-automation-skills

Strip HTML to clean plain text with the standard library only (no BeautifulSoup/lxml): drops script/style, turns block tags into line breaks, and collapses whitespace. Use when the user wants readable text from an HTML page/email, to clean scraped HTML before sending it to an LLM, or to build a text index from web content.

ai-automation-skills란 무엇인가요?

ai-automation-skills is a Gemini CLI agent skill that strip HTML to clean plain text with the standard library only (no BeautifulSoup/lxml): drops script/style, turns block tags into line breaks, and collapses whitespace. Use when the user wants readable text from an HTML page/email, to clean scraped HTML before sending it to an LLM, or to build a text index from web content.

지원 대상~Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/baronguyen001/ai-automation-skills/tree/HEAD/skills/html-to-text

즐겨 사용하는 AI에게 물어보기

이 에이전트 스킬이 미리 로드된 새 채팅을 엽니다.

문서

HTML to Text

Use this skill when you have HTML - a scraped page, a newsletter body, a feed entry's content - and you want clean, readable plain text without adding beautifulsoup4 or lxml. A small HTMLParser subclass removes <script>/<style>, converts block elements into newlines, and collapses whitespace, which is exactly the shape you want before sending content to an LLM (fewer tokens) or into a search index.

When to invoke

  • User says: "get the text out of this HTML", "clean this scraped page", "strip tags before summarizing", "convert this email HTML to text".
  • HTML is about to be fed to a model or a digest and the markup is just noise.

When NOT to invoke

  • You need to keep structure (tables, links as Markdown) - use a real HTML-to-Markdown converter.
  • The page is rendered by JavaScript - fetch it with a headless browser first, then convert.

Concrete example

User input:

Turn this article HTML into plain text so I can summarize it with fewer tokens.

Output:

# Copy assets/htmltext.py into your project, then:
from htmltext import html_to_text

text = html_to_text(article_html)
summary = call_llm(f"Summarize:\n{text}")

<script> and <style> blocks are dropped entirely, block tags become line breaks, and runs of spaces/blank lines are collapsed, so the output reads like prose.

Pattern to apply

  1. Subclass html.parser.HTMLParser with convert_charrefs=True so entities decode automatically.
  2. Track a skip-depth for script/style/head so their text never leaks into the output.
  3. Emit newlines around block tags, then collapse intra-line whitespace and multi-blank runs.
  4. Keep it pure (string in, string out) so it is trivial to unit test.

Reference: assets/htmltext.py.

Source

Distilled from the author's scraping and digest pipelines. v1.0.0. See also: [[rss-feed-reader]], [[gemini-cost-tracker]], [[pipeline-orchestrator]].

→ Build the full runnable bot with Trawlkit.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

baronguyen001/ai-automation-skills

Schedule any script to run on a recurring schedule on Windows (Task Scheduler) or Linux (cron) - register, list, and remove jobs from one command, with logging to a file and a guard against overlapping runs. Use for schedule a script, run nightly, set up a cron job, windows task scheduler, or run on a timer.

baronguyen001/ai-automation-skills

Turn a run's list of result dicts into a schema'd CSV and a Markdown table from one column spec - declare columns once, emit both, with stable ordering and safe escaping, stdlib only, no pandas. Use when the user asks to write results to CSV, export a report, make a markdown summary table, or save a run's output as a spreadsheet.

baronguyen001/ai-automation-skills

Persist records between scheduled runs as append-only JSON Lines (one object per line) with streaming reads and optional key-based dedup, stdlib only. Use when the user wants to log run results to JSONL, append events to a file, dedup records by id across runs, or keep a simple durable history without a database.

baronguyen001/ai-automation-skills

Extract text and simple table-like rows from a PDF for downstream AI without OCR binaries. Use when the user asks to read a PDF, turn a PDF into text, pull simple tables from statements/reports, or feed PDF content into an LLM pipeline.

baronguyen001/ai-automation-skills

Rotate a pool of HTTP/SOCKS proxies with round-robin selection, failure tracking, and a cooldown that benches dead proxies before retrying. BYO proxy list via env - none are shipped. Use for rotate proxies, spread requests across proxies, avoid IP bans, retry through a different proxy, or proxy health check.

baronguyen001/ai-automation-skills

Parse an RSS 2.0 or Atom feed into normalized item dicts (title, link, id, published, summary) with the standard library only - no feedparser. Use when the user asks to read an RSS/Atom feed, poll a blog/news feed, extract feed entries, or watch a site that publishes a feed.

baronguyen001/ai-automation-skills

Upload and download run artifacts from S3-compatible storage with BYO bucket and credentials from env. Use when the user asks to store scraper output, archive reports, publish artifacts to object storage, fetch a prior run file, or use MinIO/R2/Spaces/S3 without hardcoding credentials.

baronguyen001/ai-automation-skills

Draft a short-form TikTok/Reels/Shorts or long-form video script from a topic and creator persona file - hook variants, beat-structured body, CTA, b-roll cues, caption overlays, and [STORY]/[NUMBER] slots so the model never fabricates personal facts. Use for write a video script, tiktok script, youtube script, or /shortform-script.

baronguyen001/ai-automation-skills

Give scheduled scripts memory between runs with one SQLite file - a seen-set for dedup, a key/value cursor to resume where you left off, and order-preserving new-item filtering. Use for dedup across runs, don't re-alert the same item, remember the last id, resume a scraper, or persist state between cron runs.

관련 스킬