Communitygithub.com

baronguyen001/ai-automation-skills

Parse an RSS 2.0 or Atom feed into normalized item dicts (title, link, id, published, summary) with the standard library only - no feedparser. Use when the user asks to read an RSS/Atom feed, poll a blog/news feed, extract feed entries, or watch a site that publishes a feed.

Was ist ai-automation-skills?

ai-automation-skills is a Claude Code agent skill that parse an RSS 2.0 or Atom feed into normalized item dicts (title, link, id, published, summary) with the standard library only - no feedparser. Use when the user asks to read an RSS/Atom feed, poll a blog/news feed, extract feed entries, or watch a site that publishes a feed.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/baronguyen001/ai-automation-skills/tree/HEAD/skills/rss-feed-reader

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

RSS / Atom Feed Reader

Use this skill when a script needs to read a feed and you want one normalized shape across both RSS 2.0 and Atom, with no third-party dependency. It does a single ElementTree pass and pulls title, link, id, published, and summary from either format, so downstream code (dedup, digest, alert) never has to branch on feed type.

When to invoke

  • User says: "read this RSS feed", "poll a blog feed", "get the latest entries from an Atom feed", "watch a site for new posts".
  • Code already fetched feed bytes and needs structured entries.

When NOT to invoke

  • The source has a real JSON API - call that instead of scraping a feed.
  • You need full feedparser quirks handling (every malformed dialect) - install feedparser.

Concrete example

User input:

Read https://example.com/feed.xml and print each post title and link.

Output:

# Copy assets/feed.py into your project, then:
from urllib.request import urlopen
from feed import parse_feed

with urlopen("https://example.com/feed.xml", timeout=20) as r:
    for item in parse_feed(r.read()):
        print(item["title"], "->", item["link"])

Every item is a dict with the same keys whether the feed is RSS or Atom, so dedup by id and sort by published work without special-casing.

Pattern to apply

  1. Parse once with xml.etree.ElementTree; iterate item (RSS) and {atom}entry (Atom).
  2. Normalize to a stable key set so the rest of the pipeline is feed-format agnostic.
  3. Fall back from guid/atom id to the link so every item has a stable identity for dedup.
  4. Keep network out of the parser: pass in bytes/text so it stays pure and testable.

Reference: assets/feed.py.

Source

Distilled from the author's news-aggregation projects. v1.0.0. See also: [[gmail-imap-digest]], [[sqlite-state]], [[pipeline-orchestrator]].

→ Build the full runnable bot with Trawlkit.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

baronguyen001/ai-automation-skills

Schedule any script to run on a recurring schedule on Windows (Task Scheduler) or Linux (cron) - register, list, and remove jobs from one command, with logging to a file and a guard against overlapping runs. Use for schedule a script, run nightly, set up a cron job, windows task scheduler, or run on a timer.

baronguyen001/ai-automation-skills

Turn a run's list of result dicts into a schema'd CSV and a Markdown table from one column spec - declare columns once, emit both, with stable ordering and safe escaping, stdlib only, no pandas. Use when the user asks to write results to CSV, export a report, make a markdown summary table, or save a run's output as a spreadsheet.

baronguyen001/ai-automation-skills

Strip HTML to clean plain text with the standard library only (no BeautifulSoup/lxml): drops script/style, turns block tags into line breaks, and collapses whitespace. Use when the user wants readable text from an HTML page/email, to clean scraped HTML before sending it to an LLM, or to build a text index from web content.

baronguyen001/ai-automation-skills

Persist records between scheduled runs as append-only JSON Lines (one object per line) with streaming reads and optional key-based dedup, stdlib only. Use when the user wants to log run results to JSONL, append events to a file, dedup records by id across runs, or keep a simple durable history without a database.

baronguyen001/ai-automation-skills

Extract text and simple table-like rows from a PDF for downstream AI without OCR binaries. Use when the user asks to read a PDF, turn a PDF into text, pull simple tables from statements/reports, or feed PDF content into an LLM pipeline.

baronguyen001/ai-automation-skills

Rotate a pool of HTTP/SOCKS proxies with round-robin selection, failure tracking, and a cooldown that benches dead proxies before retrying. BYO proxy list via env - none are shipped. Use for rotate proxies, spread requests across proxies, avoid IP bans, retry through a different proxy, or proxy health check.

baronguyen001/ai-automation-skills

Upload and download run artifacts from S3-compatible storage with BYO bucket and credentials from env. Use when the user asks to store scraper output, archive reports, publish artifacts to object storage, fetch a prior run file, or use MinIO/R2/Spaces/S3 without hardcoding credentials.

baronguyen001/ai-automation-skills

Draft a short-form TikTok/Reels/Shorts or long-form video script from a topic and creator persona file - hook variants, beat-structured body, CTA, b-roll cues, caption overlays, and [STORY]/[NUMBER] slots so the model never fabricates personal facts. Use for write a video script, tiktok script, youtube script, or /shortform-script.

baronguyen001/ai-automation-skills

Give scheduled scripts memory between runs with one SQLite file - a seen-set for dedup, a key/value cursor to resume where you left off, and order-preserving new-item filtering. Use for dedup across runs, don't re-alert the same item, remember the last id, resume a scraper, or persist state between cron runs.

Verwandte Skills