Communitygithub.com

baronguyen001/ai-automation-skills

Give scheduled scripts memory between runs with one SQLite file - a seen-set for dedup, a key/value cursor to resume where you left off, and order-preserving new-item filtering. Use for dedup across runs, don't re-alert the same item, remember the last id, resume a scraper, or persist state between cron runs.

Qu'est-ce que ai-automation-skills ?

ai-automation-skills is a Cursor agent skill that give scheduled scripts memory between runs with one SQLite file - a seen-set for dedup, a key/value cursor to resume where you left off, and order-preserving new-item filtering. Use for dedup across runs, don't re-alert the same item, remember the last id, resume a scraper, or persist state between cron runs.

Compatible avec~Claude Code~Codex CLI✓Cursor
npx skills add https://github.com/baronguyen001/ai-automation-skills/tree/HEAD/skills/sqlite-state

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

SQLite State

Use this skill when a recurring job must remember what it already did - which items it has alerted on, the last cursor it processed - so it does not re-notify or re-scrape on the next tick. One stdlib SQLite file gives durable, file-locked state with no server and no extra dependency.

When to invoke

  • User says: "don't alert the same thing twice", "dedup across runs", "remember the last id", "resume where it left off".
  • Code in the conversation is a cron/scheduled script that currently re-processes everything each run.

When NOT to invoke

  • State is tiny and ephemeral within a single run - a plain set/dict is enough.
  • Multiple machines must share state concurrently; reach for a real server DB instead of a single file.

Concrete example

User input:

My news scraper re-sends the same headlines every hour. Make it only send new ones.

Output:

from state import filter_new, mark_seen

db = "news.db"
fresh = filter_new(db, [a["url"] for a in articles])   # only unseen URLs
for url in fresh:
    send_alert(url)
mark_seen(db, *fresh)                                   # remember them for next run

Pattern to apply

  1. Key each item by something stable and unique (URL, id, content hash), not by array position.
  2. Use filter_new to decide what to act on, then mark_seen only after the action succeeds.
  3. Store progress as a named cursor (set_cursor/get_cursor) so a resumable scraper restarts mid-stream.
  4. Enable WAL mode for durable writes with concurrent readers; keep a single writer.
  5. Back up or version the one .db file - it is the whole memory of the job.

Reference: assets/state.py.

Source

Distilled from production use across the author's automation projects. v1.0.0. See also: [[pipeline-orchestrator]], [[cron-dispatch]], [[webhook-receiver]].

→ Build the full runnable bot with Trawlkit.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

baronguyen001/ai-automation-skills

Schedule any script to run on a recurring schedule on Windows (Task Scheduler) or Linux (cron) - register, list, and remove jobs from one command, with logging to a file and a guard against overlapping runs. Use for schedule a script, run nightly, set up a cron job, windows task scheduler, or run on a timer.

baronguyen001/ai-automation-skills

Turn a run's list of result dicts into a schema'd CSV and a Markdown table from one column spec - declare columns once, emit both, with stable ordering and safe escaping, stdlib only, no pandas. Use when the user asks to write results to CSV, export a report, make a markdown summary table, or save a run's output as a spreadsheet.

baronguyen001/ai-automation-skills

Strip HTML to clean plain text with the standard library only (no BeautifulSoup/lxml): drops script/style, turns block tags into line breaks, and collapses whitespace. Use when the user wants readable text from an HTML page/email, to clean scraped HTML before sending it to an LLM, or to build a text index from web content.

baronguyen001/ai-automation-skills

Persist records between scheduled runs as append-only JSON Lines (one object per line) with streaming reads and optional key-based dedup, stdlib only. Use when the user wants to log run results to JSONL, append events to a file, dedup records by id across runs, or keep a simple durable history without a database.

baronguyen001/ai-automation-skills

Extract text and simple table-like rows from a PDF for downstream AI without OCR binaries. Use when the user asks to read a PDF, turn a PDF into text, pull simple tables from statements/reports, or feed PDF content into an LLM pipeline.

baronguyen001/ai-automation-skills

Rotate a pool of HTTP/SOCKS proxies with round-robin selection, failure tracking, and a cooldown that benches dead proxies before retrying. BYO proxy list via env - none are shipped. Use for rotate proxies, spread requests across proxies, avoid IP bans, retry through a different proxy, or proxy health check.

baronguyen001/ai-automation-skills

Parse an RSS 2.0 or Atom feed into normalized item dicts (title, link, id, published, summary) with the standard library only - no feedparser. Use when the user asks to read an RSS/Atom feed, poll a blog/news feed, extract feed entries, or watch a site that publishes a feed.

baronguyen001/ai-automation-skills

Upload and download run artifacts from S3-compatible storage with BYO bucket and credentials from env. Use when the user asks to store scraper output, archive reports, publish artifacts to object storage, fetch a prior run file, or use MinIO/R2/Spaces/S3 without hardcoding credentials.

baronguyen001/ai-automation-skills

Draft a short-form TikTok/Reels/Shorts or long-form video script from a topic and creator persona file - hook variants, beat-structured body, CTA, b-roll cues, caption overlays, and [STORY]/[NUMBER] slots so the model never fabricates personal facts. Use for write a video script, tiktok script, youtube script, or /shortform-script.

Skills associés