Communitygithub.com

curiositech/windags-skills

Use when designing a background-job system, choosing between BullMQ / Sidekiq / RQ / Temporal / SQS, deciding queue-vs-workflow, sizing concurrency vs rate limits, building dead-letter queues, or making handlers idempotent. Triggers: jobs running twice on retry, lost jobs after worker crash, DLQ filling up, Redis OOM from job backlog, exactly-once requested, "do we need Temporal?", visibility timeout / lockDuration confusion, exponential backoff vs jitter, fan-out fan-in workflows. NOT for outbound webhook publishing (different concerns), receiver-side webhook handling (different concerns), event-streaming/Kafka topology, or in-process async (event loop only).

windags-skills란 무엇인가요?

windags-skills is a Claude Code agent skill that use when designing a background-job system, choosing between BullMQ / Sidekiq / RQ / Temporal / SQS, deciding queue-vs-workflow, sizing concurrency vs rate limits, building dead-letter queues, or making handlers idempotent. Triggers: jobs running twice on retry, lost jobs after worker crash, DLQ filling up, Redis OOM from job backlog, exactly-once requested, "do we need Temporal?", visibility timeout / lockDuration confusion, exponential backoff vs jitter, fan-out fan-in workflows. NOT for outbound webhook publishing (different concerns), receiver-side webhook handling (different concerns), event-streaming/Kafka topology, or in-process async (event loop only).

지원 대상~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/curiositech/windags-skills/tree/HEAD/skills/background-job-queue-design

즐겨 사용하는 AI에게 물어보기

이 에이전트 스킬이 미리 로드된 새 채팅을 엽니다.

문서

Background Job Queue Design

A background-job system has three failure modes you have to design against from day one: lost jobs, duplicate execution, and runaway retries. Pick the wrong primitive and you'll spend the next year fixing all three. The recurring industry consensus, across BullMQ, Sidekiq, SQS, and Temporal docs alike, is at-least-once is the only honest delivery semantic — design every handler to be idempotent and stop chasing exactly-once. (Caduh — Queues 101)

Jump to your fire:

When to use

  • Designing a new background-job system in Node, Ruby, Python, or Go.
  • Migrating from setTimeout / cron / homegrown queues to a real broker.
  • Worker crashes are losing jobs.
  • Production retries are causing duplicate side effects.
  • Choosing between BullMQ (Node + Redis), Sidekiq (Ruby + Redis), RQ / Celery / Dramatiq (Python), Asynq (Go + Redis), SQS (managed), or Temporal (durable execution).

Core capabilities

Pick the primitive that matches the work

You haveReach forReason
Short, side-effecting jobs (send email, resize image) on NodeBullMQRedis-native, mature DLQ, flow patterns, ~~5k jobs/sec/node typical
Same on RubySidekiq (Pro for reliability)Same shape; Pro super_fetch adds reliable fetch via LMOVE (Sidekiq wiki)
Multi-step workflows w/ external API fan-out, hours-to-days durationsTemporalDurable execution; replays from event history (Temporal docs)
AWS-native, simple fanout, don't want to run RedisSQS + LambdaManaged at-least-once; visibility timeout + DLQ first-class (AWS SQS DLQ)
Cross-region, multi-tenant isolationTemporal or SQSSingle Redis won't survive multi-region cleanly
Need exactly-once across DB + email + paymentsNone — design for at-least-once + idempotent handlers"Exactly-once" inside the broker doesn't extend across systems (Caduh)

The hidden criterion: how complex is the recovery story? A simple BullMQ queue with retries is fine for "send a receipt." A 12-step purchase workflow that pauses for 24 hours waiting for a webhook is a Temporal workflow, not 12 chained queue jobs.

Idempotency

The mandatory property. Every popular queue's docs lead with this. From the BullMQ idempotent-jobs page: "it should not make a difference to the final state of the system if a job successfully completes on its first attempt, or if it fails initially and succeeds when retried." (BullMQ idempotent-jobs)

Practical patterns:

// 1. Idempotency key derived from the business event, not the queue.
//    Use the upstream event_id if you have one (Stripe event.id, Shopify hook id).
const job = await queue.add(
  'send-receipt',
  { orderId, amount },
  {
    jobId: `send-receipt:${orderId}`,   // BullMQ dedups by jobId in the queue.
    attempts: 5,
    backoff: { type: 'exponential', delay: 1000 },
    removeOnComplete: { age: 3600, count: 1000 },
    removeOnFail:     { age: 86400 },
  }
);

// 2. Inside the handler, dedupe at the side-effect site too.
async function processSendReceipt(job) {
  const { orderId, amount } = job.data;
  // If we already sent, do nothing.
  const sent = await db.queryOne(
    'INSERT INTO email_sends (idempotency_key, order_id) VALUES ($1, $2) ON CONFLICT DO NOTHING RETURNING id',
    [`send-receipt:${orderId}`, orderId]
  );
  if (!sent) return { skipped: 'already-sent' };
  await emailProvider.send({ to: ..., subject: ..., body: render(amount) });
}

The jobId dedupes adds while a duplicate is still queued. The DB unique constraint dedupes across worker crashes, retries, and replays. Both are required; jobId alone is insufficient because BullMQ removes completed jobs and the dedup window evaporates.

Reliable fetch and visibility timeout

When a worker pulls a job and crashes, what happens?

SystemDefault behaviorRecovery
BullMQJob has a lockDuration (default 30s). Stalled-job checker re-queues after expiry (BullMQ production)Workers MUST extend the lock or finish within lockDuration; otherwise it runs twice
Sidekiq OSSPop is non-atomic; if worker dies between pop and process, job is lostUpgrade to Sidekiq Pro super_fetch (Redis 6.2+ LMOVE to a private working queue)
Sidekiq Pro super_fetchLMOVE to per-process working list. Heartbeat expires at 60s; orphan check sweeps and re-enqueues (Sidekiq wiki)Built-in
SQSVisibility timeout (default 30s, max 12h). Message reappears after timeout if not deletedSet timeout > p99 handler latency; extend with ChangeMessageVisibility for long jobs
TemporalActivities have heartbeat + retry policies. Workflow history survives worker deathReplay from history; you don't think about this

The trap: lockDuration / visibility timeout shorter than handler p99 → job runs twice. Longer than acceptable recovery time → crashed work waits too long. Measure your p99 first, set the timeout to ~3x that, and have the handler heartbeat / extend the lock for genuinely-long work.

Error classification and DLQ

Not every error should retry forever. From SQS docs: "Set the maximum receives... and the redrive policy to send messages to the DLQ once threshold is exceeded." (AWS SQS DLQ)

// Classify before throwing. Permanent errors should NOT retry.
class PermanentError extends Error { constructor(m: string) { super(m); this.name = 'PermanentError'; } }
class TransientError extends Error { constructor(m: string) { super(m); this.name = 'TransientError'; } }

async function process(job) {
  try {
    const user = await db.user(job.data.userId);
    if (!user) throw new PermanentError('user-not-found');   // → DLQ on first try
    if (user.suspended) throw new PermanentError('user-suspended');
    await externalApi.call(user);
  } catch (e) {
    if (e instanceof PermanentError) {
      // Move to DLQ immediately — retrying won't help.
      job.discard();    // BullMQ: prevents retry
      throw e;
    }
    throw e;            // Network/timeout → retry per attempts policy
  }
}

DLQ then needs:

  • An alert on rising count (not just nonzero — "DLQ has 3 things forever" is fine; "DLQ grew 100/min" is a fire).
  • A replay tool (CLI or admin UI) so a human can fix the upstream bug, replay the dead-lettered jobs, and clear the DLQ.
  • Dashboards (see grafana-dashboard-builder).

Backoff and jitter

Exponential backoff alone isn't enough. If 1000 jobs all fail at 12:00:00 because a downstream service blipped, plain exponential backoff has all 1000 retry at exactly 12:00:01, 12:00:03, 12:00:07 — a thundering herd that may keep the downstream service down. Add full jitter:

// AWS-style "full jitter": delay = random(0, exp_backoff)
const baseDelay = Math.min(2 ** attempt * 1000, 30_000);
const delay = Math.random() * baseDelay;

For the BullMQ recommended baseline, the going-to-production doc suggests retry intervals between 1s and 20s with retryStrategy. (BullMQ production)

Concurrency and limiter

Per-worker concurrency runs N jobs in parallel from one process. Multiple workers multiply that. Without a queue-level limiter, you can't cap aggregate calls to a downstream:

// BullMQ — global rate limit across all workers on this queue.
const queue = new Queue('outbound-emails', {
  connection,
  limiter: { max: 100, duration: 1000 }, // 100 emails/sec across the whole fleet
});

The right shape: many concurrent workers + a queue limiter for downstream contract limits. Sidekiq has equivalent rate limiters; SQS uses Lambda concurrency or per-message-rate.

Memory hygiene

The single most-emphasized BullMQ production gotcha: maxmemory-policy noeviction on the Redis instance. Anything else (allkeys-lru, volatile-lru) will silently evict queued jobs under pressure. (BullMQ production)

Also:

  • removeOnComplete aggressive (e.g. { age: 3600, count: 1000 }).
  • removeOnFail even more aggressive once you've moved to a DLQ pattern.
  • Enable Redis AOF persistence — appendfsync everysec is the documented sweet spot for BullMQ. (BullMQ production)
  • Worker maxRetriesPerRequest: null to prevent ioredis from throwing during transient disconnects. (BullMQ production)
  • Queue side: enableOfflineQueue: false so producer fails fast instead of buffering.

Graceful shutdown

process.on('SIGTERM', async () => {
  await worker.close();   // stop accepting; finish in-flight jobs.
  await connection.quit();
  process.exit(0);
});
process.on('SIGINT', async () => { /* same */ });

The BullMQ production guide explicitly notes: close workers before stopping the process; default stalling timeout is ~30s. (BullMQ production)

Queue vs workflow decision

Useful framing from Temporal's docs: "Unlike message queues which move data between services, Temporal orchestrates entire processes... knows where you are in a workflow, what's completed, what's pending, and what needs to retry." (Temporal blog)

flowchart TD
  A[Background work] --> B{Single side-effect or pipeline?}
  B -->|Single| C[BullMQ / Sidekiq / SQS]
  B -->|Multi-step| D{Steps span seconds or hours+ ?}
  D -->|Seconds| E[BullMQ flows / Sidekiq batches]
  D -->|Hours+ or human-in-loop| F[Temporal / DBOS / Restate]
  C --> G{Need cross-region or multi-tenant isolation?}
  G -->|Yes| H[SQS or Temporal — Redis-single-cluster won't]
  G -->|No| I[Run with Redis cluster + AOF + noeviction]

Anti-patterns

Idempotency via Redis SET-NX only

Symptom: During a Redis blip or restart, the same job runs twice and creates duplicate side effects (charge, email, ticket). Diagnosis: Redis is a cache; it can be evicted, restarted, or partitioned. SET-NX dedup that lives only in Redis is best-effort. Fix: Idempotency key with a DB unique constraint on the side-effect record (e.g. email_sends.idempotency_key). The DB transaction that creates the side-effect IS the dedup primitive.

lockDuration shorter than handler p99

Symptom: Stalled-job log entries; same job processed by two workers. Customers complain. Diagnosis: Default lockDuration is 30s; handler sometimes takes 45s; the second worker picks it up. Fix: Measure p99, set lockDuration to ~3x p99, and call job.extendLock() inside long handlers. Or split the work.

Plain exponential backoff (no jitter)

Symptom: A downstream blip becomes a downstream outage; 10k jobs hammer the recovering service in lockstep. Diagnosis: Without jitter, retries are correlated. Fix: Full-jitter backoff. Cap maximum delay to bound DLQ time-to-failure.

maxmemory-policy allkeys-lru on Redis

Symptom: Queue mysteriously loses jobs under load. No errors. (BullMQ production) Diagnosis: Redis evicted job keys to make room. Fix: CONFIG SET maxmemory-policy noeviction AND assert it at startup. Workers should refuse to start otherwise.

Pretending exactly-once across systems

Symptom: Charged customer twice; sent two of the same email. Engineer is sure "the broker is exactly-once." Diagnosis: Even a FIFO/exactly-once broker doesn't extend its guarantee to your DB + email vendor + Stripe + ledger. Fix: Treat the queue as at-least-once. Idempotency at every side-effect boundary.

DLQ with no replay tool

Symptom: "We have 3,200 jobs in the DLQ. We don't know what's in there." DLQ becomes an unmonitored graveyard. Diagnosis: Built the DLQ, didn't build the operator tool. Fix: Replay-by-id, replay-by-time-range, drain-with-confirmation. CLI or admin UI. Tested.

Long-blocking work on the event loop

Symptom: BullMQ worker process hangs; heartbeats stop; lock expires; job re-runs. Diagnosis: CPU-bound work in the same process as the queue heartbeat; event loop blocks for > lockDuration. Fix: Spawn a child process or worker thread for CPU-bound work. Or run a sandboxed processor (BullMQ supports per-job process sandboxing).

Quality gates

  • Test: chaos test — kill a worker mid-job; assert the job re-runs and the side-effect is not duplicated.
  • Test: retry-storm test — fail downstream for 60s; assert backoff + jitter spreads retries (no thundering herd in metrics).
  • Every handler is idempotent and uses a stable idempotency key derived from the business event.
  • DB-level unique constraint backs up any in-Redis dedup (jobId, SET-NX).
  • lockDuration / visibility timeout ≥ 3× measured p99 handler latency, AND long handlers extend the lock.
  • Permanent errors (PermanentError class or equivalent) bypass retry and go straight to DLQ.
  • DLQ alerts: page when growth rate > N per minute, not just count > 0.
  • DLQ replay tool exists and is tested (one-off + range).
  • Redis: maxmemory-policy=noeviction, AOF enabled (appendfsync everysec), asserted on worker startup. (BullMQ production)
  • removeOnComplete and removeOnFail configured. Job count metric trended in grafana-dashboard-builder.
  • Graceful shutdown on SIGTERM / SIGINT with worker.close() before exit.
  • OTel spans around handler with queue.name, job.id, job.attempts, job.outcome (see opentelemetry-instrumentation).
  • Concurrency × worker-count × downstream-rate-limit math reviewed; queue-level limiter in place where needed.

NOT for

  • Outbound webhook publishing — different domain (delivery guarantees, customer secrets). No dedicated skill yet.
  • Receiver-side webhook handling — different domain. → webhook-receiver-design.
  • Event-streaming / Kafka topology — different abstraction (log, not queue). No dedicated skill.
  • In-process async only (no broker, single replica) — different operational profile. → python-asyncio-pitfalls for Python event-loop concurrency.
  • Cron-style scheduled jobs without retry/DLQ semantics — simpler tooling fits.

Sources

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

curiositech/windags-skills

Expert in 2000s-era music visualization (Milkdrop, AVS, Geiss) and modern WebGL implementations. Specializes in Butterchurn integration, Web Audio API AnalyserNode FFT data, GLSL shaders for audio-reactive visuals, and psychedelic generative art. Activate on "Milkdrop", "music visualization", "WebGL visualizer", "Butterchurn", "audio reactive", "FFT visualization", "spectrum analyzer". NOT for simple bar charts/waveforms (use basic canvas), video editing, or non-audio visuals.

curiositech/windags-skills

Expert legal research agent for finding and scraping expungement data state by state. Knows authoritative sources, URL patterns, Firecrawl configuration, and 2026 legal landscape.

curiositech/windags-skills

Expert in 3D computer vision labeling tools, workflows, and AI-assisted annotation for LiDAR, point clouds, and sensor fusion. Covers SAM4D/Point-SAM, human-in-the-loop architectures, and vertical-specific training strategies. Activate on '3D labeling', 'point cloud annotation', 'LiDAR labeling', 'SAM 3D', 'SAM4D', 'sensor fusion annotation', '3D bounding box', 'semantic segmentation point cloud'. NOT for 2D image labeling (use clip-aware-embeddings), general ML training (use ml-engineer), video annotation without 3D (use computer-vision-pipeline), or VLM prompt engineering (use prompt-engineer).

curiositech/windags-skills

Implement WCAG 2.2 AA/AAA compliance with automated testing, keyboard navigation, screen reader support, and focus management. Activate on: accessibility audit, WCAG compliance, keyboard navigation, screen reader, aria attributes, axe-core, focus trap. NOT for: design-level accessibility review (use design-accessibility-auditor), color contrast only (use css-in-js-architect).

curiositech/windags-skills

Time-blind friendly planning, executive function support, and daily structure for ADHD brains. Specializes in realistic time estimation, dopamine-aware task design, and building systems that actually work for neurodivergent minds.

curiositech/windags-skills

Designs digital experiences for ADHD brains using neuroscience research and UX principles. Expert in reducing cognitive load, time blindness solutions, dopamine-driven engagement, and compassionate design patterns. Activate on 'ADHD design', 'cognitive load', 'accessibility', 'neurodivergent UX', 'time blindness', 'dopamine-driven', 'executive function'. NOT for general accessibility (WCAG only), neurotypical UX design, or simple UI styling without ADHD context.

curiositech/windags-skills

>- Apply crisis decision-making research to agent routing, uncertainty triage, and coordination failure analysis in time-pressured systems. Use when diagnosing handoff failures, analytical paralysis, or expert judgment under incomplete information. NOT for routine coding, simple CRUD design, or static single-agent tasks with complete information.

curiositech/windags-skills

Extend and modify the admin dashboard, developer portal, and operations console. Use when adding new admin tabs, metrics, monitoring features, or internal tools. Activates for dashboard development, analytics, user management, and internal tooling.

curiositech/windags-skills

Conversation patterns and interaction protocols for multi-agent systems. Covers request/response, pub/sub, blackboard, delegation chains, debate, critique, consensus, fan-out/fan-in, supervisor-worker, and peer negotiation. Deep analysis of AutoGen conversation patterns, CrewAI delegation, LangGraph state passing, and FIPA-ACL performatives. Teaches how to design what agents say to each other and in what order. Activate on: "agent conversation", "agent protocol", "multi-agent debate", "agent delegation", "supervisor worker pattern", "agent voting", "consensus protocol", "fan-out fan-in", "agent negotiation", "blackboard pattern", "agent dialogue", "conversation topology", "agent handoff". NOT for: wire format or serialization (use agent-interchange-formats), orchestration infrastructure (use agentic-infrastructure-2026), single agent behavior (use agentic-patterns).

curiositech/windags-skills

Meta-agent for creating new custom agents, skills, and MCP integrations. Expert in agent design, MCP development, skill architecture, and rapid prototyping. Activate on 'create agent', 'new skill', 'MCP server', 'custom tool', 'agent design'. NOT for using existing agents (invoke them directly), general coding (use language-specific skills), or infrastructure setup (use deployment-engineer).

curiositech/windags-skills

AI-powered calendar management and agent-based scheduling coordination. Covers calendar APIs (Google Calendar, CalDAV/iCal), AI scheduling assistants (Reclaim, Clockwise, Motion, Cal.com), building custom calendar agents with MCP, multi-calendar merging, timezone management, focus block protection, meeting fatigue detection, and agent-to-agent meeting negotiation protocols. Activate on: "calendar agent", "AI scheduling", "calendar coordination", "meeting scheduling", "calendar API", "focus time protection", "calendar optimization", "Google Calendar MCP", "Reclaim", "Clockwise", "Motion", "Cal.com", "smart scheduling", "calendar-aware agent", "timezone scheduling", "agent negotiation meetings". NOT for: manual calendar UI component design (use form-validation-architect), project management scheduling or Gantt charts (use project-management-guru-adhd), general time-tracking or pomodoro apps (use adhd-daily-planner for time-awareness), building the agent itself from scratch (use agent-creator).

curiositech/windags-skills

Build and adopt production AI agent infrastructure in 2026. Covers framework selection (LangGraph, CrewAI, AutoGen, MCP), orchestration patterns, evaluation, observability, memory systems, and tool use. Also covers the SOCIAL dimension: how to sell agent infrastructure internally, change management, measuring ROI, building trust in autonomous systems, and scaling adoption across teams. Activate on: "agent infrastructure", "agent framework comparison", "which agent framework", "sell AI tools internally", "agent adoption", "agent observability", "agent evaluation", "MCP architecture", "agentic mesh", "enterprise AI agents", "AI change management", "agent ROI". NOT for: building specific agents (use ai-engineer), designing agent behavior patterns (use agentic-patterns), prompt tuning (use prompt-engineer).

curiositech/windags-skills

Fundamental patterns for effective agentic behavior. Teaches decomposition, tool orchestration, error recovery, context management, quality self-assessment, and knowing when to stop. Model-agnostic principles that make any agent more effective regardless of domain. Activate on: "how should I structure this agent", "agentic workflow", "agent patterns", "multi-step task", "tool orchestration", "/agentic-patterns", "decompose this", "agent best practices", "chain of actions", "when should the agent stop", "agent loop design". NOT for: creating agent infrastructure (use agent-creator), building DAGs (use windags-architect), specific tool implementation.

curiositech/windags-skills

Automated discovery and matching of agent skills for dynamic task routing and capability assessment

curiositech/windags-skills

Cryptographic security for agentic systems — zero-trust agent networking, signed message envelopes (JWS/JWE), capability-based security (ocaps), Merkle tree audit trails, WASM sandboxing, and formal verification. Covers CLI dev tool security, mTLS between agents, permission boundaries (least privilege for AI agents), and supply chain security for skills/plugins. Activate on: "agent security", "zero trust agents", "secure agent communication", "capability-based security", "ocap", "signed messages between agents", "agent audit trail", "sandbox agent execution", "agent permissions", "mTLS agents", "cryptographic verification", "agent supply chain", "OWASP agentic", "prove agent did X", "tamper-proof agent logs". NOT for: application-level SAST scanning (use security-auditor), network firewall rules (use infrastructure), SOC2/HIPAA compliance (organizational), or prompt injection defense (use prompt-engineer).

curiositech/windags-skills

Data structures and serialization formats for agent-to-agent communication. Covers message envelopes, structured output schemas, capability declarations, task handoff payloads, error/retry signaling, and context windows as data structures. Deep comparison of A2A protocol, MCP, OpenAI function calling, and LangChain message types. Teaches when to use rigid schemas vs free-form with validation, typed vs untyped, streaming vs batch. Activate on: "agent message format", "agent communication schema", "agent-to-agent protocol", "A2A protocol", "MCP message format", "structured output for agents", "agent interop", "interchange format", "agent serialization", "task handoff format", "capability declaration". NOT for: what agents say to each other (use agent-conversation-protocols), orchestration topology (use multi-agent-coordination), building agent infrastructure (use agentic-infrastructure-2026).

curiositech/windags-skills

Logic-based agent programming language implementing BDI architecture for practical autonomous agent development

curiositech/windags-skills

>- Design AgentSpeak(L)-style BDI agents with context-guarded plans, selection functions, and intention stacks. Use for interruptible autonomy, agent policy, and multi-agent orchestration in dynamic environments. NOT for simple rule engines, static planners, or centralized workflows.

curiositech/windags-skills

Foundational concurrent computation model where actors communicate exclusively through asynchronous message passing

curiositech/windags-skills

license: Apache-2.0 NOT for unrelated tasks outside this domain.

관련 스킬