Communitygithub.com

SkillMedev/skills

Classifies incident severity (SEV1-4) using impact, scope, and urgency signals and decides who to page. Use when an alert fires or a report comes in and a severity call must be made quickly.

O que é skills?

skills is a Claude Code agent skill that classifies incident severity (SEV1-4) using impact, scope, and urgency signals and decides who to page. Use when an alert fires or a report comes in and a severity call must be made quickly.

Funciona com~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/SkillMedev/skills/tree/HEAD/skills/sev-triage

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

SEV Triage

Severity is a paging decision, not a feelings meter. Assign the highest SEV that any single signal justifies, then downgrade only with evidence.

Severity Definitions

  • SEV1 - Revenue-impacting or total service loss. More than 5% of users cannot complete a critical flow. Data loss or breach possible. Page IC + engineering lead + exec on-call immediately.
  • SEV2 - Significant degradation. Core feature broken for a subset, elevated error rate above SLO breach threshold, or a workaround exists but is unacceptable long-term. Page IC + team on-call within 5 minutes.
  • SEV3 - Partial or minor degradation. Non-critical feature affected, SLO still within budget, no customer escalations yet. Ticket created, team notified async, fix before next business day.
  • SEV4 - Cosmetic or edge case. No user impact, caught proactively. Ticket only.

Signal Checklist

Answer these to land on a SEV. First 'yes' that matches wins.

  1. Is a payment, auth, or data-integrity flow broken for any user? -> SEV1 candidate.
  2. Is error rate above your SLO threshold for the last 10 minutes? -> SEV2 at minimum.
  3. Are multiple regions or availability zones affected? -> Escalate one level.
  4. Is there active data exfiltration or corruption risk? -> SEV1 regardless of user count.
  5. Is a single non-critical endpoint slow but the rest healthy? -> SEV3.

Scope Multipliers

  • Blast radius: estimate percentage of users affected. Below 1% rarely justifies SEV1 unless the 1% are paying customers or the data risk is severe.
  • Growth rate: is impact spreading? A SEV3 that doubles every 10 minutes is a SEV2.
  • Detection lag: if the issue started more than 30 minutes ago undetected, assume blast radius is larger than current signals show.

Who to Page

  • SEV1: IC (incident commander), service owner, on-call SRE, and engineering lead. Open a war room immediately.
  • SEV2: IC and service owner on-call. Invite others as needed.
  • SEV3: Service team on-call async. No war room.
  • SEV4: Ticket, no page.

Worked example: making the call

Bad: "Checkout seems broken for some users. Feels like a SEV3 - I'll dig into the cause first and page someone if it turns out to be serious."

Good: "Checkout error rate at 12% for the last 15 minutes (SLO threshold: 1%). Payment flow broken -> signal 1 -> SEV1. Paged IC, service owner, and on-call SRE at 14:32; war room open. Blast radius estimate: ~12% of checkout attempts, growth unclear. Cause unknown - irrelevant to the SEV. Re-triage at 15:00."

The bad version inverts the order (cause before severity), hedges on a broken payment flow, and delays paging. The good version cites the signal that set the level, pages immediately, records blast radius, and schedules the re-triage - all before anyone knows why it broke.

Deliverable

Produce a triage record - posted to the incident channel or ticket - containing: the assigned SEV, a one-line justification naming the checklist signal that set it, the blast radius estimate (percentage of users or requests), who was paged and when, and the scheduled re-triage time. The record is what lets the next responder trust the level instead of re-deriving it.

Quality bar

  • The SEV was assigned from impact signals before any root-cause investigation began.
  • The justification names a specific checklist signal, not a gut feel.
  • Paging happened on assignment, not after confirmation - a potential SEV1 pages on one data point.
  • A re-triage time is set (every 30 minutes) and the level moves with the evidence, both up and down.

Do / Don't

  • Do: assign SEV before investigating root cause. Severity is about impact now, not cause.
  • Do: re-triage every 30 minutes. A SEV2 resolved is a SEV4; a SEV3 spreading is a SEV1.
  • Don't: under-SEV to avoid waking people. False SEV1s are cheaper than missed ones.
  • Don't: wait for a second data point before paging on a potential SEV1.

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

SkillMedev/skills

Designs REST API surfaces - resource naming, HTTP method and status-code semantics, error shapes, pagination, and filtering - and delivers an endpoint spec a consumer can build against without asking questions. Use when someone asks "how should I name this endpoint", "what status code should this return", "should this be PUT or PATCH", "how do I paginate this list", or is reviewing an API before it ships to external consumers. Do NOT use for planning breaking-change rollouts and deprecation windows - use api-versioning-strategist instead; for GraphQL type and resolver design - use graphql-schema instead; for generating client SDKs from an existing spec - use api-client-generator instead; for designing inbound webhook endpoints - use webhook-receiver-hardener instead.

SkillMedev/skills

Turns data and charts into a decision-driving narrative structured as headline finding, trend, implication, and recommended action - with finding-led chart titles, context for every number, annotation guidance, and honest flags on any conclusion the data cannot support. Use when someone says "turn these numbers into a story", "what's the takeaway from this data", "help me present these results to leadership", or has charts but no narrative. Do NOT use for compressing a long document into a one-pager - use executive-summary instead - or for running the analysis that produces the findings - use eda-playbook instead.

SkillMedev/skills

Use when a task needs live or historical money data - "convert USD to EUR", "current/past exchange rate", "FX rate on this date / over this range", or "current price of Bitcoin/Ethereum, market cap, 24h change". Frankfurter (ECB reference rates, no key) is the FX default; CoinGecko's free keyless tier covers crypto. Do NOT use for stock quotes or equities - no keyless stock API survives verification, say so instead of guessing; do NOT use for country economic indicators like GDP or inflation series - use government-open-data instead; if the request is a vague "I need live data", route through public-data-api-picker.

SkillMedev/skills

Builds a driver-based FP&A operating model linking business inputs to P&L, balance sheet, and cash flow outputs. Use when building an annual plan, preparing investor materials, running scenario analysis, or stress-testing the business.

SkillMedev/skills

Use when a task needs live geographic lookups - "geocode this address", "what's at these coordinates" (reverse geocoding), "lat/lon for this city", "which country/state is this ZIP or postal code in", or "country facts: capital, currency, population, flag". Nominatim (OpenStreetMap) is the geocoding default; Zippopotam for postal codes; APICountries for country facts. All keyless. Do NOT use for weather at a location - use weather-climate instead; do NOT use for country-level statistics over time (GDP, population trends) - use government-open-data instead; if the request is a vague "I need live data", route through public-data-api-picker.

SkillMedev/skills

Runs the full Getting Things Done loop - capture, clarify, organize, reflect, engage - building a trusted system of context lists, a projects list with defined next actions, and a weekly review habit. Use when someone says "I'm overwhelmed and things are slipping through the cracks", "set up GTD for me", "help me do a brain dump and organize it", or "my to-do list is a mess". Do NOT use for just running the weekly review ritual itself - use weekly-review instead - or for clearing an email backlog - use inbox-zero.

SkillMedev/skills

Processes any email backlog to zero using the 4Ds - Delete, Delegate, Defer, Do - with a mass-archive strategy for the obvious, a touch-each-email-once discipline, and a keep-it-clear system of batched processing windows, ruthless unsubscribing, filters, and a minimal folder setup. Use when someone says "I have 5,000 unread emails", "help me get to inbox zero", "email is eating my whole day", or treats their inbox as a to-do list. Do NOT use for drafting the reply emails themselves or prioritization rules for an ongoing support queue - use email-triage instead - or for protecting focus time around the email windows - use deep-work-planner instead.

SkillMedev/skills

Runs structured coaching sessions using values clarification and the GROW model, ending every session with one committed action, a deadline, and an if-then plan for the likely obstacle. Use when someone says "I feel stuck in my life", "help me figure out what I want", "hold me accountable to my goals", or "coach me through this decision". Do NOT use for building a stress toolkit - use stress-management instead - or a journaling practice - use journal-framework; for a standing goal-tracking system, use goals-accountability. Coaching, not therapy: signs of clinical distress route to a licensed professional.

SkillMedev/skills

Use the Skill Me catalog from inside any conversation - discover, install, and manage Claude skills through the Skill Me MCP, and load installed skills automatically each session.

SkillMedev/skills

Writes and tunes PySpark jobs - join strategy and broadcast size limits, shuffle-partition sizing, skew diagnosis and salting, UDF avoidance, caching, and output file layout - with concrete size and skew thresholds. Use when someone asks "why is my Spark job slow", "should I broadcast this join", "one task takes forever while the rest finish", "my job OOMs during a join", or is writing a new PySpark ETL job. Do NOT use for Kafka topic, consumer-group, or streaming-pipeline design - use kafka-pipelines instead; do NOT use for single-machine dataframe work that fits in memory - use pandas-expert instead.

SkillMedev/skills

Builds clean, performant, accessible SwiftUI views with correct state ownership, scoped invalidation, and smooth list scrolling, and reviews existing SwiftUI code against a concrete frame-time and re-render budget. Use when someone asks "why does my SwiftUI list stutter", "should this be @State or @Observable", "my whole screen re-renders when one row changes", "how do I animate this transition", or wants a SwiftUI view built or refactored. Do NOT use for cross-platform React Native apps - use react-native-pro instead; do NOT use for Flutter widget trees - use flutter-widget-architect instead; do NOT use for Android Compose UIs - use jetpack-compose-builder instead.

Habilidades Relacionadas