Communitygithub.com

amplitude/builder-skills

Monitors all active and recently completed experiments across Amplitude projects, triages them by importance, then runs deep analysis and reporting on the most impactful ones. Use when the user asks to "check on experiments", "experiment status", "experiment review", "what experiments are running", or wants a periodic experiment health report.

builder-skills 是什麼?

builder-skills is a Claude Code agent skill that monitors all active and recently completed experiments across Amplitude projects, triages them by importance, then runs deep analysis and reporting on the most impactful ones. Use when the user asks to "check on experiments", "experiment status", "experiment review", "what experiments are running", or wants a periodic experiment health report.

相容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/amplitude/builder-skills/tree/HEAD/analytics-skills/skills/monitor-experiments

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

Experiment Monitor & Report Generator

Scan active and recently completed experiments, surface what needs attention, and report on the ones that matter.

This is a monitoring skill — keep output accessible to non-experts. Avoid statistical jargon (no p-values, no power analysis). For deep-dive analysis of a specific experiment, use the analyze-experiments skill instead.


CRITICAL: Managing Response Sizes

  1. get_experiments: 3-5 IDs max per call. Filter using search results BEFORE fetching.
  2. query_experiment responses are large. Extract only summary objects and validity flags. Ignore timeseries, xValues, bulk arrays.
  3. Metric name resolution: search does NOT match metric IDs in queries. Search with entityTypes: ["METRIC"], empty queries, limitPerQuery: 50, scoped to project. Match IDs from results.

Report Structure

The report has two parts:

  1. Summary & Actions (top) — One table + action items. Someone should be able to read just this and know the full picture.
  2. Details (bottom) — Deep-dives on experiments that need attention, briefs on recently decided, one-liners for monitoring experiments, and a needs-setup list.

Do NOT duplicate information between the summary table and the details. The summary table is the single source of truth for the portfolio state. Details expand on specific experiments.


Instructions

Step 1: Context & Discovery

  1. Call Amplitude:get_context. If multiple projects, ask which to monitor.
  2. Search for experiments:
Amplitude:search({
  entityTypes: ["EXPERIMENT"],
  appIds: [projectId],
  queries: [],
  sortOrder: "lastModified",
  sortDirection: "DESC",
  limitPerQuery: 50
})
  1. Filtering rules — include experiments that are:
    • Running and not stale: Any experiment in a running state that is NOT marked as stale. Stale experiments have gone idle and should be excluded.
    • Recently decided: Completed experiments that have a decision recorded AND were modified within the last 14 days. These are worth reviewing to confirm the decision or share learnings.
  2. Exclude: Drafts, disabled experiments, and stale experiments.

Step 2: Fetch Metadata, Metric Names & Primary Results

  1. Call Amplitude:get_experiments in batches of 3-5 IDs (only the filtered set from Step 1).
  2. Extract: state, dates/duration, variants, metric IDs (primary where recommendation=true), decision, owner.
  3. Apply the filtering rules from Step 1 again using the detailed metadata (some fields like stale status or decision may only be available here).
  4. Resolve metric names via Amplitude:search({ entityTypes: ["METRIC"], appIds: [projectId], queries: [], limitPerQuery: 50 }). Build { id: name } mapping.
  5. For experiments that have metrics configured, call Amplitude:query_experiment({ id: "<id>" }) (no metricIds) to get primary metric results and data quality flags. This data feeds both the summary table and the deep-dives.

CRITICAL: Never show raw metric IDs to the user. If a metric name can't be resolved, describe it by its role (e.g., "primary metric", "guardrail metric") — do NOT display strings like "rrgtky08" or "Metric gvgb5efj". Metric IDs are internal identifiers that mean nothing to a human reader.


Step 3: Summary & Actions

This is the top of the report. It should be self-contained — someone reading only this section gets the full picture.

Summary table

## Experiment Monitor: [Project Name]
Date: [Today] | Project: [Name] ([ID])

| Experiment | State | Duration | Lift | Verdict |
|------------|-------|----------|------|---------|

Column definitions:

  • Experiment — Human-readable name (NOT a URL, NOT an ID)
  • State — Running or Completed
  • Duration — Days since experiment started
  • Lift — Primary metric relative lift if available, "—" if no data. For multi-variant, show range (e.g., "-17% to -20%")
  • Verdict — What to do (see vocabulary below)

Verdict vocabulary:

VerdictWhen to use
ShipSignificant positive primary, guardrails clean, data quality good
IteratePositive signal but guardrail regression or quality concern
MonitorRunning, not yet significant, nothing to act on
AbandonSignificant negative primary, or critical data quality issues
Decided: ShipRecently completed, team decided to ship
Decided: Don't shipRecently completed, team decided not to ship
Fix configQuery error, broken exposure event, or severely imbalanced traffic
Needs metricsRunning without metrics — can't evaluate
N/ASurvey/nudge deployment, not a feature experiment

Action items (immediately after the table)

**Act now:**
- [Experiment] — [specific action with owner name]

**Keep watching:**
- [Experiment] — [brief status + when to check back]

**Needs setup:**
- [Experiment] — [what's missing + owner name]

Every experiment from the table should appear in exactly one action category. If there are no experiments in a category, omit that category.


Step 4: Details

Below the summary, provide additional detail organized into sections. Each experiment appears in exactly ONE section — no duplication.

4a: Deep-Dives (experiments needing attention)

Deep-dive on experiments with Ship, Iterate, or Abandon verdicts (up to 3). These have significant results worth examining.

For each, output the full report in a single pass:

---
## [Experiment Name]
[State] | [Duration]d | Control: [N] / Treatment: [N] | [Link to experiment]

Data Quality (first — before results):

Check data quality BEFORE reporting metric results. If there are critical issues, the reader needs that context before interpreting any numbers.

If everything passes, one line: "Data quality checks all pass." Then proceed to primary metric.

If issues found, report before the primary metric:

### Data Quality: [N issues found]

SRM (Sample Ratio Mismatch):

  • Use the srmDetected field from the API response
  • If srmDetected: true: always report first and prominently
  • Report: actual split vs. expected split with specific percentages

Traffic allocation changes:

  • Compare the current variant weights against the cumulative exposure distribution
  • If they don't match (e.g., current weights are 100/0/0 but cumulative is 15/70/15), the allocation was changed mid-experiment
  • Flag this prominently — it means some variants may no longer be receiving new traffic

Sample size:

  • <100 per variant: Flag as insufficient — too early for any conclusions
  • 100–1,000 per variant: Directional signals only, not enough for confident decisions
  • 1,000+: Adequate for analysis

Validity flags — only report failures:

FlagWhat it means when it failsSeverity
isVariancePositive = falseMetric data is invalid — statistical tests can't runCritical
isConfidenceIntervalNotFlipped = falseCalculation error in results — don't trust the numbersCritical
isMeanValid = falseMetric values are broken (NaN/infinite) — can't analyzeCritical
statsAssumptionsMetForWholeExperiment = falseStatistical assumptions aren't met — results may be unreliableHigh
hasSuspiciousUplift = trueUnusually large effect — may be a measurement error, not a real changeHigh
isPointEstimateInsideConfidenceInterval = falseInternal math inconsistency — results may be wrongHigh
isStandardErrorLargeEnough = falseToo much noise to get reliable estimatesMedium

If SRM is detected or a Critical flag fails:

⚠️ Results below should be interpreted with caution due to [issue].

Primary Metric:

### [Metric Name]: [Plain-language verdict]

| Variant | Value | Lift | 95% CI |
|---------|-------|------|--------|
| Control | [X] | — | — |
| Treatment | [Y] | [+Z%] | [A% to B%] |

[One sentence plain-language interpretation.]

Framing rules:

  • Do NOT include p-values. Non-experts don't know what they mean.
  • Use "not significant" instead of "no effect." There may be an effect — the experiment just can't detect it at this sample size.
  • Lead with what happened in plain language.

Verdict words:

  • Significant positive — CI entirely above zero, positive lift
  • Significant negative — CI entirely below zero, negative lift
  • Not significant — CI includes zero; we can't confidently say there's a real difference yet
  • Trending positive/negative — Directional signal but CI includes zero; worth watching

Add one sentence on practical significance only if lift is very small (<2%) or very large (>20%).

Guardrails & Secondary Metrics:

Call Amplitude:query_experiment({ id: "<id>", metricIds: [...] }).

Focus on catching regressions, not narrating every metric.

If no significant regressions: "All guardrails and secondary metrics are clean — no significant regressions detected."

If significant regressions found:

- **[Metric Name]:** Significant regression ([−X%], CI: [A% to B%]) — [one sentence on impact]

If large directional regressions (>20% relative lift) that aren't yet significant:

- **[Metric Name]:** Not significant, but directionally negative ([−X%]) — worth watching

Small directional changes that aren't significant should be omitted. If 5+ metrics, note that multiple comparisons increase the chance of false positives.

Feedback (only if primary is significant):

Call Amplitude:get_feedback_insights with experiment-related keywords. Report 2-3 themes in one line each. If no relevant feedback, skip the section.

Verdict:

### Verdict: [SHIP / ITERATE / MONITOR / ABANDON]
1. [Primary metric finding — one sentence]
2. [Quality / guardrail status — one sentence]
3. [What to do next — one sentence]

Decision matrix (internal reference, don't output):

SignalShipIterateMonitorAbandon
PrimarySignificant positive + practicalSignificant positive but guardrail regressionNot yet significant, still runningSignificant negative
QualityAll passMinor flagsAll pass or insufficient dataSRM or critical flags
GuardrailsClean1 regression needs fixingClean or not enough dataSignificant regression on critical metric
Experiment stateCompletedCompletedRunningCompleted or running

4b: Running Experiments (monitoring, no action needed)

For experiments with Monitor verdict, provide a one-line status each:

---
### Running Experiments

**[Name]** — Running [X]d, Control: [N] / Treatment: [N]. [Primary metric status — e.g., "trending positive (+3.3%) but not yet significant"]. Data quality clean. [Link]

**[Name]** — Running [X]d, Control: [N] / Treatment: [N]. Too early for conclusions — [reason, e.g., "sample sizes are 100-1,000 range"]. [Link]

Keep these tight — one line per experiment. Include the link at the end for anyone who wants to check in Amplitude.

4c: Recently Decided (brief summaries)

For experiments with Decided: Ship or Decided: Don't ship verdict:

---
### Recently Decided

- **[Name]** — [Shipped/Rolled back] on [date] after [X] days. [One sentence on outcome or rationale]. [Link]

If the experiment was shipped without metrics or without statistical significance, note that briefly.

4d: Needs Setup

For experiments with Fix config, Needs metrics, or N/A verdict:

---
### Needs Setup

- **[Name]** — [Issue description]. Owner: [owner].
- **[Name]** — Survey/nudge deployment, not a feature experiment. No analysis needed.

Each experiment appears in exactly one detail section. Do NOT list an experiment in both the summary table actions AND a detail section with the same information — the detail section expands on the action item, it doesn't repeat it.


Edge Cases

  • No experiments: Report clearly, suggest checking other projects
  • No data (<24hrs): Note the experiment just started, check back in 7 days
  • SRM detected: Lead with this in data quality — it can invalidate everything else
  • 10+ experiments: Summary table for all, deep-dive top 3, offer to analyze more
  • Response too large / saved to disk: Extract summary + validity flags only
  • No metrics configured: List in "Needs Setup" section, don't deep-dive
  • All experiments are stale: Report that there are no active experiments to monitor
  • Query error from query_experiment: Report the error clearly with the experiment owner's name. Flag as "Fix config" in the summary. Common causes: broken exposure event property, misconfigured experiment key.
  • Traffic allocation changed mid-experiment: Flag in data quality. Compare current weights to cumulative exposure distribution. Note which variants are no longer receiving traffic.
  • Survey/nudge experiments (NPS, etc.): Note in summary as N/A. No analysis needed — skip deep-dive.
  • Severely imbalanced traffic (e.g., 5%/95% or 0%/100%): Flag as "Fix config" — can't run a valid experiment without a meaningful control group. Note in needs-setup with owner.

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

amplitude/builder-skills

Performs deep analysis of a specific Amplitude chart to explain trends, anomalies, and likely drivers. Use when a metric looks unusual, investigating a spike or drop, or understanding the "why" behind numbers.

amplitude/builder-skills

Deeply analyze Amplitude dashboards by analyzing key charts, surfacing top areas for concern and takeaways, identify anomalies, then explain changes using customer feedback trends.

amplitude/builder-skills

Designs A/B tests with proper metrics and variants, analyzes running or completed experiments, and interprets results with statistical rigor. Use when setting up experiments, checking experiment status, analyzing results, or making ship decisions.

amplitude/builder-skills

Synthesizes customer feedback into actionable themes including feature requests, bugs, pain points, and praise. Use when planning product roadmap, understanding user sentiment, investigating specific issues, or preparing voice-of-customer reports.

amplitude/builder-skills

Analyze MCP server usage instrumented with Amplitude's MCP Analytics SDK: break usage and errors down by tool, read the rationales within each tool to see what callers are trying to do, and produce a prioritized write-up of actionable fixes. Use this skill whenever the user asks to understand how their MCP server is being used, what agents/users are trying to do with it, why tool calls are failing, what to fix or improve in their MCP server, or asks for an "MCP usage report", "tool error analysis", "intent analysis", "rationale clustering", or "MCP insights". Also trigger when the user mentions [MCP]-prefixed events, tool rationale, tool call errors, or just finished instrumenting their MCP server and wants to see what the data says. Requires the Amplitude MCP connector.

amplitude/builder-skills

Read lost deals and churned accounts from your CRM, extract reasons clustered by theme (missing features, pricing, competitors, UX), and write a prioritized weekly analysis with product improvement recommendations. Use before roadmap planning or to build the case for prioritizing retention work.

amplitude/builder-skills

Creates Amplitude charts from natural language descriptions, handling event selection, filters, groupings, and visualization choices. Use when you know what you want to measure but prefer not to build the chart manually.

amplitude/builder-skills

Guide an Amplitude user through building a custom agent by suggesting use cases grounded in their role and data, shaping the idea into a well-formed spec, and generating a ready-to-run Global Agent deeplink that creates it. Use to create, build, or set up a custom agent, automate a recurring analysis, or put a repeated report on a schedule.

amplitude/builder-skills

Builds comprehensive Amplitude dashboards from requirements or goals, organizing charts into logical sections with appropriate layouts. Use when creating a complete dashboard from scratch or assembling existing charts into a cohesive view.

amplitude/builder-skills

Pull Intercom tickets and Slack support messages from the past 7 days, classify each signal, enrich with CRM data (ARR, plan, renewal), score by customer value and churn risk, and output a tiered priority report saved to Drive. Use when you need a fast, data-driven view of what support signals matter most.

amplitude/builder-skills

Use this skill whenever a user wants to improve existing pages on their website to get cited more by AI models — whether they say "our pages aren't getting cited", "improve this page for AI visibility", "which of our pages should we update", "make this article more cite-worthy", "our competitors are getting cited instead of us", "update our content for AI search", or any variation where the goal is improving an existing asset rather than creating something new. This skill pulls owned pages from AI Visibility, identifies which ones have citation potential but are underperforming, compares them against the external pages that are winning citations on the same topics, and produces section-level rewrites or a full-page update — then pushes the revision to the CMS as a draft. Trigger even if the user just says "help me get cited more" or "why is [competitor] getting cited instead of us".

amplitude/builder-skills

Use this skill whenever a user wants to win AI citations on prompts that competitors currently dominate — whether they say "competitors are getting cited instead of us", "we're losing on these prompts", "how do I outrank [competitor] in AI answers", "find prompts where we should be winning", "create content to beat [competitor]", or any variation where the goal is capturing AI share on prompts a competitor currently owns. This skill pulls competitor visibility data from AI Visibility, identifies the specific prompts where competitors win and Amplitude is absent, clusters them by intent, and produces targeted comparison pages, alternatives content, or rebuttal assets — then pushes drafts to CMS. Trigger on any mention of competitor, prompt hijack, outrank, or "why is [competitor] getting cited instead of us".

amplitude/builder-skills

Use this skill whenever a user wants to turn AI Visibility data into published content — whether they say "find content gaps", "what should we write about", "which topics have low visibility", "help me get cited by AI models", "create a blog post from our AI Visibility gaps", "we're losing to competitors on these prompts", or any variation where they want to go from AI visibility weakness to a draft article, landing page, or FAQ. This skill connects directly to Amplitude AI Visibility data (topics, prompts, visibility scores, citations, competitor data, full LLM responses and sources) and produces a publish-ready content brief plus full article draft. If the user mentions CMS (WordPress, Webflow, Contentful, Sanity, HubSpot, Ghost, Shopify), also trigger this skill to push the draft directly. Trigger even if they just say something vague like "what content should we create?" in an AI Visibility context.

amplitude/builder-skills

Use this skill whenever a user wants to test content variants before publishing to find which one will get cited most by AI models — whether they say "which version of this content will perform better", "test this article before we publish", "simulate how AI will respond to this content", "which angle should we use", "generate content variants and pick the winner", "run a simulation before publishing", or any variation where the goal is data-driven content selection rather than gut-feel publishing. This skill takes an identified content opportunity, generates 2–3 distinct variants with different angles or structures, scores them against actual AI model responses from AI Visibility, references the Simulate Changes feature for pre-publish validation, and produces a clear recommendation on which variant to publish — then pushes the winner to CMS. Trigger on any mention of "simulate", "test variants", "which performs better", "A/B content", or "before we publish".

amplitude/builder-skills

Use this skill whenever a user wants to understand which external sources are being cited by AI models on topics relevant to their brand, and wants to create content that will outrank those sources — whether they say "what sources are AI models citing", "why is [third-party site] being cited instead of us", "we want to be the definitive source on X", "build something that gets cited more than G2 or TechRadar", "create an authoritative asset", or any variation where the goal is producing a new reference asset (definition page, benchmark, methodology, glossary, comparison hub) designed to beat existing top-cited sources. This skill analyzes AI Visibility source data, reverse-engineers what makes top-cited pages authoritative, and produces a superior source asset — then pushes it to CMS as a draft. Trigger on any mention of "sources", "third-party citations", "authoritative content", "definitional pages", or "outrank".

amplitude/builder-skills

Instrument a Node/TypeScript MCP server with Amplitude's @amplitude/mcp-analytics SDK so tool calls, sessions, and rationale are tracked as Amplitude events. Use this skill whenever the user wants to add Amplitude analytics to their MCP server, mentions "MCP Analytics", "@amplitude/mcp-analytics", "instrument my MCP server", "track MCP tool calls", "add rationale to my MCP tools", or wants agent traffic (Claude, Cursor, ChatGPT) attributed back to Amplitude. Also use for adding UTM tagging to MCP-returned links, or for troubleshooting identity/user_id mismatches between MCP events and web/mobile Amplitude data.

amplitude/builder-skills

Instruments a pull request with Amplitude analytics that conform to the project's existing taxonomy. Reads the tracking plan via the Amplitude MCP server (events, properties, naming conventions), analyzes the PR diff to find the few user actions genuinely worth tracking, detects the codebase's SDK and tracking patterns, and adds instrumentation that matches both. Optionally (opt-in) stages new events and properties on an Amplitude tracking-plan branch for data-governance review. Use when asked to "instrument this PR", "add analytics to this change", "add tracking", "add Amplitude events", "instrument this feature", or "what should I track here".

amplitude/builder-skills

Diagnoses product health by cross-referencing Amplitude analytics (dashboards, charts, funnels, feedback, AI agent analytics), optionally Datadog (errors, latency, stack traces), and optionally Slack (qualitative feedback, bug reports, feature requests). Identifies what's broken, what's working, and what to do about it — with root causes, not just symptoms. Use when asked to "diagnose my product", "what's going on", "product health check", "what's broken", "where are users struggling", "give me a product diagnosis", or "what should I focus on".

amplitude/builder-skills

Turn one or more meeting transcripts, notes, or Slack threads into concise takeaways and clear action items with DRIs. Works with a single meeting or a batch from the whole week.

amplitude/builder-skills

Summarizes B2B account health by analyzing usage patterns, engagement trends, risk signals, and expansion opportunities. Use for customer success reviews, renewal preparation, QBRs, or account prioritization.

相關技能