Communitygithub.com

codehealth-mcp

Real-time structural Code Health via CodeScene MCP — review before edits, verify score deltas after changes, gate commits and PRs. Use when reviewing code quality, refactoring, checking if AI changes degraded a file, or before commit/PR.

codehealth-mcp 是什么?

codehealth-mcp is a Claude Code agent skill that real-time structural Code Health via CodeScene MCP — review before edits, verify score deltas after changes, gate commits and PRs. Use when reviewing code quality, refactoring, checking if AI changes degraded a file, or before commit/PR.

兼容平台Claude Code~Codex CLI~Cursor
npx skills add https://github.com/affaan-m/everything-claude-code/tree/main/skills/codehealth-mcp

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

Code Health MCP (CodeScene)

Structural maintainability feedback for AI-assisted coding. Complements style/lint skills (coding-standards, plankton-code-quality) with design-level health scores and regression gates.

Upstream: codescene-oss/codescene-mcp-server Package: @codescene/codehealth-mcp (stdio via npx)

Security and boundaries

Opt-in (ECC): The codescene block in mcp-configs/mcp-servers.json is a template only. ECC plugin installs do not auto-enable bundled MCP servers. Copy the entry into your config only if you want it. You can exclude it during ECC install/sync with ECC_DISABLED_MCPS=codescene,....

Credentials: No bundled token. Set CS_ACCESS_TOKEN yourself (see getting-a-personal-access-token.md in the upstream repo). Never commit tokens to the repo.

What the tools read: When invoked, tools analyze files and git state in the local repository you point them at (paths you pass, plus branch context for analyze_change_set). They do not run by themselves. For standalone mode, follow upstream privacy docs: codescene-mcp-server README and CodeScene policies. Do not use this skill for secrets, credentials, or paths you do not want analyzed.

If the MCP is unavailable (offline, bad token, server crash): Do not invent Code Health scores. Tell the user the check was skipped. Continue only with explicit user approval. Prefer lint/tests/verification-loop for gating when MCP is down. Re-enable checks once the server connects.

When to Use

  • User asks to review code quality, refactor a file, or check if AI changes degraded maintainability
  • Before editing a hotspot, legacy module, or unfamiliar file
  • Before commit or pull request when you need a maintainability safeguard
  • After a large agent-written diff — verify Code Health did not regress
  • Pair with verification-loop, tdd-workflow, or /quality-gate as a structural check (not a replacement for tests/lint)

When to Activate

Same triggers as When to Use above — this heading is what ECC uses for skill auto-activation.

How It Works

1. Connect the MCP server

Copy the codescene entry from mcp-configs/mcp-servers.json into your harness MCP config.

Claude Code (~/.claude.jsonmcpServers):

"codescene": {
  "command": "npx",
  "args": ["-y", "@codescene/codehealth-mcp"],
  "env": {
    "CS_ACCESS_TOKEN": "YOUR_CS_ACCESS_TOKEN_HERE"
  }
}

Project-scoped: merge the same block into .mcp.json at the repo root.

Token setup is documented in the upstream repo (link above). Standalone mode does not require a paid CodeScene platform account for the four tools listed below. Restart the session and confirm the codescene server is connected before relying on scores.

2. Call standalone tools only

ToolWhen to use
code_health_reviewFull structural analysis before modifying a file
code_health_scoreQuick numeric score after each change (delta check)
pre_commit_code_health_safeguardBlock commits that introduce Code Health regressions
analyze_change_setBranch-level check before opening a PR

Do not call platform-only tools (e.g. repository-wide technical debt hotspot lists). Do not reference delta_analysis — not available on standalone.

3. Interpret scores (1–10)

RangeMeaningAgent behavior
9.0–10.0Green — healthySafer to extend; still prefer vertical slices
4.0–8.9Yellow — debtTread carefully; no drive-by refactors
1.0–3.9Red — severe debtNarrow scope only

4. Run the feedback loop

Before touching a file

  1. Run code_health_review on the target path.
  2. Record baseline score and listed code smells.
  3. Plan the smallest change that addresses the task.

Scope by score: below 5 — minimal diff only; 5–7 — no broad refactors; above 7 — safer to refactor, still verify after each edit.

After each change

  1. Run code_health_score on the same file.
  2. Compare to the baseline from code_health_review.
  3. If the score regressed, fix before continuing. Never mark the task done while the score is lower than when you started.

Before every commit — run pre_commit_code_health_safeguard on the repository path.

Before a PR — run analyze_change_set against the base branch (e.g. main).

Examples

Example: Flask maintainability improvement

On pallets/flask, an agent loop using only standalone tools:

  1. code_health_review on a target module (baseline 4.82)
  2. Targeted refactor addressing listed smells
  3. code_health_score after each edit
  4. pre_commit_code_health_safeguard before commit
  5. analyze_change_set before PR

Result: Code Health 4.82 → 9.1 (free standalone token only).

Example: AGENTS.md enforcement block

Paste into the project AGENTS.md or CLAUDE.md:

## Code Health (CodeScene MCP)

Before modifying any file: run `code_health_review`, note score and issues.

- Score below 5: problematic range — scope changes narrowly.
- Score 5–7: warning range — no broad refactors.

After each change: run `code_health_score` to verify delta.

- If score regressed: fix before continuing; never declare done if score dropped.

Before every commit: run `pre_commit_code_health_safeguard`.

Before PR: run `analyze_change_set`.

Example: anti-patterns vs correct loop

# BAD: Edit first, check later
[large refactor without code_health_review]

# BAD: Ignore score drop
"Tests pass" → mark task done while Code Health decreased

# BAD: Broad refactor on red-score file (below 5)
Drive-by cleanup across the module

# GOOD: review → small change → score → commit safeguard → analyze_change_set

Pairing with ECC

ECC skill / flowCode Health MCP role
coding-standardsStyle/naming; Code Health = structure/complexity
plankton-code-qualityWrite-time lint/format; Code Health = pre/post edit structural gate
verification-loop / /quality-gateAdd structural regression check before "done"
security-reviewSecurity vs maintainability — use both when relevant
tdd-workflowTests pass ≠ healthy design — check score after refactors

Context tip: ECC recommends keeping MCP count low. Enable codescene when doing substantive edits; disable when not needed.

Related Skills

  • coding-standards — baseline conventions
  • plankton-code-quality — write-time lint/format hooks
  • verification-loop — build/test/lint gate
  • tdd-workflow — test-first development
  • security-review — security checklist
  • documentation-lookup — library docs via Context7 (orthogonal)

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

accessibility

Design, implement, and audit inclusive digital products using WCAG 2.2 Level AA. Use when building or auditing UI that must meet WCAG 2.2 Level AA, or when reviewing a change for keyboard, contrast, or screen-reader support.

affaan-m/claude-api

Anthropic Claude API patterns for Python and TypeScript. Covers Messages API, streaming, tool use, vision, extended thinking, batches, prompt caching, and Claude Agent SDK. Use when building applications with the Claude API or Anthropic SDKs.

affaan-m/everything-claude-code

Development conventions and patterns for everything-claude-code. JavaScript project with conventional commits.

affaan-m/everything-claude-code-conventions

Development conventions and patterns for everything-claude-code. JavaScript project with conventional commits.

affaan-m/frontend-design

Create distinctive, production-grade frontend interfaces with high design quality. Use when the user asks to build web components, pages, or applications and the visual direction matters as much as the code quality.

affaan-m/gget

gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.

affaan-m/literature-review

Systematic literature-review workflow for academic, biomedical, technical, and scientific topics, including search planning, source screening, synthesis, citation checks, and evidence logging.

affaan-m/motion-ui

Production-ready UI motion system for React/Next.js. Use when implementing animations, transitions, or motion patterns.

affaan-m/project-guidelines-example

Example project-specific skill template based on a real production application.

affaan-m/pubmed-database

Direct PubMed and NCBI E-utilities search workflows for biomedical literature, MeSH queries, PMID lookup, citation retrieval, and API-backed literature monitoring.

affaan-m/scholar-evaluation

Structured scholarly-work evaluation for papers, proposals, literature reviews, methods sections, evidence quality, citation support, and research-writing feedback.

affaan-m/uspto-database

USPTO patent and trademark data workflow for official record lookup, PatentSearch queries, TSDR checks, assignment data, and reproducible IP research logs.

agent-architecture-audit

Full-stack diagnostic for agent and LLM applications. Audits the 12-layer agent stack for wrapper regression, memory pollution, tool discipline failures, hidden repair loops, and rendering corruption. Produces severity-ranked findings with code-first fixes. Essential for developers building agent applications, autonomous loops, or any LLM-powered feature. Use when an agent or LLM feature misbehaves and the failing layer is unknown, or before shipping an agent stack.

agent-eval

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.

agent-harness-construction

Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates. Use when defining or revising an agent

agentic-engineering

Operate as an agentic engineer using eval-first execution, decomposition, and cost-aware model routing. Use when planning or executing engineering work that agents will carry out end to end.

agentic-os

Build persistent multi-agent operating systems on Claude Code. Covers kernel architecture, specialist agents, slash commands, file-based memory, scheduled automation, and state management without external databases. Use when building a persistent multi-agent system on Claude Code with its own memory, commands, and scheduling.

agent-introspection-debugging

Structured self-debugging workflow for AI agent failures using capture, diagnosis, contained recovery, and introspection reports. Use when an agent run fails and you need a reproducible diagnosis instead of a retry.

agent-payment-x402

Add x402 payment execution to AI agents with per-task budgets, spending controls, and non-custodial wallets. Supports Base through agentwallet-sdk and X Layer through OKX Payments / OKX Agent Payments Protocol. Use when an agent must pay for something itself and needs per-task budgets, spending controls, and a non-custodial wallet.

agent-self-evaluation

Use after completing any non-trivial task. The agent self-rates its output on 5 axes — accuracy, completeness, clarity, actionability, conciseness — with concrete evidence per criterion. Produces a structured 1-5 scorecard with specific improvement suggestions.

相关技能