agent-eval-harness 是做什么的?
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
一个用于在真实 GitHub Issue 上对比 AI 编码智能体的实时开源基准测试工具。
`agent-eval-harness` 是一个开源的基准测试工具,旨在实时评估和比较不同 AI 编码助手(如 Claude Code、Cursor 等)在解决真实 GitHub Issue 时的性能。它通过从 GitHub 仓库中提取真实问题,并让不同的 AI 智能体独立尝试解决,从而提供客观的量化对比数据。该工具不仅支持多种主流的 AI 编码智能体,还允许用户自定义测试任务,非常适合开发者研究、比较或优化 AI 编码智能体的实际工作能力。其核心优势在于其真实性和动态性:它使用真实世界的编程任务(而非人工构建的玩具问题),并持续更新,确保测试结果能反映当下实际场景中的表现。
npx skills add linny006/agent-eval-harnessInstalled? Explore more 研究与数据分析 skills: obra/superpowers, affaan-m/quarkus-verification, affaan-m/uspto-database · View all 6 →
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
Agent skill repository discovered by 10x-chat research.
Verification loop for Quarkus projects: build, static analysis, tests with coverage, security scans, native compilation, and diff review before release or PR.
USPTO patent and trademark data workflow for official record lookup, PatentSearch queries, TSDR checks, assignment data, and reproducible IP research logs.
Structured scholarly-work evaluation for papers, proposals, literature reviews, methods sections, evidence quality, citation support, and research-writing feedback.
Systematic literature-review workflow for academic, biomedical, technical, and scientific topics, including search planning, source screening, synthesis, citation checks, and evidence logging.
Evidence-first current-state research workflow for ECC. Use when the user wants fresh facts, comparisons, enrichment, or a recommendation built from current public evidence and any supplied local context.