Community研究&データ分析github.com

linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

agent-eval-harness とは?

agent-eval-harness is a Claude Code agent skill that live, open-source benchmark for comparing AI coding agents on real GitHub issues.

対応~Claude Code~Codex CLI~Cursor
npx skills add linny006/agent-eval-harness

Installed? Explore more 研究&データ分析 skills: obra/superpowers, affaan-m/quarkus-verification, affaan-m/uspto-database · View all 6 →

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

agent-eval-harness は何をしますか?

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

関連スキル