Community研究&データ分析github.com

tonydzi/persona-portability-benchmark

One persona, one frozen memory, N models: how much of an agent's character survives a model swap? Harness + blind multi-lens judge panel + cross-vendor rank control + contamination probe. Results included: 7 models, spread 4.80 -> 2.20.

persona-portability-benchmark とは?

persona-portability-benchmark is a Claude Code agent skill that one persona, one frozen memory, N models: how much of an agent's character survives a model swap? Harness + blind multi-lens judge panel + cross-vendor rank control + contamination probe. Results included: 7 models, spread 4.80 -> 2.20.

対応~Claude Code~Codex CLI~Cursor
npx skills add tonydzi/persona-portability-benchmark

Installed? Explore more 研究&データ分析 skills: obra/superpowers, affaan-m/quarkus-verification, affaan-m/uspto-database · View all 6 →

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

persona-portability-benchmark は何をしますか?

One persona, one frozen memory, N models: how much of an agent's character survives a model swap? Harness + blind multi-lens judge panel + cross-vendor rank control + contamination probe. Results included: 7 models, spread 4.80 -> 2.20.

関連スキル