Benchmark deep research agents across factual, quality, and process dimensions with MiroEval
Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship.
Prerequisites
Python, uv, model result JSON, required API keys for judge and retrieval services
Installation
No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.