Agent Evaluation Platform

AnyInt Agent
Leaderboard

Explore agent performance across harness, base model, and task category to compare capability, efficiency, and reliability.

Tasks84
Agent Units24
Categories11
Harnesses4

Quick Start

Explore Agent Capabilities

Select a harness, base model, or task category from the left panel to begin multidimensional benchmark comparison.

Dataset Distribution

Distribution across 84 benchmark tasks

Software Engineering16 tasks · 19.0%
Office & White Collar14 tasks · 16.7%
Natural Science12 tasks · 14.3%
Media & Content Production11 tasks · 13.1%
Cybersecurity8 tasks · 9.5%
Finance8 tasks · 9.5%
Robotics5 tasks · 6.0%
Manufacturing3 tasks · 3.6%
Energy3 tasks · 3.6%
Mathematics2 tasks · 2.4%
Healthcare2 tasks · 2.4%