Automated benchmark
Most benchmarks use synthetic questions. This one uses real tasks that real people brought to AI assistants — drawn exclusively from consented, privacy-cleaned DATA HEDGE contributions — and scores every response on the six core AgentEval metrics.
No benchmark results published yet. The first run will appear here.
Known limitation: subject and judge currently share a model family, so scores may carry self-preference bias. A cross-vendor second judge (IBM Granite) is planned, mirroring our dual-vendor privacy pipeline. Treat absolute numbers with care; trends and findings are the signal.
Results will appear here after the first run.
AgentEval is an independent evaluation platform and is not affiliated with or endorsed by Robinhood Markets, Inc., Solana Foundation, or any evaluated project.
This is an independent AgentEval benchmark of a publicly available model. It is not an Anthropic product and implies no endorsement. Claude is a trademark of Anthropic. Scores reflect performance on the sampled tasks only and do not guarantee future performance.