Preprint2026
Wildclawbench: A benchmark for real-world, long-horizon agent evaluation
Unknown
Wildclawbench introduces a benchmark for evaluating frontier AI models on long-horizon, native-runtime agent tasks, showing that current models still struggle with such evaluations.
0May 1, 2026Benchmarks
arXiv