ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
… Our results show that long-horizon planning over massive tool ecosystems remains highly challenging. While most models remain below two-thirds accuracy in the default setting, …
As LLMs are increasingly deployed as autonomous agents, their ability to plan over long horizons and select from vast tool ecosystems becomes paramount. This paper introduces PlanBench-XL, a benchmark that directly tests these capabilities at scale. The finding that most models remain below two-thirds accuracy underscores a fundamental limitation in current LLM-based planning, which is critical for applications like automated software development, scientific discovery, and enterprise workflows.
The significance lies in exposing the gap between simple tool use (e.g., single API calls) and complex, multi-step planning over hundreds or thousands of tools. This work provides a much-needed stress test for the community, pushing beyond existing benchmarks that often use small, curated tool sets.
This paper sets a new challenge for the AI community, highlighting that scaling tool ecosystems exposes fundamental weaknesses in LLM planning. It will likely spur research into hierarchical planning, retrieval-augmented tool selection, and memory-augmented agents. For practitioners, it suggests that current LLM agents are not yet reliable for complex, real-world tool orchestration tasks, guiding investment in more robust planning frameworks.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba