ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
… actually serve as reliable world models and deliver concrete bene… to evaluate LLM-based world models: (i) fidelity and … show that sufficiently trained world models capture coherent envi…
This paper tackles a fundamental question: can large language models (LLMs) go beyond pattern matching to actually model the world? By proposing a systematic evaluation of LLMs as implicit text-based world models, it addresses a critical gap in understanding whether these models possess genuine reasoning capabilities or merely mimic surface-level statistics. The findings have implications for AI safety, planning, and interactive agents that must understand environment dynamics.
The work is timely given the rapid deployment of LLMs in decision-making roles. If LLMs can reliably model world dynamics from text, they could serve as lightweight simulators for training other AI systems or for few-shot planning in novel environments. This paper provides the first structured attempt to measure that capability.
The paper introduces two key evaluation metrics:
This research opens the door to using LLMs as general-purpose world models for text-based reasoning, planning, and simulation. It could reduce the need for hand-crafted environment models in AI training pipelines. However, the limitation to text-based settings means real-world applications require further validation. The proposed metrics provide a foundation for future work on grounding LLMs in physical or multimodal environments.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba