ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
14k
Citations
2.3k
Influential Citations
Journal of the American Statistical Association
Venue
1995
Year
From the Publisher: The past decade has seen considerable theoretical and applied research on Markov decision processes, as well as the growing use of these models in ecology, economics, communications engineering, and other fields where outcomes are uncertain and sequential decision-making processes are needed. A timely response to this increased activity, Martin L. Puterman's new work provides a uniquely up-to-date, unified, and rigorous treatment of the theoretical, computational, and applied research on Markov decision process models. It discusses all major research directions in the field, highlights many significant applications of Markov decision processes models, and explores numerous important topics that have previously been neglected or given cursory coverage in the literature. Markov Decision Processes focuses primarily on infinite horizon discrete time models and models with discrete time spaces while also examining models with arbitrary state spaces, finite horizon models, and continuous-time discrete state models. The book is organized around optimality criteria, using a common framework centered on the optimality (Bellman) equation for presenting results. The results are presented in a theorem-proof format and elaborated on through both discussion and examples, including results that are not available in any other book. A two-state Markov decision process model, presented in Chapter 3, is analyzed repeatedly throughout the book and demonstrates many results and algorithms. Markov Decision Processes covers recent research advances in such areas as countable state space models with average reward criterion, constrained models, and models with risk sensitive optimality criteria. It also explores several topics that have received little or no attention in other books, including modified policy iteration, multichain models with average reward criterion, and sensitive optimality. In addition, a Bibliographic Remarks section in each chapter comments on relevant historic
This book, published in 1995, arrived at a time when Markov decision processes (MDPs) were gaining traction across ecology, economics, and communications engineering. It provided a much-needed unified and rigorous treatment that bridged theoretical foundations with practical algorithms. The work's emphasis on the Bellman equation as a central organizing principle helped standardize how researchers approach optimality criteria, making it easier to compare and extend results. Its coverage of topics like modified policy iteration and multichain models filled gaps left by earlier texts, and the recurring two-state example made complex concepts accessible.
The book's impact is evident in its citation count of over 14,000, reflecting its role as a key reference for both researchers and practitioners. In the context of modern AI, MDPs form the backbone of reinforcement learning, and this book's systematic treatment of infinite horizon models, average reward criteria, and risk-sensitive optimality remains highly relevant. It laid the groundwork for later advances in deep reinforcement learning by clarifying the theoretical underpinnings of value iteration, policy iteration, and their variants.
The abstract does not provide quantitative experimental results or benchmarks, as the work is a textbook rather than an empirical study. However, its impact is measured by its citation count (14,290) and its adoption as a standard reference in the field. The book's contributions are theoretical and algorithmic, providing proofs and convergence guarantees for methods like value iteration, policy iteration, and modified policy iteration. It also surveys applications in ecology (e.g., harvesting strategies), economics (e.g., inventory control), and communications (e.g., queueing systems), demonstrating the breadth of MDP applicability.
This book has had a profound and lasting impact on AI and operations research. It formalized the mathematical foundations of sequential decision-making under uncertainty, directly influencing the development of reinforcement learning algorithms such as Q-learning and SARSA. By emphasizing the Bellman equation, it provided a clear path from dynamic programming to modern deep RL. The coverage of constrained and risk-sensitive MDPs anticipated later work in safe reinforcement learning and robust control. For practitioners, the book remains a go-to resource for understanding convergence properties, algorithm design, and model formulation. Its legacy is evident in its continued citation in contemporary research on reinforcement learning, optimal control, and stochastic optimization.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba