ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
9.5k
Citations
2.6k
Influential Citations
ACM Transactions on Graphics
Venue
2023
Year
Radiance Field methods have recently revolutionized novel-view synthesis of scenes captured with multiple photos or videos. However, achieving high visual quality still requires neural networks that are costly to train and render, while recent faster methods inevitably trade off speed for quality. For unbounded and complete scenes (rather than isolated objects) and 1080p resolution rendering, no current method can achieve real-time display rates. We introduce three key elements that allow us to achieve state-of-the-art visual quality while maintaining competitive training times and importantly allow high-quality real-time (≥ 30 fps) novel-view synthesis at 1080p resolution. First, starting from sparse points produced during camera calibration, we represent the scene with 3D Gaussians that preserve desirable properties of continuous volumetric radiance fields for scene optimization while avoiding unnecessary computation in empty space; Second, we perform interleaved optimization/density control of the 3D Gaussians, notably optimizing anisotropic covariance to achieve an accurate representation of the scene; Third, we develop a fast visibility-aware rendering algorithm that supports anisotropic splatting and both accelerates training and allows realtime rendering. We demonstrate state-of-the-art visual quality and real-time rendering on several established datasets.
Novel-view synthesis has been revolutionized by radiance field methods like NeRF, which produce stunning visual quality but require costly neural networks for training and rendering. While faster methods emerged, they inevitably traded speed for quality, and no existing method could achieve real-time display rates (≥30 fps) for unbounded, complete scenes at 1080p resolution. This paper directly addresses this critical gap by introducing 3D Gaussian Splatting, a method that achieves state-of-the-art visual quality while enabling real-time rendering. The significance is underscored by its massive impact—over 9,400 citations in just over a year—indicating it has become a foundational technique in the field.
The paper's importance also lies in its elegant departure from neural network-heavy approaches. By representing scenes with explicit 3D Gaussians and a fast differentiable renderer, the method avoids the computational overhead of MLP-based radiance fields. This makes it highly practical for real-world applications like virtual reality, augmented reality, film production, and 3D content creation, where real-time performance is essential.
The paper introduces three key innovations:
The paper demonstrates state-of-the-art visual quality on multiple datasets (e.g., Mip-NeRF 360, Tanks and Temples, Deep Blending) while achieving real-time rendering at ≥30 fps for 1080p resolution. Compared to prior methods like Instant NGP and Plenoxels, 3D Gaussian Splatting achieves higher PSNR and SSIM metrics while rendering orders of magnitude faster. Training times are competitive with the fastest prior methods (minutes to tens of minutes), making the approach practical for real-world use.
3D Gaussian Splatting has fundamentally changed the landscape of novel-view synthesis and radiance field rendering. Its combination of high quality and real-time performance has opened up new possibilities for interactive applications, from VR/AR to real-time 3D reconstruction. The method's explicit representation and fast rendering have inspired numerous follow-up works in dynamic scene rendering, large-scale scene reconstruction, and generative modeling. With nearly 10,000 citations, it stands as one of the most influential papers in computer graphics and computer vision of the 2020s.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba