ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
344
Citations
21
Influential Citations
Annual Meeting of the Association for Computational Linguistics
Venue
2024
Year
Within the evolving landscape of deep learning, the dilemma of data quantity and quality has been a long-standing problem. The recent advent of Large Language Models (LLMs) offers a data-centric solution to alleviate the limitations of real-world data with synthetic data generation. However, current investigations into this field lack a unified framework and mostly stay on the surface. Therefore, this paper provides an organization of relevant studies based on a generic workflow of synthetic data generation. By doing so, we highlight the gaps within existing research and outline prospective avenues for future study. This work aims to shepherd the academic and industrial communities towards deeper, more methodical inquiries into the capabilities and applications of LLMs-driven synthetic data generation.
This paper addresses a critical bottleneck in deep learning: the scarcity and quality limitations of real-world data. By leveraging Large Language Models (LLMs) for synthetic data generation, the field gains a scalable, data-centric solution. However, prior work lacked a unified framework, making it difficult to compare methods or identify systematic gaps. This paper fills that void by proposing a generic workflow that covers generation, curation, and evaluation, thus providing a structured lens for both researchers and practitioners.
The timing is significant: with LLMs becoming more capable and accessible, synthetic data is increasingly used in domains like NLP, code generation, and multimodal learning. Without a coherent organizational scheme, efforts remain fragmented. This paper's taxonomy helps consolidate knowledge and directs future work toward the most pressing challenges, such as data diversity, fidelity, and bias mitigation.
As a survey and position paper, no experimental results are reported. The main output is a structured organization of the literature and a set of identified research gaps. The paper does not provide quantitative metrics or comparisons between synthetic data generation methods.
This paper serves as a roadmap for the growing field of LLM-driven synthetic data. By providing a common vocabulary and workflow, it enables more systematic comparisons and collaborations. For industry practitioners, it offers a checklist for building robust synthetic data pipelines. For academics, it highlights open problems that could drive future research. The framework's generality means it can adapt as LLMs evolve, ensuring long-term relevance.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba