ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
A 14B language model prioritizing data quality through a training process incorporating synthetic data for pretraining and midtraining, curated organic data seeds, and innovative post-training techniques like pivotal token search for DPO, resulting in strong performance on reasoning-focused benchmarks, especially in STEM, comparable to much larger models, while also addressing overfitting and data contamination concerns.
This paper challenges the prevailing trend of scaling model size by demonstrating that a 14B parameter model can achieve competitive reasoning performance, especially in STEM domains, through a focus on data quality. The use of synthetic data for pretraining and midtraining, combined with curated organic seeds, suggests that careful data curation can be a more efficient path to strong performance than simply increasing parameters. This is particularly relevant for practitioners with limited computational resources.
The introduction of pivotal token search for DPO is a notable innovation in post-training. By identifying and focusing on the most critical tokens during preference optimization, the method may improve alignment efficiency and reduce overfitting to spurious correlations in preference data. This could have broader implications for how language models are fine-tuned for specific tasks.
The abstract states that Phi-4 achieves strong performance on reasoning-focused benchmarks, particularly in STEM, and is comparable to much larger models. However, no specific metrics, benchmark names, or comparison baselines are provided. The claim of addressing overfitting and data contamination is also stated without quantitative evidence.
This work reinforces the growing recognition that data quality is a critical factor in language model performance, potentially more important than scale. The synthetic data generation and curation techniques could be adopted by other researchers to build efficient models. The pivotal token search method for DPO offers a new direction for post-training alignment, which could improve the reliability and accuracy of language models in specialized domains like STEM. However, the lack of concrete results limits the immediate impact until further validation is provided.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba