ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
arXiv.org
Venue
2025
Year
A translation model trained by fine-tuning Gemma3–4B-IT. It supports 22 Indian languages - Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Urdu, Kannada, Odia, Malayalam, Punjabi, Assamese, Maithili, Santali, Kashmiri, Nepali, Sindhi, Dogri, Konkani, Manipuri (Meitei), Bodo, Sanskrit.
This paper addresses the critical need for machine translation in India, a country with 22 official languages and hundreds of dialects. By fine-tuning a relatively small (4B parameter) instruction-tuned model, the authors demonstrate a practical approach to building a multilingual translation system without requiring massive computational resources. The inclusion of low-resource languages such as Santali, Dogri, and Bodo is particularly noteworthy, as these languages are often neglected in mainstream NLP research.
The choice of Gemma3–4B-IT as the base model is strategic: it is a publicly available, efficient model that can be fine-tuned on modest hardware, making the approach accessible to researchers and developers in India. This work could serve as a foundation for more comprehensive translation systems and encourage further research on Indian language NLP.
The abstract does not provide any quantitative results, such as BLEU scores, translation accuracy, or comparisons to existing models. This is a significant omission, as the effectiveness of the fine-tuning cannot be assessed without evaluation metrics. Future work should include benchmarks on standard translation datasets for Indian languages.
This paper contributes to the democratization of machine translation for Indian languages. By releasing a fine-tuned model that covers 22 languages, it provides a valuable resource for developers, researchers, and businesses operating in India. The work also highlights the potential of fine-tuning smaller instruction-tuned models for specialized tasks, offering a cost-effective alternative to training large models from scratch. However, the lack of evaluation metrics limits the immediate impact, and further validation is needed to establish the model's practical utility.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba