ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
410
Citations
18
Influential Citations
Journal of Applied Learning & Teaching
Venue
2023
Year
Developments in the chatbot space have been accelerating at breakneck speed since late November 2022. Every day, there appears to be a plethora of news. A war of competitor chatbots is raging amidst an AI arms race and gold rush. These rapid developments impact higher education, as millions of students and academics have started using bots like ChatGPT, Bing Chat, Bard, Ernie and others for a large variety of purposes. In this article, we select some of the most promising chatbots in the English and Chinese-language spaces and provide their corporate backgrounds and brief histories. Following an up-to-date review of the Chinese and English-language academic literature, we describe our comparative method and systematically compare selected chatbots across a multi-disciplinary test relevant to higher education. The results of our test show that there are currently no A-students and no B-students in this bot cohort, despite all publicised and sensationalist claims to the contrary. The much-vaunted AI is not yet that intelligent, it would appear. GPT-4 and its predecessor did best, whilst Bing Chat and Bard were akin to at-risk students with F-grade averages. We conclude our article with four types of recommendations for key stakeholders in higher education: (1) faculty in terms of assessment and (2) teaching & learning, (3) students and (4) higher education institutions.
This paper arrives at a critical moment when higher education is grappling with the rapid proliferation of AI chatbots. By systematically testing major English and Chinese chatbots—ChatGPT (GPT-3.5 and GPT-4), Bing Chat, Bard, and Ernie—on multi-disciplinary academic tasks, the authors cut through the hype to reveal a sobering reality: none of these models perform at an A or B level. The finding that Bing Chat and Bard score F-grade averages is particularly striking, given the publicized claims of their capabilities. This matters because institutions and educators need evidence-based assessments to inform policy, assessment redesign, and teaching strategies, rather than relying on vendor marketing or anecdotal reports.
Furthermore, the paper bridges a gap in the literature by including Chinese-language chatbots (Ernie) and reviewing both English and Chinese academic sources, offering a more global perspective. The four categories of recommendations—for faculty (assessment and teaching/learning), students, and institutions—provide actionable guidance that can be immediately applied. This makes the paper a valuable resource for decision-makers in higher education.
This paper has broad implications for the AI field and higher education. It provides a reality check for the exaggerated narratives surrounding large language models, emphasizing that current chatbots are not yet intelligent enough for high-stakes academic use. The findings encourage more rigorous evaluation of AI tools before adoption, and the recommendations offer a roadmap for responsible integration. For AI practitioners, the paper underscores the need for continued improvement in reasoning, factual accuracy, and domain-specific knowledge. It also highlights the importance of multilingual and cross-cultural evaluation, as the inclusion of Chinese chatbots reveals performance disparities that may affect global adoption.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba