ACLUE
Freean evaluation benchmark focused on ancient Chinese language comprehension.
About ACLUE
ACLUE (Ancient Chinese Language Understanding Evaluation) is a benchmark designed to evaluate the performance of large language models (LLMs) on ancient Chinese language comprehension. It consists of 15 multiple-choice tasks spanning five domains: vocabulary (e.g., polysemy, interchangeable characters, named entity recognition), syntax (sentence segmentation), semantics (couplet completion, poetry line prediction), reasoning (poetry quality assessment, reading comprehension, poetry appreciation, emotion classification), and knowledge (ancient Chinese knowledge, traditional culture, medical ancient texts, literary history, historical phonetics). The dataset covers texts from the Xia Dynasty (c. 2070 BCE) to the Ming Dynasty (1368 CE). A leaderboard tracks zero-shot performance for models like ChatGLM2-6B, ChatGPT, BLOOMZ-7B, and others. The project is open source under the MIT license, with data licensed under CC BY-NC-SA 4.0. A paper describing ACLUE was presented at the Ancient Language Processing Workshop (ALP 2023).
Key Features
Pros & Cons
- Comprehensive coverage of multiple aspects of ancient Chinese language
- Structured evaluation with a public zero-shot leaderboard
- Open source and freely available for research
- Based on a peer-reviewed academic paper
- Includes diverse task types from vocabulary to reasoning
- Limited to ancient Chinese language only, not multilingual
- Multiple-choice format may not capture nuanced generative understanding
- Relatively small number of tasks (15) compared to broader benchmarks
- Leaderboard currently shows low performance even for top models, indicating difficulty