Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model...
| Benchmark | Score | Source |
|---|---|---|
| Chatbot Arena Elo | 1141 | |
| IFEval | 73.1 | |
| MMLU-Pro | 48.3 |
| Benchmark | Score | Source |
|---|---|---|
| HumanEval |
72.6 |
| Benchmark | Score | Source |
|---|---|---|
| ARC-Challenge | 83.4 |
| Benchmark | Score | Source |
|---|---|---|
| HellaSwag | 82 | |
| WinoGrande | 78.9 |
| Benchmark | Score | Source |
|---|---|---|
| TruthfulQA | 52.4 |
Aion Labs