Benchmarking modern multiprocessors
Kai Li, Christian Bienia
This paper analyzes how pre-intervention exercise habits and baseline depression levels predict adherence, contamination, and dropout rates in walking and control groups.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Kai Li, Christian Bienia
This paper analyzes how pre-intervention exercise habits and baseline depression levels predict adherence, contamination, and dropout rates in walking and control groups.
Unknown
This paper introduces Llmc, a versatile toolkit for benchmarking LLM quantization methods, enabling standardized evaluation and comparison.
Xianfu Cheng, Shiwei Zhang, Jiyu Zhao, et al.
FinanceComplexQA is a benchmark with 2,026 deep research tasks on 1,009 financial documents for evaluating agentic reasoning in complex financial QA.
Hao Liang, Qihan Lin, Zhaoyang Han, et al.
Introduces K12-KGraph, a curriculum-aligned knowledge graph from Chinese textbooks, with benchmark and training data to improve LLMs' curriculum cognition.
Unknown
SEED-Bench is a large-scale benchmark for evaluating multimodal large language models across hierarchical capabilities.
Frank F. Xu, Yufan Song, Boxuan Li, et al.
Introduces TheAgentCompany, an extensible benchmark for evaluating LLM agents on real-world professional tasks.
Yuyao Wang, Zhongjian Zhang, Mo Chi, et al.
EvoMemBench benchmarks agent memory from a self-evolving perspective, providing a unified testbed for adaptive memory in execution-oriented tasks.
Zexue He, Yu Wang, Churan Zhi, et al.
Proposes MemoryArena, a benchmark for evaluating agent memory in multi-session, interdependent tasks.
Unknown
This paper introduces Riosworld, a benchmark for evaluating the risk of multimodal computer-use agents in real-world manipulation tasks.
Reyna Abhyankar, Qi Qi, Yiying Zhang
This paper benchmarks the temporal efficiency of 16 popular computer-use agents in realistic environments.
Xu Li, Simon Yu, Minzhou Pan, et al.
Introduces MTAgentRisk, the first multi-turn safety benchmark for tool-using agents, revealing a 16% average increase in attack success rate across multiple turns.