BeSafe-Bench: Behavioral Safety Risks of Situated Agents (2026)
FreeFirst benchmark across 4 real functional domains (Web, Mobile, Embodied VLM/VLA) with 9 safety-risk categories; even the best agent completes <40% of tasks under full safety constraints
About BeSafe-Bench: Behavioral Safety Risks of Situated Agents (2026)
BeSafe-Bench (BSB) is a benchmark designed to expose behavioral safety risks of situated agents operating in functional environments. It covers four representative domains: Web, Mobile, Embodied VLM, and Embodied VLA. The benchmark constructs a diverse instruction space by augmenting tasks with nine categories of safety-critical risks and adopts a hybrid evaluation framework combining rule-based checks with LLM-as-a-judge reasoning to assess real environmental impacts. Evaluations of 13 popular agents reveal that even the best-performing agent completes fewer than 40% of tasks while fully adhering to safety constraints, and strong task performance frequently coincides with severe safety violations. These findings underscore the urgent need for improved safety alignment before deploying agentic systems in real-world settings.
Key Features
Pros & Cons
- First comprehensive safety benchmark across multiple real-world domains
- Uses functional environments rather than low-fidelity simulations
- Hybrid evaluation (rule-based + LLM-as-a-judge) increases reliability
- Identifies critical safety gaps in state-of-the-art agents
- Open access and free (arXiv paper, code expected via links)
- Limited to four domains; may not generalize to all real-world scenarios
- Best-performing agent completes fewer than 40% of tasks safely, indicating high safety risk even for top models
- Benchmark focuses on behavioral safety risks and does not cover other safety aspects (e.g., data privacy, adversarial robustness)