ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

Abstract (EN)

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their embodied safety limits remain poorly understood. To address this gap, we introduce ForesightSafety-VLA, a diagnostic benchmark that makes safety the primary evaluation target for VLA systems. We define a 13-category safety taxonomy covering physical interaction safety (Safe-Core), instruction-side safety (Safe-Lang), and perception-side safety (Safe-Vis), and evaluate policies under three controlled dimensions of variation -- scene structure, language command, and visual observation -- so that failure sources can be diagnosed rather than hidden in a single aggregate score. Beyond binary task success, ForesightSafety-VLA measures process-level risk through cumulative safety cost (CC) and risk exposure time (RET), together with a four-quadrant decomposition of safe/unsafe success and failure. We instantiate 66 safety-augmented base scenarios in RoboTwin across 5 embodiments and report results on representative VLA baselines. Across the evaluated baselines, even the strongest policy incurs non-trivial safety cost and unsafe nominal success, while structure and visual variation induce substantially stronger safety degradation than ordinary language variation. These results suggest that embodied safety is tightly coupled to perception, grounding, and control competence rather than being reducible to post-hoc safety filtering alone.

摘要 (ZH)

在具身智能领域，安全性是机器人在物理世界中可靠部署的前提条件。当前视觉-语言-动作模型正持续向通用任务能力迈进，但其具身安全边界尚不明确。为填补这一空白，我们提出了ForesightSafety-VLA——一个以安全性为主要评估目标的诊断性基准。我们定义了涵盖物理交互安全（Safe-Core）、指令侧安全（Safe-Lang）和感知侧安全（Safe-Vis）的13类安全分类法，并在场景结构、语言指令和视觉观测三个受控变化维度下评估策略，从而能够诊断故障来源而非将其隐藏于单一综合评分中。除二元任务成功指标外，ForesightSafety-VLA通过累积安全代价和风险暴露时间测量过程级风险，并结合安全/不安全的成功与失败四象限分解。我们在RoboTwin中实例化了66个安全增强基础场景，涵盖5种具身形态，并报告了代表性VLA基线的结果。在评估的基线中，即使最强策略也会产生不可忽视的安全代价和不安全的标称成功，而结构与视觉变化比普通语言变化导致更显著的安全退化。这些结果表明，具身安全性紧密耦合于感知、基础能力与控制能力，而非仅能通过事后安全过滤来简化处理。

← Back