论文

RedLight-VLA:驾驶政策中交通规则基础和行为重点的模型

RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies

模型训练监督微调与指令调优

摘要

Behavior-cloned Vision-Language-Action (VLA) driving policies struggle with rare rule-governed maneuvers at signalized intersections.制动和启动示例对平均轨迹损失影响不大,而融合表示缺乏对交通灯和停车线状态的明确监督。我们提出了 RedLight-VLA,这是一种训练目标,它使用专家未来和自动生成的感知目标,无需额外的手动规则注释。首先,轨迹衍生的行为重新加权(BR)强调使用旋转不变的纵向动力学和比例保持减少的罕见减速和加速,在禁用时精确恢复基线。其次,并行辅助(AUX)以连续的融合后规则标记引导地面交通灯和停车线状态,无需自回归语言生成或对轨迹解码器的更改。 We evaluate on a curated set of 20 s sequences with a 5 s prediction horizon. Controlled variants share the same backbone, training data, decoder, and evaluation population.与其他相同的 VLA 基线相比,RedLight-VLA 将红灯停止线超调从 7.3% 降低至 6.8%,将停止线速度误差降低 12.7%,并将 3 s 交通灯切片 ADE/FDE 从 0.274/0.964 m 提高到 0.247/0.897 m。 Green-light false stops increase from 3.2% to 3.9%; however, combining BR with AUX supervision mitigates the larger increase observed for AUX alone (4.0%).该组合模型还将非红绿灯 ADE/FDE 从 0.268/0.956 m 提高到 0.241/0.876 m,并且在所有四个切片位移测量上均优于单独的任一机制。