论文

MotionBlind:探索视频中运动理解的错觉-LLM

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

模型评测基准与评测资源

摘要

视频 大语言模型 (Video-LLM) 越来越多地用作世界模型的感知前端,这个角色假设它们可以读取运动:物体移动的速度有多快、行进的方式、推动它的力度有多大。我们证明他们不能。 Video-LLM 可以在同一房间观看同一个人的两个片段,命名两个片段中的每个对象,但仍然无法说出哪个片段移动得更快。我们介绍 MotionBlind,这是一个自录制视频的对比基准,用于测量物理基础运动(速度、幅度和方向),这是世界模型必须预测的变量。 Each instance is a pair of near-identical clips that differ only in motion. Each clip carries two complementary yes/no questions, giving four items per instance, and a model earns credit only if all four are correct. We report Instance Accuracy(IAcc), which has a 6.25% chance floor. Single-frame, appearance, and language-only shortcuts all collapse to it. MotionBlind complements the recent TimeBlind benchmark.我们对六个开放视频和两个前沿视频 LLM 进行了对照研究,改变视频是否存在、帧是否以正确的时间顺序显示以及帧的采样方式(1 到 24 帧,四种选择策略)。 Open models sit near the 6.25% floor, and scale does not help. Removing the video drops every model to zero IAcc, and shuffling frames collapses IAcc to chance, so the task genuinely needs video in order. Neither more frames nor smarter frame selection closes the gap, because these change which frames are seen, not whether motion is read. Only Gemini3.1 Pro clears the benchmark overall, and even it fails on speed.无法区分同一动作的两种速度的前端还不是世界模型的值得信赖的监督、奖励或评估来源。