LogiScope-VQA:工业场景中物流危险识别的视觉语言模型基准测试
LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios
摘要
大型多模式模型 (LMM) 在工业仓库环境中的大规模部署特别需要模型表现出人类专家级的面向危险的感知、理解和推理能力。 However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA 包含主要来自现实世界物流园区的 2,476 张图像和 2,918 个视频,以及由人工注释者精心策划和验证的 10,274 个 VQA。基于 18 个核心对象和 20 个风险类型,我们设计了 39 个子任务,与三个主要主题相一致:工业元素感知、仓库知识理解和潜在风险推理。 Furthermore, we incorporate dynamic thinking-budget configurations and dual-dimensional risk bias analyses to elucidate the properties of LMMs.大量实验表明,即使是强大的专有模型,包括 GPT-5.5、Gemini-3.1-Pro 和 Claude-Opus-4.7,也与人类表现相比存在显着差距。联合整合感知、理解和推理来进行危险识别的独特挑战为 LogiScope-VQA 的进一步改进提供了巨大的空间。 We additionally reveal the pervasive security bias issue that impedes LLM' practical deployment in real-world settings. The industrial dataset is publicly available under the CC BY-NC-SA 4.0 license.