Cite this article as:

Shaobo Li, Qinghua Gu, Shunling Ruan, and Song Jiang, Event-level abnormal driving behavior detection in mining trucks based on temporal-aware vision-language models, Int. J. Miner. Metall. Mater., (2026). https://doi.org/10.1007/s12613-026-3451-4
Shaobo Li, Qinghua Gu, Shunling Ruan, and Song Jiang, Event-level abnormal driving behavior detection in mining trucks based on temporal-aware vision-language models, Int. J. Miner. Metall. Mater., (2026). https://doi.org/10.1007/s12613-026-3451-4
引用本文 PDF XML SpringerLink

基于时序感知视觉语言模型的矿用卡车事件级异常驾驶行为检测

摘要: 不安全驾驶行为是露天矿作业事故的重要诱因。传统监测系统往往难以捕捉此类行为的时间连续性和语义模糊性。为解决上述问题,本文提出一种基于视觉语言模型(VLM)的事件级异常驾驶行为识别方法。该方法将时序感知低秩适配(T-LoRA)与事件级监督微调相结合,实现视觉语言模型面向复杂露天矿作业环境的高效垂直领域定制。通过显式增强视觉序列的时序建模能力,所提方法能够准确识别事件级异常驾驶行为,精确定位其发生起点和持续时间,并生成可解释的语义描述。基于真实矿用卡车数据集的实验结果表明,该方法的事件级 F1 值达到 0.923,时间定位平均绝对误差低于 0.8 s。结果表明,所提方法能够有效捕捉连续行为模式,提高识别准确性和时间定位精度,为智能安全监测提供可靠方案。此外,该系统支持基于事件级预警的主动干预,有助于提升矿山作业的安全性与效率。

 

Event-level abnormal driving behavior detection in mining trucks based on temporal-aware vision-language models

Abstract: Unsafe driving behaviors significantly contribute to accidents in open-pit mining operations. Conventional monitoring systems often fail to capture the temporal continuity and semantic ambiguity of such behaviors. To address these challenges, an event-level abnormal driving behavior recognition method based on a vision-language model (VLM) is proposed. The method integrates Temporal-Aware Low-Rank Adaptation (T-LoRA) with event-level supervised fine-tuning to enable efficient vertical-domain customization of the VLM toward complex open-pit operating environments. By explicitly enhancing temporal modeling of visual sequences, the proposed method enables accurate event-level recognition of abnormal driving behaviors, precise localization of their onset and duration, and the generation of interpretable semantic descriptions. Experiments conducted on a real-world mining truck dataset demonstrate that the proposed approach achieves an event-level F1-score of 0.923, while maintaining a mean absolute error for temporal localization below 0.8 s. These results indicate that the proposed method effectively captures continuous behavioral patterns and improves both recognition accuracy and temporal precision, providing a reliable solution for intelligent safety monitoring. Furthermore, the system supports proactive intervention through event-level early warnings, contributing to safer and more efficient mining operations.

 

/

返回文章
返回