吴文欢,陈杰,王文舒.DPM-YOLO:面向自动驾驶复杂场景的多尺度自适应目标检测算法[J].电子测量与仪器学报,2026,40(5):296-307
DPM-YOLO:面向自动驾驶复杂场景的多尺度自适应目标检测算法
DPM-YOLO: A multi-scale adaptive object detection algorithm forcomplex scenarios in autonomous driving
  
DOI:
中文关键词:  机器视觉  YOLOv11n  动态注意力金字塔  并行融合卷积  MPDWIoU损失函数
英文关键词:machine vision  YOLOv11n  dynamic attention pyramid  parallel fusion convolution  MPDWIoU loss
基金项目:湖北省自然科学基金创新发展联合基金项目(2025AFD239 )、汽车动力传动与电子控制湖北省重点实验室开放基金项目(ZDK12025A02)资助
作者单位
吴文欢 1.湖北汽车工业学院智能网联汽车学院十堰442000;2. 湖北汽车工业学院汽车动力传动与 电子控制湖北省重点实验室十堰442000 
陈杰 湖北汽车工业学院智能网联汽车学院十堰442000 
王文舒 湖北汽车工业学院智能网联汽车学院十堰442000 
AuthorInstitution
Wu Wenhuan 1.School of Intelligent Connected Vehicle, Hubei University of Automotive Technology, Shiyan 442000, China; 2.Key Laboratory of Automotive Power Train and Electronics, Hubei University of Automotive Technology, Shiyan 442000, China 
Chen Jie School of Intelligent Connected Vehicle, Hubei University of Automotive Technology, Shiyan 442000, China 
Wang Wenshu School of Intelligent Connected Vehicle, Hubei University of Automotive Technology, Shiyan 442000, China 
摘要点击次数: 232
全文下载次数: 89
中文摘要:
      目标检测是自动驾驶环境感知的核心技术,其任务是识别并定位车辆周围环境中的关键物体。针对自动驾驶场景中小目标漏检、遮挡误判及动态适应性不足的挑战,提出基于YOLOv11n的算法模型DPM-YOLO。首先,在骨干网络(Backbone)中,构建动态注意力金字塔模块(DAP)替换原快速空间金字塔池化模块(SPPF),利用参数共享空洞卷积与分解式空间注意力高效地捕获不同粒度特征,从而更好地聚合目标的上下文信息;其次,在颈部网络(Neck)中,设计并行融合卷积模块(PFConv),采用异构多分支卷积动态解耦通道并融合多尺度特征;最后,提出一种改进的MPDWIoU损失函数,融合动态梯度聚焦与角点约束,结合批次统计特征自适应策略优化定位精度。实验结果表明,DPM-YOLO在KITTI数据集上的mAP@0.5达92.7%,较基准模型YOLOv11n提升2.0个百分点,在Cityscapes和BDD100K上的mAP@0.5较YOLOv11n分别提升了1.6和0.8个百分点,在极端天气数据集Foggy-Cityscapes上的mAP@0.5较YOLOv11n提升了1.2个百分点。本研究方法较好地实现了对复杂场景检测的精度和实时性平衡。
英文摘要:
      Object detection is a core technology in the environmental perception of autonomous driving, and its task is to identify and locate key objects in the surrounding environment of the vehicle. To address the challenges of missed detection of small targets, misjudgment due to occlusion, and insufficient dynamic adaptability in autonomous driving scenarios, an improved object detection algorithm DPM-YOLO based on YOLOv11n is proposed. Firstly, in the Backbone network, a dynamic attention pyramid module is constructed to replace the original rapid spatial pyramid pooling structure. By utilizing parameter-shared dilated convolutions and a decomposed spatial attention mechanism, multi-granularity features are efficiently captured and the context information of each target can be better aggregated. Secondly, in the Neck network, a parallel fusion convolution module is designed. It can achieve channel-wise dynamic decoupling and multi-scale features fusion by heterogeneous multi-branch convolutions. Finally, an improved MPDWIoU loss function is proposed, which integrates a dynamic gradient focusing mechanism with corner constraints and optimizes bounding box localization accuracy by combining a batch statistics adaptive strategy. Experimental results demonstrate that the mAP@0.5 of DPM-YOLO on the KITTI dataset is 92.7%, which is 2.0 percentage points higher than that of YOLOv11n. The mAP@0.5 of DPM-YOLO on the Cityscapes and BDD100K datasets is improved by 1.6 and 0.8 percentage points respectively compared with YOLOv11n. The mAP@0.5 of DPM-YOLO on the extreme weather dataset Foggy-Cityscapes increases by 1.2 percentage points over YOLOv11n. The results indicate that the proposed method effectively balances detection accuracy and real-time performance in complex driving scenarios.
查看全文  查看/发表评论  下载PDF阅读器