张立成,王蕾,宋逸菲,谢耀华,周竹萍,陈春成.基于多源扩散特征融合的行人目标检测算法研究[J].电子测量与仪器学报,2026,40(6):49-63
基于多源扩散特征融合的行人目标检测算法研究
Research on pedestrian target detection algorithmbased on multi-source diffusion feature fusion
  
DOI:
中文关键词:  行人检测  图像融合  扩散模型  红外与可见光  YOLOv8
英文关键词:pedestrian detection  image fusion  diffusion model  infrared and visible  YOLOv8
基金项目:国家自然科学基金(52472314)、陕西省自然科学基金(2025JC-YBMS-408)、浙江省交通运输厅重大交通建设工程科研项目(2023-GCKY-09)资助
作者单位
张立成 长安大学信息工程学院西安710064 
王蕾 空军工程大学空管领航学院西安710051 
宋逸菲 中航电测仪器(西安)有限公司西安710119 
谢耀华 招商局重庆交通科研设计院有限公司重庆400067 
周竹萍 南京理工大学自动化学院南京210094 
陈春成 长安大学信息工程学院西安710064 
AuthorInstitution
Zhang Licheng School of Information Engineering,Chang′an University, Xi′an 710064, China 
Wang Lei School of Air Traffic Control and Navigation, Air Force Engineering University, Xi′an 710051, China 
Song Yifei Zhonghang Electronic Measuring Instruments Co., Ltd., Xi′an 710119, China 
Xie Yaohua China Merchants Chongqing Communications Technology Research & Design Institute Co., Ltd., Chongqing 400067, China 
Zhou Zhuping School of Automation, Nanjing University of Science and Technology, Nanjing 210094, China 
Chen Chuncheng School of Information Engineering,Chang′an University, Xi′an 710064, China 
摘要点击次数: 130
全文下载次数: 17
中文摘要:
      复杂环境下的行人检测是自动驾驶场景面临的重要技术挑战。提出一种基于红外与可见光多模态特征融合的行人目标检测方法。首先,针对跨模态特征融合难题,设计了一种基于扩散模型的红外与可见光图像融合算法,通过可逆全局捕获模块和轻量局部捕获模块实现多模态特征的有效提取与融合,提升融合图像的细节保真度和结构一致性。其次,针对低照度场景检测需求,提出了一种改进的YOLOv8n检测网络,通过引入MobileNetV4主干网络和颈部网络进行轻量化改进,并融入注意力机制和损失函数,增强关键特征提取能力和网络收敛性能。实验结果表明,所提方法在LLVIP和MSRS数据集上的融合图像质量指标显著优于现有方法,检测精度达到95.3%,较基础模型提升1.7%,模型参数量减少42.6%,计算效率提升17.4%。系统在复杂环境下的具备鲁棒性和实时性,为自动驾驶和智能安防应用提供了有效的技术解决方案。
英文摘要:
      Pedestrian detection in complex environments presents a critical technical challenge for autonomous driving systems. This paper proposes a pedestrian target detection method based on multimodal feature fusion of infrared and visible light images. First, to address the challenge of cross-modal feature fusion, we propose an infrared-visible image fusion algorithm (DIVIF) based on a multi-source diffusion model. The algorithm incorporates a reversible feature decoupling module (IGCM) to capture target information distributions across modalities and extract large-scale structural features, along with a lightweight local capture module (LLCM) to enhance detail fidelity through attention-driven dynamic interaction. A global-local iterative optimization mechanism is introduced to achieve effective feature transfer and deep fusion. Second, for low-illumination scenarios, we propose an improved YOLOv8n detection network with lightweight architecture and attention mechanisms. The model integrates MobileNetV4 and GSConv into the backbone and neck networks for efficiency, while incorporating EMA attention to strengthen key feature extraction and MPDIoU loss to optimize convergence. Furthermore, a complete pedestrian detection system implementation is designed for practical applications. Experimental results demonstrate that the proposed method effectively fuses multimodal features, achieving precise detection in low-light and occluded scenarios. Compared to mainstream detection networks, our algorithm outperforms in all evaluation metrics.
查看全文  查看/发表评论  下载PDF阅读器