Abstract:Addressing the challenges of large parameter sizes, slow inference speeds, and the difficulty in balancing model lightweighting with detection performance in existing steel surface defect detection models, this paper proposes a UAL-YOLO model based on YOLOv11n. Firstly, a parameter-free unified 3D attention module (UAM) is integrated into the backbone. By deriving 3D weights from second-order statistics, it identifies outlier defect features via linear separability, significantly enhancing the signal-to-noise ratio of semantic information. Secondly, a novel attention-guided fusion (AGF) neck network is designed, which utilizes deep features to generate guiding weights, intelligently and selectively aggregating multi-scale features to improve the efficiency and accuracy of feature fusion. Finally, a lightweight efficient offset-based up sampler (LEO) is developed using content-aware offsets to recover fine textures of slender cracks and minute pits. This approach improves the extraction of low-contrast edges and effectively minimizes the missed detection rate of tiny defects in industrial scenarios. Experimental results demonstrate that compared to the baseline YOLOv11 model, the UAL-YOLO model achieves a 2.1% improvement in mAP@0.5 and a 3.7% increase in recall rate on the NEU-DET dataset, while reducing the number of parameters, computational load, and storage capacity by 26.36%, 15.87%, and 23.64%, respectively. The UAL-YOLO model significantly enhances precision while achieving notable lightweighting, providing a new solution for the deployment of high-precision, low-cost steel defect detection technologies in industrial scenarios.