DEF-Net: A dual-modal feature enhancement and fusion network for infrared and visible object detection

X Xiaoming Guo F Fengbao Yang L Linna Ji

Abstract

Infrared-visible object detection in complex dynamic environments often suffers from weak feature representation and underutilized cross-modal complementarity, leading to missed and false detections. To address these issues, we propose a Dual-modal Enhanced Feature Enhancement and Fusion Network (DEF-Net). To enhance the model’s focus on informative features within both infrared and visible modalities, a feature interaction enhancement module is designed to effectively highlight and reinforce salient information. Furthermore, to better exploit the complementary characteristics of the two modalities, a transformer-based fusion architecture incorporating a cross-attention mechanism is introduced, enabling deep inter-modal feature integration. Experiments on SYUGV and LLVIP datasets show that DEF-Net outperforms existing methods in accuracy while maintaining real-time processing speed.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 4
Published April 01, 2026
Pages e0345815
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (3)

X

Xiaoming Guo

F

Fengbao Yang

L

Linna Ji