A multi-modal co-attention model for accurate drug-target interaction prediction

W Wanjun Ma W Wenjun Li M Mengyun Yang Z Zhengdong Pu X Xiwei Tang

Abstract

Accurate prediction of drug-target interactions (DTIs) plays a crucial role in modern drug discovery and repositioning. Despite recent advances in deep learning, existing methods often fail to effectively integrate heterogeneous data, such as molecular structures and protein sequences, into a unified representation. To address this limitation, we propose MMCA (Multi-Modal Co-Attention), a novel deep learning framework that introduces a multi-modal co-attention mechanism to dynamically align and fuse graph-based drug features with sequence-based protein embeddings. Our model leverages parallel encoding pathways to capture both structural and semantic information, followed by a context-aware fusion module that adaptively weighs cross-modal dependencies. Evaluation on three benchmark datasets—BioSNAP, BindingDB, and Human STRING—demonstrates that MMCA outperforms state-of-the-art methods in terms of AUC, AUPR, and F1-score, achieving up to 98.4% AUC. Ablation studies confirm the significance of our co-attention fusion mechanism in enhancing both accuracy and robustness. Case studies of high-confidence predictions reveal biologically plausible drug-protein interactions, supporting MMCA’s potential for prioritizing candidates for experimental validation. By enabling end-to-end multi-modal reasoning, MMCA provides a powerful framework for advancing DTI prediction systems and offers broad applicability for various bioinformatics tasks. The source code of MMCA is publicly available at https://github.com/Join-xiaobai/MMCA .

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 6
Published June 22, 2026
Pages e0351880
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (5)

W

Wanjun Ma

W

Wenjun Li

M

Mengyun Yang

Z

Zhengdong Pu

X

Xiwei Tang