Dual-stage framework with soft-label distillation and spatial prompting for image-text retrieval

R Ran Jin Z Zhengang Li F Fang Deng Y Yanhong Zhang M Min Luo (College of Life Sciences, Anhui Normal University) T Tao Jin (Department of Chemistry, University of Basel, St. Johanns-Ring 19, Basel 4056, Switzerland) T Tengda Hou C Chenjie Du X Xiaozhe Gu J Jie Yuan

Abstract

Vision-language pre-training (VLP) methods have significantly advanced cross-modal tasks in recent years. However, image-text retrieval still faces two critical challenges: inter-modal matching deficiency and intra-modal fine-grained localization deficiency. These issues significantly impede the accuracy of image-text retrieval. To address these challenges, we propose a novel dual-stage training framework. In the first stage, we employ Soft Label Distillation (SLD) to align the contrastive relationships between images and texts by mitigating the overfitting problem caused by hard labels. In the second stage, we introduce Spatial Text Prompt (STP) to enhance the model’s visual grounding capabilities by incorporating spatial prompt information, thereby achieving more precise fine-grained alignment. Extensive experiments on standard datasets show that our method outperforms state-of-the-art approaches in image-text retrieval.The code and supplementary files can be found at https://github.com/Leon001211/DSSLP.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 10
Published October 10, 2025
Pages e0333084
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (10)

R

Ran Jin

Z

Zhengang Li

F

Fang Deng

Y

Yanhong Zhang

M

Min Luo

College of Life Sciences, Anhui Normal University

T

Tao Jin

Department of Chemistry, University of Basel, St. Johanns-Ring 19, Basel 4056, Switzerland

T

Tengda Hou

C

Chenjie Du

X

Xiaozhe Gu

J

Jie Yuan