Huang, C., Hu, Z. (2025). A multimodal transformer-based visual question answering method integrating local and global information. PLoS ONE. https://doi.org/10.1371/journal.pone.0324757