Adaptive cognitive driven cross modal network for few shot fine grained recognition
Abstract
Abstract Current few-shot learning (FSL) methods struggle with fine-grained texture loss, inefficient cross-modal knowledge integration, and catastrophic forgetting. To resolve these bottlenecks, we propose the Adaptive cognitive driven cross modal network (ACD-Net) for few-shot fine-grained recognition. Inspired by human cognition, ACD-Net introduces three systematic innovations. First, the Adaptive Dual-Domain Cognitive Attention (ADCA) module employs Two-Dimensional Discrete Wavelet Transform and Gated Recurrent Units to decouple high-frequency textures from noise and dynamically localize discriminative regions. Second, the Graph-Guided Semantic-to-Visual Distillation (GSD) strategy utilizes Graph Convolutional Networks and a bilinear attention mechanism to seamlessly embed structured semantic priors into the visual space, generating exceptionally robust category prototypes. Finally, the Dynamic Balanced Anti-forgetting (DBAF) loss function mitigates catastrophic forgetting during fine-tuning by adaptively adjusting regularization weights based on gradient orthogonality between old and novel tasks. Extensive evaluations on the miniImageNet, CUB-200-2011, and Medical-44 datasets demonstrate that ACD-Net achieves state-of-the-art results, elevating average accuracy by 1.75% and 1.36% in 1-shot and 5-shot scenarios, respectively. Ultimately, ACD-Net establishes an innovative paradigm for FSL, offering pragmatic solutions for complex real-world deployments including industrial defect inspection and intelligent clinical diagnosis.
Article Details
Authors (2)
Xin He
Zeshi Wu