Contrastive learning with mutual information enhancement and negative sample augmentation combined with KAN for text clustering
Abstract
This paper proposes a text clustering model based on mutual information-enhanced contrastive learning combined with Kolmogorov-Arnold Networks (KAN) to address challenges in text clustering, including insufficient robustness of text representations, feature redundancy, and the curse of dimensionality. Text clustering is essential for organizing unstructured textual data, yet existing methods often suffer from weak feature representations and limited scalability in high-dimensional spaces. The proposed model jointly addresses these three core issues by maximizing mutual information to capture nonlinear relationships among positive samples, expanding negative samples to enhance discriminability, and employing the KAN architecture to adapt to the complex structures of high-dimensional data. Experimental results on eight benchmark datasets demonstrate that the proposed model achieves state-of-the-art accuracy on seven out of eight datasets and leads normalized mutual information on six datasets, outperforming existing text clustering baselines. To validate its practical utility, the model was applied to cluster Weibo posts collected via web crawlers using “technology” as the keyword during January and February 2025. The case study results reveal that the model effectively identifies meaningful thematic clusters—such as AI applications, technology industry trends, and consumer electronics discussions—confirming its applicability to real-world social media data.
Article Details
Authors (6)
Yuanmin Zhang
Hao Li
Chunzhi Xie
Yong Huang
National Laboratory of Solid State Microstructures, School of Physics
Zhenyi Wu
Yanjun Li
Hefei National Laboratory for Physical Sciences at Microscale and Department of Physics