An automated clustering workflow for molecular simulation data
Abstract
Efficient analysis methods are needed that are able to cope with the huge size of modern molecular simulation datasets. Typical analysis workflows involve dimensionality reduction algorithms that improve the comprehensibility of the dataset by lowering its dimensions. Other common elements of the analysis are clustering algorithms that divide the dataset into groups according to a predefined characteristic, often with an emphasis on structural homogeneity. The application of such a clustering algorithm can help identify important metastable states of the system. In this paper, we revisit a clustering workflow described by Hunkler et al. [J. Chem. Phys. 158, 144109 (2023)] and provide a refined and automated version. We apply this workflow to a dataset of atomistic simulations of the 76 amino-acid residue protein ubiquitin (Ub) to illustrate its strengths and characteristics. The automated clustering workflow combines two dimensionality reduction algorithms, cc_analysis and EncoderMap, with the clustering algorithm HDBSCAN and a root mean square deviation-based sorting criterion. Due to an iterative approach, it is especially suitable for highly efficient categorization of large datasets of structures into structurally homogeneous clusters of different densities and sizes. Small adaptations to the original clustering workflow ensure considerable improvements in the quality and homogeneity of the obtained clusters.
Article Details
Journal Info
The Journal of Chemical Physics
American Institute of Physics
Authors (3)
Elena Paulus
Department of Chemistry, University of Konstanz 1 , Konstanz,
Madlen Malcharek
Department of Chemistry, University of Konstanz 1 , Konstanz,
Christine Peter
Department of Chemistry, University of Konstanz 1 , Konstanz,