Mining lysine post-translational modification sites by integrating protein language model representations with structural context
Abstract
Lysine (Lys/K) residues serve as major hubs for post-translational modifications (PTMs) owing to the chemical versatility of their ε-amino groups, giving rise to diverse regulatory functions. Accurate and efficient identification of modified lysine residues therefore requires computational models that can effectively capture both sequence and structural information while minimizing domain-specific feature engineering. In this study, we propose a unified deep learning framework for lysine PTM site identification that integrates sequence representations derived from a protein language model with atom-level three-dimensional structural features. This framework can be consistently applied to multiple lysine PTM types using a shared modeling strategy. As an application, we used the model to predict potential PTM site on human C-type lectin domain family 12 member A (hCLEC12A) and evaluated their functional relevance through all-atom molecular dynamics simulations. The simulations indicate that the predicted lysine residues influence the stability and binding behavior of the hCLEC12A-antibody 50C1 complex. Overall, this work presents an integrative computational framework for lysine PTM site mining and functional analysis.
Article Details
Journal Info
Proceedings of the National Academy of Sciences
National Academy of Sciences
Authors (5)
Mengqi Luo
Key Laboratory of Systems Health Science of Zhejiang Province, School of Life Science, Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences
Xiaohong Zhu
Chen Bai
Warshel Institute for Computational Biology, School of Life and Health Sciences, School of Medicine, The Chinese University of Hong Kong (Shenzhen), Shenzhen 518172, China
Arieh Warshel
Department of Chemistry
Luonan Chen