Void-X: A generative void-filling model for predicting atomic packing in proteins

J Jing Yang J Junying Yuan J James J. Chou

Abstract

Generative AI algorithms such as the transformer and diffusion models have greatly empowered de novo design of proteins capable of specifically interacting with designated structural sites on another protein. Most of these design methods employ a top–down approach, in which an overall protein shape is generated by an AI model to pack against a given structural site, followed by sequence design to optimize the interaction. Despite being trained on limited protein complex structures available in the database, the top–down approach has yielded encouraging results. Here, we propose a bottom–up approach that generates atom clusters for optimal packing against a specified structured region for informing the design of protein–protein interactions. To this end, we trained a masked discrete diffusion model, named Void-X, that uses the diffusion transformer to learn atomic-level interactions and fill atomic voids in protein interaction interfaces. Void-X was trained using 8.7 million spherical clusters of atoms from experimental structures in the Protein Data Bank. In each cluster, ~70% of the atoms are used as context (or prompt), and ~30% are masked for information recovery (or answer). By training the model with 172 million parameters, Void-X achieves an overall accuracy of 78.3% and 68.2% for intra- and interchain spherical clusters, respectively. Furthermore, we find that information entropy is a reliable indicator of the prediction accuracy for Void-X. This level of performance allows de novo generation of molecular interactions at the atomic level, offering an alternative approach of protein design complementary to the existing ones.

Article Details

Volume / Issue Vol. 123, Issue 24
Published June 16, 2026
ISSN 0027-8424
Publisher National Academy of Sciences

Authors (3)

J

Jing Yang

J

Junying Yuan

J

James J. Chou