Discovery of a phenazine–thiol conjugase from sparse data using genome-informed machine learning

X Xiaoyu Shan (Division of Biology and Biological Engineering, California Institute of Technology) I Inês B. Trindade (Division of Biology and Biological Engineering, California Institute of Technology) N Nathaniel R. Glasser (Resnick Sustainability Institute, California Institute of Technology) K Korbinian O. Thalhammer (Division of Geological and Planetary Sciences, California Institute of Technology) M Matthew Scurria (Department of Chemistry and Biochemistry, University of California) A Ariane Mora (AITHYRA GmbH, Research Institute for Biomedical Artificial Intelligence of the Austrian Academy of Sciences) S Stuart J. Conway (Department of Chemistry, Chemistry Research Laboratory, University of Oxford, Mansfield Road, Oxford OX1 3TA, U.K.) D Dianne K. Newman (Division of Biology and Biological Engineering, California Institute of Technology)

Abstract

Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (phenazine–thiol conjugase), an enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through nonenzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.

Article Details

Volume / Issue Vol. 123, Issue 31
Published August 04, 2026
ISSN 0027-8424
Publisher National Academy of Sciences

Authors (8)

X

Xiaoyu Shan

Division of Biology and Biological Engineering, California Institute of Technology

I

Inês B. Trindade

Division of Biology and Biological Engineering, California Institute of Technology

N

Nathaniel R. Glasser

Resnick Sustainability Institute, California Institute of Technology

K

Korbinian O. Thalhammer

Division of Geological and Planetary Sciences, California Institute of Technology

M

Matthew Scurria

Department of Chemistry and Biochemistry, University of California

A

Ariane Mora

AITHYRA GmbH, Research Institute for Biomedical Artificial Intelligence of the Austrian Academy of Sciences

S

Stuart J. Conway

Department of Chemistry, Chemistry Research Laboratory, University of Oxford, Mansfield Road, Oxford OX1 3TA, U.K.

D

Dianne K. Newman

Division of Biology and Biological Engineering, California Institute of Technology