Pichia-CLM: A language model–based codon optimization pipeline for <i>Komagataella phaffii</i>
Abstract
The preference in synonymous codon usage—the so-called codon usage bias (CUB)—is governed by several factors such as the host organism, context and function of the gene, and the position of the codon within the gene itself. We demonstrated that this mapping can be learned from the host’s genome using language models and subsequently applied for codon optimization of heterologous proteins expressed by the host. This pipeline called Pichia–Codon language model (Pichia-CLM) was applied to the industrial host organism, Komagataella phaffii. With this approach, production of heterologous proteins was enhanced up to threefold compared to their native sequences. Furthermore, Pichia-CLM consistently yielded constructs with enhanced productivity for proteins of varied complexity, compared to commercially available tools. Finally, we showed that Pichia-CLM generates sequences resembling the properties of codon usage found in the host’s intrinsic host cell proteins and learned features such as avoiding negative cis-regulatory and repeat elements based on patterns in the genome data. These results show the potential of language models to unbiasedly learn patterns and design robust sequences for improved protein production.
Article Details
Journal Info
Proceedings of the National Academy of Sciences
National Academy of Sciences
Authors (2)
Harini Narayanan
Koch Institute for Integrative Cancer Research, Massachusetts Institute of Technology
J. Christopher Love
Koch Institute for Integrative Cancer Research, Massachusetts Institute of Technology