Machine learning-based detection of multiple cancers using 5-hydroxymethylcytosine at cancer mutation hotspot as marker.

J Julia Lu (Anchor Molecular Inc, Pleasanton, CA) Y Yabin Lu (Anchor Molecular Inc, Pleasanton, CA)

Abstract

806 Background: 5-hydroxymethylcytosine (5hmC) is an important DNA epigenetic modification that has been linked to gene regulation and cancer pathogenesis. Several recent studies using Machine Learning have evaluated and demonstrated the potential usefulness of 5hmC as cancer marker. To simply the massive amount of NGS data, these studies used combined or averaged 5hmC signals across long chromosome regions (>2kbs). This strategy is useful to look at the overall picture of 5hmC patterns. But it may mask localized specific 5hmC markers with changes that are easier to quantify or to model. In this study, we used Machine Learning to evaluate diagnostic potential of individual 5hmC located at specific cancer hotspots. Methods: Data of 5hmC purified from Cell-free DNA from 12 healthy patients and 27 cancer patients (15 lung cancer, 5 gastric cancer, and 7 pancreatic cancer) were used in the study. Two or three of each group were used as testing group and the rest as training. Based on frequency of detection of 5hmC, around 1200 most frequent 5hmC cancer mutation hotspots with either (C) or (G) were selected for the study. Classic machine learning approaches applied in this research/study included logistic regression, random forest, support vector machine(SVM), and XGboost models. These models were evaluated for two primary objectives. (1) binary classification of healthy versus cancer samples, and (2) multi-class classification of cancer types. The model performance was assessed using four standard evaluation matrics: precision, recall (sensitivity), F1-score and support. The SHAP feature importance ranking plot is generated. Deep learning typically requires large datasets, therefore, we implemented a feedforward neural network (FNN) as a representative model. Results: Due to the limited size of our datasets, the classic machine learning models generally outperformed the deep learning FNN model. For binary classification, the accuracies of models ranged from 80% to 100%. For the multi-class classification, the SVM model reached 100% accuracy, while the XGboost model has 80% accuracy. Notably, the SHAP feature importance analysis highlighted chromosome 11 position 77936165 as a strong feature. Model optimization with more data is on-going. The detailed comparison on the model performance and discussion will be presented during conference. Conclusions: The result indicates that with machine learning model based on only around 1200 hotspot, 5hmC signature can be used to detect gastric, lung and pancreatic cancer from healthy patients.

Article Details

Volume / Issue Vol. 44, Issue 2_suppl
Published January 10, 2026
Pages 806-806
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (2)

J

Julia Lu

Anchor Molecular Inc, Pleasanton, CA

Y

Yabin Lu

Anchor Molecular Inc, Pleasanton, CA