Machine learning–based classification of cancer types using genomic profiling data from the Australian Molecular Screening and Therapeutics (MoST) program.
Abstract
e13683 Background: Mutational patterns offer diagnostic value for cancer type determination, particularly in patients with cancer of unknown primary site (CUP) and synchronous metastases from multiple primaries. Despite increasing adoption of comprehensive genomic profiling (CGP), systematic evaluation of mutational pattern analysis as a standalone diagnostic classifier for cancer type determination remains necessary. Methods: From the MoST pan-cancer program (ACTRN12616000908437) through December 2023, we developed machine learning (ML) cancer type classifiers incorporating age, sex, and CGP data (including TSO500 or FoundationOne CDx panels), pathogenic variants, tumour mutational burden and microsatellite status. Variants were stratified by type and functional impact as features. L2-regularised logistic regression models estimated posterior probabilities in binary classification, implementing a 1-vs-rest strategy across cancer types and hierarchical levels with n≥10 cases. Stratified 10-fold cross-validation was used, with area under the receiver operating characteristic curve (AUC) as the primary metric. Concordance between model predictions and pathology review was assessed in the MoST CUP cohort, where cases with posterior probability or likelihood ratio (LR) methods (threshold > 2.0) generated differential diagnoses and were compared against a pathologist's review. Results: The cohort consisted of 4,990 patients with solid tumours, from which 209 cancer types and hierarchical subtype models were developed. Median AUC for classification across all cancer types was 0.857 (bootstrapped 95% CI: 0.844-0.875), with 70 cancer types (33%) achieving AUC > 0.9. Models demonstrated robust performance for major cancer types: breast (AUC 0.959, 95% CI: 0.922-0.965), colorectal (0.967, 95% CI: 0.946-0.967), prostate (0.953, 95% CI: 0.924-0.964), pancreatic (0.926, 95% CI: 0.905-0.948), gynaecologic (0.884, 95% CI: 0.877-0.890), and non-small-cell lung (0.878, 95% CI: 0.804-0.899) cancers. Analysis of the CUP cohort (n = 153) revealed that model-generated differential diagnoses showed concordance with pathologist assessment in 129 cases (84.3%, 95% CI: 77.6-89.7) for ≥1 broad cancer category and 104 cases (68.0%, 95% CI: 60.0-75.3) for specific diagnostic classifications. The LR method showed concordance in 97 cases (73.9%, 95% CI: 66.1-80.6) for broad categories and 113 cases (63.4%, 95% CI: 55.2-71.0) for specific classifications. Conclusions: Cancer type classification using CGP demonstrated high discriminative performance in selected tumour types. The concordance study suggests that leveraging mutational patterns through ML could provide information beyond pathognomonic alterations, supplementing multidisciplinary assessment in diagnostically challenging cases, particularly CUPs.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (9)
Frank Po-Yen Lin
Garvan Institute of Medical Research, Sydney, NSW, Australia
Min Li Huang
Garvan Institute of Medical Research, Sydney, NSW, Australia
John P. Grady
Centre for Molecular Oncology, University of New South Wales, Sydney, NSW, Australia
Subotheni Thavaneswaran
The Kinghorn Cancer Centre, St Vincent's Hospital, Darlinghurst, NSW, Australia
Maya Kansara
Centre for Molecular Oncology, University of New South Wales, Sydney, NSW, Australia
Christine Napier
Omico, Kensington, Australia
Mandy L Ballinger
Omico, Sydney, NSW, Australia
John Simes
NHMRC Clinical Trials Centre, University of Sydney, NSW, Australia (J.S.).
David Morgan Thomas
Centre for Molecular Oncology, University of New South Wales, Sydney, NSW, Australia