Automated discovery and curation of public spatial transcriptomic datasets across multiple cancer indications.
Abstract
e15010 Background: Public spatial transcriptomic datasets provide critical insights into tumor architecture and the tumor microenvironment, which is central to target discovery, biomarker identification and patient stratification. However, their integration for oncology research is limited by fragmentation across repositories, platforms, and inconsistent metadata and molecular data across studies. Scalable approaches are needed to systematically identify and organize these datasets across cancer indications. Methods: The Rancho Data Crawler, an automated discovery tool, was used to systematically identify publicly available spatial transcriptomic studies across seven thoracic and gastrointestinal oncologic indications. Searches were conducted across GEO, EGA, dbGaP, ArrayExpress, SRA, and Zenodo. Initial discovery focused on 10x Genomics Visium and Xenium platforms, with NanoString GeoMx and CosMx technologies incorporated to expand coverage. Candidate studies underwent manual review, followed by harmonized study-, donor-, and sample-level metadata curation and quality control. Raw and author-processed data were downloaded and delivered using a standardized cloud-based structure. Results: Automated crawling enabled large-scale, systematic identification of spatial transcriptomic studies across the targeted cancer indications, from which 25 studies were prioritized for detailed curation. The curated dataset was enriched for thoracic and aerodigestive tissues, with lung and head-and-neck samples comprising the largest tissue groups, alongside substantial representation of liver, pancreas, colorectum, and stomach. Samples were predominantly derived from malignant tumors, with a notable proportion representing advanced and metastatic disease. Tumor-proximal and immune-relevant tissues, including lymph node and peritoneal samples, were also captured. Data delivery comprised 27 datasets totaling approximately 4 TB and more than 8,500 files, organized with manifest files to support traceability and data integrity. Conclusions: Automated discovery using the Rancho Data Crawler, combined with rigorous metadata harmonization, enables scalable access to spatial transcriptomic datasets enriched for clinically and translationally relevant tumor tissues and disease states, supporting downstream translational oncology research.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Amrita Bhattacharya
Department of Metallurgical Engineering and Materials Science, Indian Institute of Technology Bombay 1 , Mumbai 400076,
Corey Oravetz
Rancho Biosciences, Rancho, Santa Fe, CA
Anne Cooley
Rancho Biosciences, Rancho Santa Fe, CA
Dzmitry Fedarovich
Rancho BioSciences, San Diego, CA
Kenneth Chan
RanchoBioSciences, Rancho Santa Fe, CA
Kai Ding
Emre Arslan
Department of Genomic Medicine and MDACC Epigenomics Therapy Initiative, The University of Texas MD Anderson Cancer Center
Somnath Bandyopadhyay
Takeda Pharmaceutical Company Limited, Cambridge, MA