Automated discovery and curation of public spatial transcriptomic datasets across multiple cancer indications.

A Amrita Bhattacharya (Department of Metallurgical Engineering and Materials Science, Indian Institute of Technology Bombay 1 , Mumbai 400076,) C Corey Oravetz (Rancho Biosciences, Rancho, Santa Fe, CA) A Anne Cooley (Rancho Biosciences, Rancho Santa Fe, CA) D Dzmitry Fedarovich (Rancho BioSciences, San Diego, CA) K Kenneth Chan (RanchoBioSciences, Rancho Santa Fe, CA) K Kai Ding E Emre Arslan (Department of Genomic Medicine and MDACC Epigenomics Therapy Initiative, The University of Texas MD Anderson Cancer Center) S Somnath Bandyopadhyay (Takeda Pharmaceutical Company Limited, Cambridge, MA)

Abstract

e15010 Background: Public spatial transcriptomic datasets provide critical insights into tumor architecture and the tumor microenvironment, which is central to target discovery, biomarker identification and patient stratification. However, their integration for oncology research is limited by fragmentation across repositories, platforms, and inconsistent metadata and molecular data across studies. Scalable approaches are needed to systematically identify and organize these datasets across cancer indications. Methods: The Rancho Data Crawler, an automated discovery tool, was used to systematically identify publicly available spatial transcriptomic studies across seven thoracic and gastrointestinal oncologic indications. Searches were conducted across GEO, EGA, dbGaP, ArrayExpress, SRA, and Zenodo. Initial discovery focused on 10x Genomics Visium and Xenium platforms, with NanoString GeoMx and CosMx technologies incorporated to expand coverage. Candidate studies underwent manual review, followed by harmonized study-, donor-, and sample-level metadata curation and quality control. Raw and author-processed data were downloaded and delivered using a standardized cloud-based structure. Results: Automated crawling enabled large-scale, systematic identification of spatial transcriptomic studies across the targeted cancer indications, from which 25 studies were prioritized for detailed curation. The curated dataset was enriched for thoracic and aerodigestive tissues, with lung and head-and-neck samples comprising the largest tissue groups, alongside substantial representation of liver, pancreas, colorectum, and stomach. Samples were predominantly derived from malignant tumors, with a notable proportion representing advanced and metastatic disease. Tumor-proximal and immune-relevant tissues, including lymph node and peritoneal samples, were also captured. Data delivery comprised 27 datasets totaling approximately 4 TB and more than 8,500 files, organized with manifest files to support traceability and data integrity. Conclusions: Automated discovery using the Rancho Data Crawler, combined with rigorous metadata harmonization, enables scalable access to spatial transcriptomic datasets enriched for clinically and translationally relevant tumor tissues and disease states, supporting downstream translational oncology research.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (8)

A

Amrita Bhattacharya

Department of Metallurgical Engineering and Materials Science, Indian Institute of Technology Bombay 1 , Mumbai 400076,

C

Corey Oravetz

Rancho Biosciences, Rancho, Santa Fe, CA

A

Anne Cooley

Rancho Biosciences, Rancho Santa Fe, CA

D

Dzmitry Fedarovich

Rancho BioSciences, San Diego, CA

K

Kenneth Chan

RanchoBioSciences, Rancho Santa Fe, CA

K

Kai Ding

E

Emre Arslan

Department of Genomic Medicine and MDACC Epigenomics Therapy Initiative, The University of Texas MD Anderson Cancer Center

S

Somnath Bandyopadhyay

Takeda Pharmaceutical Company Limited, Cambridge, MA