Research Internship
Indian Institute of Technology, Roorkee (IIT Roorkee)
Exploring Breast Cancer Transcriptomics Across Global and Indian Cohorts: A Comparative Study
Authors: Priyal Tripathi (Supervisor: Dr. Aparajita Khan)
Comparative transcriptomic analysis across TCGA (n=815) and Indian cohorts (n=81), identifying CFH, DST, and COPZ2 as candidate universal biomarkers across subtypes.
Project Overview
Undergraduate Research Internship conducted at the Indian Institute of Technology, Roorkee (IIT Roorkee) under the supervision of Dr. Aparajita Khan.
Most public cancer transcriptomic datasets represent predominantly Western demographic populations. In this study, we evaluated whether gene expression landscapes and candidate diagnostic biomarkers translate consistently to Indian breast cancer cohorts.
Methodology & Findings
- Cohort Integration & Preprocessing:
- Analyzed global TCGA breast cancer RNA-seq data (n = 815) alongside an Indian patient cohort (n = 81).
- Applied batch-correction and normalization pipelines to ensure cross-cohort comparability.
- Unsupervised Clustering & Differential Expression (DEA):
- Evaluated transcriptional heterogeneity using UMAP dimensionality reduction.
- Performed differential expression analysis across clinically defined molecular subtypes (Estrogen Receptor
ER, Progesterone ReceptorPR, andHER2).

- Candidate Universal Biomarkers:
- Identified 3 shared DEGs—
CFH,DST, andCOPZ2—consistently dysregulated across both global and Indian populations regardless of subtype heterogeneity. - These genes represent potential pan-cohort prognostic candidates warranting further clinical validation.

- Identified 3 shared DEGs—
- Indian Cohort Specific Biomarkers:
- Genes such as KIT and SFRP1, which were mostly present in the Indian dataset, could indicate regional specificity, highlighting the value in incorporating diverse groups of people in genomic analysis

- Longitudinal samples from Indian patients might be used to confirm the prognostic utility of identified biomarkers
- Genes such as KIT and SFRP1, which were mostly present in the Indian dataset, could indicate regional specificity, highlighting the value in incorporating diverse groups of people in genomic analysis