Product
Supplier
Encyclopedia
Inquiry
Home > News > Pharma News > Cancer Res | new algorithm for decoding cancer's genetic ancestors

Cancer Res | new algorithm for decoding cancer's genetic ancestors

yaozh.com 2023-01-18

Many published large-scale cancer genomic studies have shown differences in the molecular composition of diseases between groups of different ancestral backgrounds, and that race and ethnicity are important determinants of the incidence, clinical course, and outcome of many cancers. There are two main sources of ancestral traits for cancer-derived data: the patient's self-identified race and ethnicity (SIRE) and the patient's cancer-free genotype. But SIRE is often incomplete or inaccurate and often not tied to genetic ancestry, which prevents doctors from capturing complete patient ancestry information, especially in the case of mixed ancestry. Genotyping a patient's DNA from cancer-free tissue often yields more accurate and detailed ancestral traits, but it is not applicable to all types of cancer, such as leukemia. In this case, it is necessary to infer the genetic ancestry of the patient from the nucleic acid sequence of the tumor itself.

 

Recently, the research team of Cold Spring Harbor Laboratory (CSHL) published an article entitled "Genetic Ancestry Inference from Cancer-Derived Molecular Data across Genomic and Transcriptomic Platforms" in the journal Cancer Research. The research team revealed lineage associations between cancer and race/ethnicity and developed a new algorithm capable of accurately and reliably infering a patient's genetic ancestry from tumor DNA and RNA in the absence of matching cancer-free genomic data. The research helps clinicians develop new strategies for early cancer detection and personalized treatment.

 

71f76ddd0c8a4171e8d1447523a75768.png

Article published in Cancer Research

 

The research team developed a data synthesis framework to infer genetic ancestry from cancer-derived data, including whole exomes, transcriptomes, and targeted genomes (the algorithm flow is shown in Figure 1). The algorithm first performs data synthesis on genomic data from patient samples and known ancestors. The research team applied the established ancestor inference method to the algorithm and compared the resulting results with known ancestor data to generate multiple synthetic data to evaluate its accuracy in inferring the genetic ancestors of patients. In addition, by using synthetic data, the research team was able to optimize the algorithm's inference process based on the parameters it relied on.

 

WeChat image_20230117134829.png

 

Figure 1. An overview of inferring genetic ancestry from cancer-derived molecular data using data synthesis. Source: Cancer Research

 

The research team included data from four datasets, including TCGA-Ovarian Cystadenocarcinoma (TCGA-OV), TCGA-Breast Cancer Ancestral Diversity Subset (TCGA-BRCA), Beat AML Clinical Trial (Beat AML), and a Pancreatic Ductal Adenocarcinoma Study Using PDO (PDAC) (Figure 2), and aggregated the data used in the form of Venn plots, including cancer DNA (whole exon or whole genome) sequences, Cancer RNA sequences and matching normal DNA (whole exon or whole genome) sequences. In addition, the research team used the 1,000 Genome Project (1KG) dataset as a reference, comparing it with patient molecular data to infer global ancestry at the continental level. The latter is defined as a categorical variable with five values: Africa (AFR), East Asia (EAS), Europe (EUR), the Americas (AMR), and South Asia (SAS).

THE RESEARCH TEAM PERFORMED PRELIMINARY DATA PROCESSING ON THE 1KG DATASET, LABELING THE GENOMIC (HFS) LOCATIONS OF ITS HIGH-FREQUENCY ALTERNATIVE VARIANTS AS THE BASIS FOR ANCESTOR INFERENCE, AND THE SUBSET OF HFS LOCATIONS IS CALLED THE HIGH CONFIDENCE GENOTYPE (HCG) SET. Further, the research team pruned the HCG genomic locations to reduce correlations between adjacent genotypes, resulting in a pruned set of high-confidence genotype (PHCG) locations.

 

WeChat image_20230117134834.png

 

Figure 2. Molecular data used in the study. Source: Cancer Research

 

The flow of genetic ancestry inference is shown in Figure 3, and the research team used a combination of principal component analysis (PCA) and K-nearest neighbor classification. For a subset of patients in each cohort, the research team evaluated the ancestry inference performance of the parameters K and D functions separately, and retained the number of primary dimensions to evaluate them based on data synthesis.

 

WeChat image_20230117134838.png

 

Figure 3. Flowchart of genetic ancestry inference. Source: Cancer Research

 

To verify the effectiveness of the algorithm, the research team studied four cancer types, namely pancreatic adenocarcinoma (PDAC), ovarian cystadenocarcinoma, epithelial tumors represented by breast cancer, and hematopoietic malignancies represented by acute myeloid leukemia (AML). The research team selected D and K value pairs in the optimal range and applied them to cancer-free WES spectra in TCGA-OV and TCGA-BRCA patients. The results show that the ancestor presumption results obtained by the algorithm are consistent with the database data. The team also compared it to matching ancestor inferences based on cancer-free genotypes, and for patients with Beat AML, TCGA-OV and TCGA-BRCA, the presumption of ancestors was consistent with database data. The above results show that the algorithm shows high accuracy in all cohorts and analysis modes.

 

WeChat image_20230117134842.png

 

Figure 4. AMR-specific dependence of AUROC on inferred parameters D and K. Source: Cancer Research

 

In summary, the research team developed a computational method for accurate and robust ancestor inference from cancer-derived molecular data. This method combines a PCA-based ancestry inference technique with a method for inference parameter optimization using synthetic data to facilitate genetic lineage-oriented cancer research. The research team also created sample maps from cancers of known background and unrelated cancer-free genomes, and verified the performance of the algorithm using pancreatic, ovarian, breast and blood cancer samples of known ancestry, showing that the algorithm was more than 95% accurate.

 

Disclaimer: ECHEMI reserves the right of final explanation and revision for all the information.

Looking for chemical products? Let suppliers reach out to you!

Comment
Comment
  • Life Sciences Industry Overview

    The coverage spans the global life sciences industry across pharmaceuticals and food & nutrition, tracking the shift from lowest-cost sourcing to supply continuity, quality, and risk management, along with product trends and the growing edge of differentiated, globally capable players.
    Published in: June.2026

Trade Alert

Delivering the latest product trends and industry news straight to your inbox.
(We'll never share your email address with a third-party.)

Scan the QR Code to Share

Feedback & Suggestions
Send Message

Thank you for your feedback. If you require further assistance, please contact us by email at info@echemi.com or call us at +86-532-55729510.