A groundbreaking computational algorithm for characterizing cancer cells and tracking their evolution has surpassed previous methods, demonstrating superior speed and accuracy
This web page was produced as an assignment for an undergraduate course at Davidson College.
Single-cell sequencing allows researchers to examine the genetic makeup of individual cells, providing insights into cellular diversity and function within complex biological systems. Over the years, scientists have been trying to better understand the genetic makeup of cancer cells. Researchers use single-cell data coupled with analysis technologies to distinguish between cancerous and non-cancerous cells and understand their genetic changes. Many computational methods have been used to tackle this problem. However, these methods have shortcomings due to the lack of accuracy and speed in analyzing single-cell data.
The research group, De Falco et al. 20231, proposed using single-cell variational ANeuploidy analysis (SCEVAN), an improved method that can automatically and accurately characterize data from single-cell sequencing. Their method can also track evolutionary relationships between different groups of cancer cells. SCEVAN can give us essential insights into how tumors grow and how we might develop better cancer treatments.

Image created using BioRender
RNA-seq data is filtered by selecting highly confident normal cells with significant gene expression. Normal cells are used to establish a baseline, serving as a reference for comparison with other cells. The baseline is subtracted from the dataset to normalize data. Next, a smoothing program will clean the data, and SCEVAN will segment to depict new normalized data accurately. Now, SCEVAN categorizes normal and tumor cells by classifying them into two clusters: normal and tumorous. The tumorous cells are further categorized based on “subclones,” which share genetic alterations. The subclones are further segmented individually based on the number of gene copies. Lastly, SCEVAN identifies and characterizes alterations common to all subclones (truncal), those shared among some subclones, and alterations unique to specific subclones. It achieves this by comparing different clusters to understand the evolutionary relationships among subclones.
Putting SCEVAN to the test, the researchers aimed to assess how accurately SCEVAN could determine whether a cell was cancerous or non-cancerous. For this experiment, they made datasets with known classifications of tumor and normal cells. These datasets mimicked real single-cell RNA-seq data. They then applied SCEVAN and CopyKAT, another method for classifying malignant and non-malignant cells in RNA-seq datasets (Gao et al. 2021)2. De Flaco et al. assessed the performance of each method by calculating the F1 score, a measure that combines precision and recall. The results showed that SCEVAN consistently outperformed CopyKAT, with higher F1 scores indicating better accuracy in classifying cells as cancerous or non-cancerous.
They also used SCEVAN and CopyKAT to analyze real sequencing data that had already been labeled to differentiate between normal and tumor cells (Yu et al. 2020)3. Their analysis revealed that SCEVAN achieved a better classification score than CopyKAT in 63% of the samples, while CopyKAT performed better in 23%. Overall, SCEVAN obtained an average F1 score of 0.90 across all samples, while CopyKAT’s average F1 score was 0.63. These results demonstrate that SCEVAN can accurately distinguish between tumor and normal cells in various solid tumors from scRNA-seq data.
Next, the team compared SCEVAN’s computational speed with that of other tools, specifically in two aspects: the classification of malignant cells and the segmentation step, which involves identifying regions of the genome that have similar patterns of copy number changes, such as amplifications or deletions of genetic material (Stankiewicz and Lupski 2002)4 They found that SCEVAN was 2–7 times faster than other methods in distinguishing malignant and non-malignant cells. Meaning SCEVAN can identify cancerous cells more quickly than alternative approaches. These results indicate that SCEVAN’s algorithm for segmentation is notably efficient compared to other tools commonly used for inferring copy number alterations from single-cell RNA sequencing data.
Having used SCEVAN to identify cancer cells, they also wanted to learn about the diversity between these cancer cells. They investigated the heterogeneity of glioblastoma(a type of brain tumor) across different regions of tumors, as a single biopsy may not capture the entire tumor’s complexity. By analyzing multiple biopsies using SCEVAN, they identified distinct clonal structures within each sample, even finding a group with seven biopsies. For instance, biopsies taken from the tumor periphery showed different copy number alterations compared to those from the core, with variations in chromosome amplifications. This analysis revealed evolutionary relationships between tumor regions, highlighting their clonal architecture and dynamics.
Continuing, SCEVAN was used to compare primary tumors and lymph node metastases(cancer cells that have spread) in patients with head and neck squamous cell carcinoma (HNSCC). In one patient (HNSCC5), the clonal structure differed between the primary tumor and metastasis, with a notable absence of chromosome 7 amplification in the metastasis. This region includes the GPNMB gene, which is downregulated in metastasis and is associated with tumor growth and spread. The clonal structures of primary tumors and metastases were similar for other patients, showing a high correlation. This highlights SCEVAN’s utility in studying the clonal evolution and spread of metastatic cancer.
In short, De Flaco et al. 2023 see SCEVAN as a robust tool for studying single-cell datasets, helping to characterize tumor cells and identify specific genomic changes within cancer cells. With the power of SCEVAN, scientists can develop better medicinal compounds to combat various cancers known to plague humans worldwide. The algorithm has a good blend of speed and accuracy that can be applied to most RNA-seq datasets. However, there were some cases in which SCEVAN misclassified some cells, indicating that the algorithm was not wholly perfect in characterizing particular cells. A limitation is that SCEVAN assumes that cancer cells can be identified based on their genetic alterations, particularly aneuploidy, which refers to an abnormal number of chromosomes. However, there are certain types of cancers, such as leukemia (a type of blood cancer), pediatric cancers, and Ependymomas (a type of brain tumor), that may have very few genetic alterations. It is vital to understand that the sole use of SCEVAN is impractical as it is lacking in some areas. Although SCEVAN may not excel in characterizing some cancer cells, it still has its place in genomic studies. Instead, It’s essential to use multiple methods to understand complex data in cancer studies and other disciplines. Also, it is crucial to have diverse RNA-seq datasets to represent different groups. RNA-seq data can vary across diverse populations due to genetic background, environmental factors, and lifestyle differences. These variations can influence gene expression levels and patterns, leading to differences in RNA-seq data profiles between populations. Age, sex, ethnicity, and geographical location can also contribute to variability in RNA-seq data. Therefore, it’s essential to consider population-specific factors when analyzing RNA-seq data and interpreting the results in studies involving diverse populations.
References:
- De Falco A., F. Caruso, X.-D. Su, A. Iavarone, and M. Ceccarelli, 2023 A variational algorithm to detect the clonal copy number substructure of tumors from scRNA-seq data. Nat Commun 14: 1074. https://doi.org/10.1038/s41467-023-36790-9
- Gao R., S. Bai, Y. C. Henderson, Y. Lin, A. Schalck, et al., 2021 Delineating copy number and clonal substructure in human tumors from single-cell transcriptomes. Nat Biotechnol 39: 599–608. https://doi.org/10.1038/s41587-020-00795-2
- Yu K., Y. Hu, F. Wu, Q. Guo, Z. Qian, et al., 2020 Surveying brain tumor heterogeneity by single-cell RNA-sequencing of multi-sector biopsies. National Science Review 7: 1306–1318. https://doi.org/10.1093/nsr/nwaa099
- Stankiewicz P., and J. R. Lupski, 2002 Genome architecture, rearrangements and genomic disorders. Trends in Genetics 18: 74–82. https://doi.org/10.1016/S0168-9525(02)02592-1
Contact author here
© Copyright 2022 Department of Biology, Davidson College, Davidson, NC 28036
It’s truly amazing how much potential computation algorithms have in detecting normal versus abnormal cell types based on their RNA expression, but even this study faced the challenge of sampling bias from biopsies. Would taking multiple biopsy samples contribute to novel research discoveries and be feasible to sequence multiple biopsies for each patient in the future, or would this come at the financial cost for conducting multiple sequencing analysis for each patient? Additionally, as the tumor periphery was found to have differing copy number alterations compared to the core, I would be interested to learn how this would affect possible gene therapies and how this discovery may give insight to developing cancer treatments.