Research on Preprocessing and Downstream Analysis of Single-Cell Transcriptome Data Based on Seurat and Harmony
DOI:
https://doi.org/10.54097/f7dfsv24Keywords:
Single-cell transcriptome, Seurat, Harmony, Batch effect correction, Cell clustering, Cell annotation, Differential expression analysis.Abstract
To systematically decipher the biological information embedded in single-cell transcriptomic data, this study utilized the 10X Genomics PBMC 3k dataset as the research subject and integrated the Seurat package in R with the Harmony algorithm to conduct a comprehensive analytical workflow—from raw count matrices to cell annotation and differential expression analysis. The study first applied quality filtering at both gene and cell levels to remove low-quality data, followed by effective batch effect correction using the Harmony algorithm. Subsequent steps included linear dimensionality reduction via PCA, non-linear dimensionality reduction through UMAP, and cell clustering based on Shared Nearest Neighbor (SNN) graph partitioning. Cell type identification was accomplished by combining automated annotation with SingleR and manual validation via marker genes, ultimately leading to the screening of differentially expressed genes across distinct cell subpopulations. Results demonstrated that after quality filtering, 4,200 high-quality cells and 25,000 valid genes were retained; the Harmony algorithm completely eliminated batch biases between stimulated and control groups. Clustering at a resolution of 0.8 yielded 11 biologically meaningful PBMC subpopulations, successfully identifying major immune cell types—such as T cells, B cells, and monocytes—along with their characteristic marker genes. The established single-cell transcriptomic data analysis pipeline provides a robust methodological framework for resolving cellular heterogeneity and characterizing feature genes of cell subtypes, thereby laying a data foundation for functional studies of immune cells.
Downloads
References
[1] Stegle O, Teichmann SA, Marioni JC. Computational and analytical challenges in single-cell transcriptomics. Nature Reviews Genetics, 2015, 16(3):133-145.
[2] Zheng GXY, Terry JM, Belgrader P, et al. Massively parallel digital transcriptional profiling of single cells. Nature Communications, 2017, 8(1):14049.
[3] Luecken MD, Theis FJ. Current best practices in single‐cell RNA‐seq analysis: a tutorial. Molecular Systems Biology, 2019, 15: MSB188746.
[4] Hao Y, Hao S, Andersen-Nissen E, et al. Integrated analysis of multimodal single-cell data. Cell, 2021, 184(13):3573-3587.
[5] Korsunsky I, Millard N, Fan J, et al. Fast, sensitive and accurate integration of single - cell data with Harmony. Nature Methods, 2019, 16(12):1289 - 1296.
[6] Becht E, McInnes L, Healy J, et al. Dimensionality reduction for visualizing single-cell data using UMAP. Nature Biotechnology, 2019, 37(1):38-44.
[7] Xu C, Su Z. Identification of cell types from single-cell transcriptomes using a novel clustering method. Bioinformatics, 2015, 31(12): 1974-1980.
[8] Aran D, Looney AP, Liu L, et al. Reference-based analysis of lung single-cell sequencing reveals a transitional profibrotic macrophage. Nature Immunology, 2019, 20(2):163-172.
[9] Butler A, Hoffman P, Smibert P, et al. Integrating single-cell transcriptomic data across different conditions, technologies, and species[J]. Nature Biotechnology, 2018, 36(5):411-420.
[10] Haghverdi L, Lun ATL, Morgan MD, et al. Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors[J]. Nature Biotechnology, 2018, 36(5):421-427.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

