A Novel Feature Fusion Method Based on RSCU and K-mer to Classify the SARS-Cov-2
DOI:
https://doi.org/10.54097/yay6zs23Keywords:
SARS-Cov-2, Support vector machine, K-mer, Relative Synonymous Codon Usage.Abstract
The SARS-Cov-2 virus exhibits a high mutation rate, which makes the prediction and classification of its genetic evolution and variation trends highly significant. Accurate classification methods not only contribute to epidemiological studies of the virus, but also play a crucial role in vaccine development and antiviral drug discovery. This study aims to systematically evaluate the accuracy and generalization capability of the RSCU (Relative Synonymous Codon Usage) and K-mer encoding techniques in the classification of the SARS-CoV-2 genome. We extracted genomic data from two major SARS-CoV-2 variants, Alpha and Beta, and applied the Support Vector Machine (SVM) classification algorithm to train the data and assess the impact of different feature encoding methods on classification performance. Furthermore, we introduce a novel multi-feature fusion method, KRSCU, which combines the sequence position information from K-mer with the synonymous codon compositions from RSCU. This method effectively captures subtle differences in genomic data, significantly improving both the accuracy and generalization capability of the classification model. Experimental results demonstrate that the KRSCU method outperforms traditional single-feature encoding approaches in SARS-CoV-2 subtype classification tasks. Our research offers new insights into genomic data analysis, with potential applications in viral mutation monitoring.
Downloads
References
[1] Zhou, P., Yang, X. L., Wang, X. G., Hu, B., Zhang, L., Zhang, W., ... & Shi, Z. L. (2020). A pneumonia outbreak associated with a new coronavirus of probable bat origin. nature, 579(7798), 270-273.
[2] Korber, B., Fischer, W. M., Gnanakaran, S., Yoon, H., Theiler, J., Abfalterer, W., ... & Montefiori, D. C. (2020). Tracking changes in SARS-CoV-2 spike: evidence that D614G increases infectivity of the COVID-19 virus. Cell, 182(4), 812-827.
[3] Zhang, J., & Liu, B. (2019). A review on the recent developments of sequence-based protein feature extraction methods. Current Bioinformatics, 14(3), 190-199.
[4] Bzhalava, Z., Tampuu, A., Bała, P., Vicente, R., & Dillner, J. (2018). Machine Learning for detection of viral sequences in human metagenomic datasets. BMC bioinformatics, 19, 1-11.
[5] Song, W., Ji, C., Chen, Z., Cai, H., Wu, X., Shi, C., & Wang, S. (2022). Comparative analysis the complete chloroplast genomes of nine Musa species: genomic features, comparative analysis, and phylogenetic implications. Frontiers in Plant Science, 13, 832884.
[6] Oh, J. W., & Beer, M. A. (2024). Gapped-kmer sequence modeling robustly identifies regulatory vocabularies and distal enhancers conserved between evolutionarily distant mammals. Nature communications, 15(1), 6464.
[7] Kaniwa, F. (2018). A kmer-based parallel algorithm for pattern searching in DNA sequences on shared-memory model (Doctoral dissertation, Botswana International University of Science & Technology (Botswana)).
[8] Sharp, P. M., & Li, W. H. (1987). The codon adaptation index-a measure of directional synonymous codon usage bias, and its potential applications. Nucleic acids research, 15(3), 1281-1295.
[9] Novoa, E. M., & de Pouplana, L. R. (2012). Speeding with control: codon usage, tRNAs, and ribosomes. Trends in Genetics, 28(11), 574-581.
[10] Khandia, R., Gurjar, P., Kamal, M. A., & Greig, N. H. (2024). Relative synonymous codon usage and codon pair analysis of depression associated genes. Scientific Reports, 14(1), 3502
[11] Schneider, T. D., & Stephens, R. M. (1990). Sequence logos: a new way to display consensus sequences. Nucleic acids research, 18(20), 6097-6100.
[12] He, C., Washburn, J. D., Schleif, N., Hao, Y., Kaeppler, H., Kaeppler, S. M., ... & Liu, S. (2024). Trait association and prediction through integrative k‐mer analysis. The Plant Journal, 120(2), 833-850.
[13] Moeckel, C., Mareboina, M., Konnaris, M. A., Chan, C. S., Mouratidis, I., Montgomery, A., ... & Georgakopoulos-Soares, I. (2024). A survey of k-mer methods and applications in bioinformatics. Computational and Structural Biotechnology Journal.
[14] Van Etten, J., Stephens, T. G., & Bhattacharya, D. (2023). A k-mer-based approach for phylogenetic classification of taxa in environmental genomic data. Systematic biology, 72(5), 1101-1118.
[15] Tahara, S., Tsuchiya, T., Matsumoto, H., & Ozaki, H. (2023). Transcription factor-binding k-mer analysis clarifies the cell type dependency of binding specificities and cis-regulatory SNPs in humans. BMC genomics, 24(1), 597.
[16] Gupta, S., & Shankar, R. (2024). Comprehensive analysis of computational approaches in plant transcription factors binding regions discovery. Heliyon, 10(20).
[17] Nguyen, E., Poli, M., Faizi, M., Thomas, A., Wornow, M., Birch-Sykes, C., ... & Baccus, S. (2024). Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution. Advances in neural information processing systems, 36.
[18] Sievers, A., Bosiek, K., Bisch, M., Dreessen, C., Riedel, J., Froß, P., ... & Hildenbrand, G. (2017). K-mer content, correlation, and position analysis of genome DNA sequences for the identification of function and evolutionary features. Genes, 8(4), 122.
[19] Hadfield, J., Megill, C., Bell, S. M., Huddleston, J., Potter, B., Callender, C., ... & Neher, R. A. (2018). Nextstrain: real-time tracking of pathogen evolution. Bioinformatics, 34(23), 4121-4123.
[20] Endrullat, C., Glökler, J., Franke, P., & Frohme, M. (2016). Standardization and quality management in next-generation sequencing. Applied & translational genomics, 10, 2-9.
[21] Giovanetti, M., Slavov, S. N., Fonseca, V., Wilkinson, E., Tegally, H., Patané, J. S. L., ... & Covas, D. T. (2022). Genomic epidemiology of the SARS-CoV-2 epidemic in Brazil. Nature Microbiology, 7(9), 1490-1500.
[22] Beerenwinkel, N., Günthard, H. F., Roth, V., & Metzner, K. J. (2012). Challenges and opportunities in estimating viral genetic diversity from next-generation sequencing data. Frontiers in microbiology, 3, 329.
[23] Cortes, C. (1995). Support-Vector Networks. Machine Learning.
[24] Varma, S., & Simon, R. (2006). Bias in error estimation when using cross-validation for model selection. BMC bioinformatics, 7, 1-8.
[25] Griffel, L. M., Delparte, D., & Edwards, J. (2018). Using Support Vector Machines classification to differentiate spectral signatures of potato plants infected with Potato Virus Y. Computers and electronics in agriculture, 153, 318-324.
[26] Busia, A., Dahl, G. E., Fannjiang, C., Alexander, D. H., Dorfman, E., Poplin, R., ... & DePristo, M. (2018). A deep learning approach to pattern recognition for short DNA sequences. BioRxiv, 353474.
[27] Gunasekaran, H., Ramalakshmi, K., Rex Macedo Arokiaraj, A., Deepa Kanmani, S., Venkatesan, C., & Suresh Gnana Dhas, C. (2021). Analysis of DNA sequence classification using CNN and hybrid models. Computational and Mathematical Methods in Medicine, 2021(1), 1835056.
[28] Mandal, I. (2015). A novel approach for predicting DNA splice junctions using hybrid machine learning algorithms. Soft Computing, 19, 3431-3444.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Academic Journal of Science and Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.








