Identification of Traditional Chinese Medicine Based on KNN Algorithm and Random Forest
DOI:
https://doi.org/10.54097/hset.v60i.10357Keywords:
Principal component analysis, Hierarchical clustering, Random forest.Abstract
Medicinal materials are a system of components and complex mixtures, and spectroscopic principles can provide in-depth analysis of the composition mechanism of traditional Chinese medicinal materials through the determination of their material structures. Chinese medicinal materials of different origins and varieties exhibit different spectral characteristics due to differences in chemical composition and organic matter. However, spectral features have high-dimensional attributes, so PCA principal component dimensionality reduction is considered to condense spectral features into representative feature variables. Then, a hierarchical clustering method with a clear hierarchy can be used to classify multiple Chinese medicinal materials into three categories based on their spectral characteristics; At the same time, in order to solve the problem of origin identification, based on the situation that there are many classifications of origins, the KNN algorithm is combined to achieve the requirements of origin identification; Using both mid infrared and near infrared spectral data and using KNN algorithm to verify the origin of Chinese medicinal materials is more accurate; Due to the small number of categories, the prediction of the types of Chinese medicinal materials is implemented using commonly used random forests. The realization of the above methods demonstrates the feasibility of identifying Chinese medicinal materials through infrared spectroscopy, and it is also worth further exploration and research.
Downloads
References
Dimension reduction and visualization in principal component analysis. Anal Chemistry 2008:80 (13): 4933-4944.
Beattie JR. Esmonde White FWL Principal Component Analysis: Using spectroscopy to intuitively derive principal component analysis pectrosc2021:75(4):361-375. doi: 101177/0003702820987847
Bu Jun, Liu Wei, Pan Zhen, Ling Kun. Comparative study on hydrochemical classification based on Dif é rent clustering analysis method. International Journal of Environment and Public Health. 2020:17(24):9515. Issued on December 18, 2020.
Research on Liu Xingbo's Cohesive Hierarchical Clustering Algorithm [J. Science and Technology Information (Science Teaching and Research), 2008 (11): 202
Geng Lijuan, Li Xingyi Research on KNN Algorithm for Big Data Classification [J]. Computer Application Use Research, 2014,31 (05): 1342-1344+1373.
Dou Xiaofan. Overview of KNN Algorithm [J]. Communication World, 2018 (10): 273-274.
Comprehensive genetic and epigenetic prediction of coronary artery type Y diabetes in Framingham radiotherapy study, PLoS One.2018; 13(1): e0190549
Le, S., Josse, J.Husson, F.(2008) FactoMineR: The R package for multivariate analysis. Journal of Statistical Software. 25(1)
Hechenbichler K, Schliep K. Weighted K-nearest neighbor technique and ordered classification [J]. Discussion Paper Sfb, 2004.
Breiman, L. (2001), Random Forest, Machine Learning 45 (1), 5-32
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







