Prediction of Scrabble game players based on ARIMA and Machine Learning under big data

Authors

  • Xinhua Lu

DOI:

https://doi.org/10.54097/hset.v60i.10363

Keywords:

Big Data, ARIMA, Random Forest Regressor, Word Feature Mining.

Abstract

Wordle is a favored puzzle. Firstly, to predict its daily player count, based on analyzing the inherent trend of the data and verifying its stationarity, the effectiveness of using the ????????????????????((2,3,5),2,(2,3,5)) model to predict the number of reported results is proved. Before predicting, this thesis preprocesses the data from two parts: Number of reported results and Number in hard mode. Then, the 95% prediction interval for March 1, 2023, is [8626 16199]. Secondly, to predict the distribution of reported results percentages, this thesis uses 7 random forest regressors, their feature variables have 5 dimensions (Contest number, Frequency of word, Number of repeated types, Number of letter types, Maximum number of repetitions), and the response variable takes one of the percentages (7 dimensions) in turn. Results show that the distribution of reported results percentages of “SALSA” on March 1,2023 is [0, 2, 15, 35, 31, 15, 2].

Downloads

Download data is not yet available.

References

WU Peijin. Journal of Chifeng University(Natural Science Edition), 2022,38(06):22-26.DOI:10.13398/j.cnki.issn1673-260x.2022.06.013.

Wu Jinxia. Text Difficulty Judgment for English Learning[D]. Harbin Institute of Technology, 2007.

Tuo Siwei. Global Sales Forecast of Home Game Consoles and Review of Entering China——Analysis Based on ARIMA Model and PEST Framework[J]. China Market, 2015, No.866(51):130-134.DOI:10.13939/j. cnki.zgsc.2015.51.130.

Ye Congzhou, Xiao Penglin, Qin Jun, Zhang Chengxiong, Chen Lie, Wang Ke.Research on Energy Consumption Load Forecasting Model of Super Large Buildings Based on Clustering and Random Forest Regression[J].Green Building,2022,14(05):48- 51+55.

Xu Shasha, Ma Chao. Analysis on the Writing Questions of College Entrance Examination (Zhejiang Volume) Based on Text Feature Mining [J]. Foreign Language Testing and Teaching, 2021, No.43(03):60-64.

Study Definition & Meaning | Dictionary.com

Cheng Hongliang, Liu Hong, Bai Zhaoxu, Rao Siwei, Zhang Jian. A missing data filling method based on the nearest neighbor KNN algorithm [P]. Shaanxi Province: CN107193876B, 2020-10-09.

Zhou Liang. Journal of Hunan University of Finance and Econom-ics,2018,34(01):72-78.DOI:10.16546/j.cnki.cn43-1510/f.2018.01.008

Li Xiufang, Huang Zhiguo, Chen Xiaowei. Application Research of Bagging Integration Method in Insurance Fraud Identification [J]. Insurance Research, 2019, No.372(04):66-84.DOI:10.13497/j.cnki.is. 2019.04.006.

English-Corpora: COCA

Liu Ziliang, Ju Xiang, Zhang Yongfang, etc. Random Forest Parameter Tuning Optimization Based on Improved Random Search Algorithm [J]. Network Security Technology and Application, 2022, No.256(04):49-51.

Downloads

Published

25-07-2023

How to Cite

Lu, X. (2023). Prediction of Scrabble game players based on ARIMA and Machine Learning under big data. Highlights in Science, Engineering and Technology, 60, 246-254. https://doi.org/10.54097/hset.v60i.10363