Walmart Sales Prediction Based on Machine Learning
DOI:
https://doi.org/10.54097/hset.v47i.8170Keywords:
Machine learning, Feature engineering, Walmart sales.Abstract
Accurate sales forecasting can improve a company's profitability while minimizing expenditures. The use of machine learning algorithms to predict product sales has become a hot topic for researchers and companies over the past few years. This report features the machine learning sales prediction model that combines the ML algorithm and meticulous feature engineering processing to predict Walmart sales. The following regressions are analyzed in this paper: linear regression, random forest regression, and XGBoost regression. The regression analysis has been tested for the same time period every year for three years from 2010 to 2012 on a continuous time basis. The experiments show that XGBoost algorithm overperforms the other machine learning methods by examining the same evaluation metric (WAME). The findings can contribute to a better understanding of the development of new decision support for the retail industry e.g., Walmart retail stores. Moreover, this paper also represents a detailed procedure to rank the feature importance for the dataset. Within the next few years, the ML algorithm is destined to become an important approach for business forecasting. However, this strategy largely ignores the time series method for accuracy.
Downloads
References
Chu, C. W., & Zhang, G. P. (2003). A comparative study of linear and nonlinear models for aggregate retail sales forecasting. International Journal of production economics, 86(3), 217-231.
Zhao, J., Tang, W., Fang, X., Wang, J., Liu, J., Ouyang, H., ... & Qiang, J. (2015, October). A novel electricity sales forecasting method based on clustering regression and time-series analysis. In Proceedings of the 2015 International Conference on Artificial Intelligence and Software Engineering.
Thomassey, S., & Fiordaliso, A. (2006). A hybrid sales forecasting system based on clustering and decision trees. Decision Support Systems, 42(1), 408-421.
Johannesen, N. J., Kolhe, M., & Goodwin, M. (2019). Relative evaluation of regression tools for urban area electrical energy demand forecasting. Journal of cleaner production, 218, 555-564.
Asuero, A. G., Sayago, A., & González, A. G. (2006). The correlation coefficient: An overview. Critical reviews in analytical chemistry, 36(1), 41-59.
Kuzlu, M., Cali, U., Sharma, V., & Güler, Ö. (2020). Gaining insight into solar photovoltaic power generation forecasting utilizing explainable artificial intelligence tools. IEEE Access, 8, 187814-187823.
Jordan, M. I., & Mitchell, T. M. (2015). Machine learning: Trends, perspectives, and prospects. Science, 349(6245), 255-260.
Uyanık, G. K., & Güler, N. (2013). A study on multiple linear regression analysis. Procedia-Social and Behavioral Sciences, 106, 234-240.
Xue, L., Liu, Y., Xiong, Y., Liu, Y., Cui, X., & Lei, G. (2021). A data-driven shale gas production forecasting method based on the multi-objective random forest regression. Journal of Petroleum Science and Engineering, 196, 107801.
Dietterich, T. G. (2000). An experimental comparison of three methods for constructing ensembles of decision trees: Bagging, boosting, and randomization. Machine learning, 40(2), 139-157.
Cleger-Tamayo, S., Fernández-Luna, J. M., & Huete, J. F. (2012, September). On the Use of Weighted Mean Absolute Error in Recommender Systems. In RUE@ RecSys (pp. 24-26).
Singh, D., & Singh, B. (2020). Investigating the impact of data normalization on classification performance. Applied Soft Computing, 97, 105524.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







