The Impact of Data Preprocessing on Federated Learning for Air Quality Prediction: A Comparative Study of Imputation Methods
DOI:
https://doi.org/10.54097/6qspgz48Keywords:
Federated learning, air quality prediction, missing data imputation, data preprocessing.Abstract
Data preprocessing plays a foundational role in machine learning, but it receives limited systematic attention in federated learning (FL) environments. This study empirically compares four missing value strategies using the UCI dataset: direct deletion, mean imputation, K-Nearest Neighbors (KNN) imputation, and Random Forest (RF) imputation. The setup uses a federated framework with five clients, each running a multi layer perceptron model. The findings show direct deletion achieves the strongest performance, with a mean absolute error of 0.173 and an of 0.966, clearly exceeding imputation methods (the errors around 0.25). The NMHC(GT) feature has an 88.4% missing rate, making imputation unreliable and introducing noise. Although direct deletion reduces the sample from 7,674 to 827 observations, it safeguards data integrity. This research shows that preserving data quality should take priority over quantity when the missing rates are exceptionally high, providing the practical guidance for robust preprocessing design in federated learning systems.
References
[1] McMahan B, Moore E, Ramage D, Hampson S, y Arcas BA. Communication-efficient learning of deep networks from decentralized data. In: Proc 20th Int Conf Artif Intell Stat (AISTATS); 2017. p. 1273–1282.
[2] Kairouz P, McMahan HB, Avent B, Bellet A, Bennis M, Bhagoji AN, et al. Advances and open problems in federated learning. Found Trends Mach Learn. 2021; 14(1–2): 1–210.
[3] Yang Q, Liu Y, Chen T, Tong Y. Federated machine learning: Concept and applications. ACM Trans Intell Syst Technol. 2019; 10(2): 1–19.
[4] Zhu L, Liu Z, Han S. Deep leakage from gradients. In: Proc Adv Neural Inf Process Syst (NeurIPS); 2019. p. 14774–14784.
[5] Li T, Sahu AK, Talwalkar A, Smith V. Federated learning: Challenges, methods, and future directions. IEEE Signal Process Mag. 2020; 37(3): 50–60.
[6] Li T, Sahu AK, Zaheer M, Sanjabi M, Talwalkar A, Smith V. FedProx: Federated learning with proximal term. In: Proc Mach Learn Syst (MLSys); 2020. p. 429–450.
[7] Emmanuel T, Maupong T, Mpoeleng D, Semong T, Mphago B, Tabona O. A survey on missing data in machine learning. J Big Data. 2021; 8(1): 1–37.
[8] Seu K, Kang MS, Lee H. An intelligent missing data imputation techniques: A review. Int J Informatics Vis. 2022; 6(1–2): 278–283.
[9] Liu M, Li S, Yuan H, Ong MEH, Ning Y, Xie F, et al. Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques. Artif Intell Med. 2023; 142: 102587.
[10] De Vito FS, Massera E, Piga M, Martinotto L, Di Francia G. Air quality data set. UCI Mach Learn Repository. 2008. Available from: https://archive.ics.uci.edu/ml/datasets/Air+Quality
[11] Khan SI, Hoque ASML. SICE: An improved missing data imputation technique. J Big Data. 2020; 7(1): 1–21.
[12] Ali SS, Shaikh F, Shah SA. Missing data imputation using random forests and its application in healthcare. In: Proc Int Conf Comput Intell (ICCI); 2020. p. 152–157.
[13] Yıldız A, Er ME, Bursalı A, Çolakoğlu T, Erkuş EC. The effects of data standardization and normalization techniques in click through rate prediction. In: Proc Int Conf Adv Technol; 2023. p. 200–203.
[14] Mora A, Bujari A, Bellavista P. Enhancing generalization in federated learning with heterogeneous data: A comparative literature review. Future Gener Comput Syst. 2024; 157: 1–15.
[15] Lee KJ, Kang JB. The proportion of missing data should not be used to guide decisions on multiple imputation. BMC Med Res Methodol. 2019; 19: 159.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







