A Hybrid-Effects Gradient Boosting and SHAP-Interpretable XGBoost Framework for Personalized NIPT Timing and Multiclass Aneuploidy Classification in High-BMI Pregnancies

Authors

  • Jinghan Wang Dalian Foreign Languages University, Dalian, Liaoning, 116044, China
  • Guanze Lin Dalian Foreign Languages University, Dalian, Liaoning, 116044, China

DOI:

https://doi.org/10.54097/wzjhk661

Keywords:

Non-invasive Prenatal Testing, Fetal Fraction, Body Mass Index, Risk-aware Timing, XGBoost, SHAP

Abstract

Background: Selecting an informative sampling time for non-invasive prenatal testing (NIPT) can be challenging in pregnancies with elevated body mass index (BMI), because fetal fraction and measurement reliability may vary across gestation. Methods: We analyzed 1,687 regional NIPT records and developed an integrated retrospective modeling framework. After gestational-age harmonization, missing-data handling, and quality screening, a gradient-boosting regression model with a mixed-effects correction was used to model male-fetus Y-chromosome fraction. Ward clustering and risk-constrained optimization produced BMI-stratified schedules; Cox modeling and bootstrap resampling assessed covariate-stratified schedules. For female fetuses, an XGBoost classifier with SMOTE balancing and SHAP interpretation was trained for five-class aneuploidy screening. Results: The male-fetus model achieved mean five-fold validation R² of 0.85 and threshold-classification accuracy of 92.3%. The optimized four-stratum schedule recommended testing at 10.7, 13.3, 17.0, and 20.8 weeks across increasing BMI strata, reducing the reported expected risk by 32% relative to the initial grouping. The female-fetus classifier achieved 94.2% accuracy and a multiclass AUC of 0.97; Z21, Z18, and Z13 were the dominant explanatory features. Conclusions: The framework links timing prediction, risk-aware stratification, and interpretable classification using one retrospective dataset. Prospective, multi-center validation is required before clinical deployment.

Downloads

Download data is not yet available.

References

[1] Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7.

[2] Kleinbaum, D. G., & Klein, M. (2020). Survival analysis: A self-learning text (4th ed.). Springer. https://doi.org/ 10.1007/ 978-3-030-37367-9.

[3] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/ 2939672. 2939785.

[4] Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30. Curran Associates.

[5] Wang, E., et al. (2018). Fetal fraction evaluation in non-invasive prenatal screening. Prenatal Diagnosis.

[6] Gerson, K. D., et al. (2015). Factors affecting levels of circulating cell-free fetal DNA in maternal plasma and their implications for noninvasive prenatal testing. Prenatal Diagnosis.

[7] Beulen, L., et al. (2017). Comparing methods for fetal fraction determination and quality control of NIPT samples. Prenatal Diagnosis, 37, 769–773.

[8] Huang, Y., et al. (2023). Maternal and fetal factors influencing fetal fraction: A retrospective analysis of 153,306 pregnant women undergoing noninvasive prenatal screening. Frontiers in Genetics.

[9] van Prooyen Schuurman, L., et al. (2024). Insights into non-informative results from non-invasive prenatal screening through gestational age, maternal BMI, and age analyses. Prenatal Diagnosis.

[10] American College of Obstetricians and Gynecologists. (2020). Screening for fetal chromosomal abnormalities. Obstetrics & Gynecology, 136, e48–e69.

[11] Gregg, A. R., et al. (2016). Noninvasive prenatal screening by next-generation sequencing. Genetics in Medicine, 18, 1056–1065.

[12] Kilpatrick, S. J., et al. (2015). Current guidance on prenatal screening. Obstetrics & Gynecology, 125, 653–662.

[13] van Buuren, S., & Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45, 1–67.

[14] Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29, 1189–1232.

[15] Cox, D. R. (1972). Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological), 34, 187–220.

[16] Efron, B., & Tibshirani, R. J. (1994). An introduction to the bootstrap. CRC Press.

[17] Chawla, N. V., et al. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357.

[18] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10. 1145/ 2939672. 2939785.

[19] Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30. Curran Associates.

[20] Collins, G. S., et al. (2015). TRIPOD statement. Annals of Internal Medicine, 162, 55–63.

[21] Wolff, R. F., et al. (2019). PROBAST. Annals of Internal Medicine, 170, 51–58.

[22] Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the area under a ROC curve. Radiology, 143, 29–36.

[23] DeLong, E. R., et al. (1988). Comparing areas under correlated ROC curves. Biometrics, 44, 837–845.

[24] Steyerberg, E. W., et al. (2010). Assessing prediction-model performance. Epidemiology, 21, 128–138.

[25] Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis. Medical Decision Making, 26, 565–574.

Downloads

Published

30-07-2026

Issue

Section

Articles

How to Cite

Wang, J., & Lin, G. (2026). A Hybrid-Effects Gradient Boosting and SHAP-Interpretable XGBoost Framework for Personalized NIPT Timing and Multiclass Aneuploidy Classification in High-BMI Pregnancies. Frontiers in Computing and Intelligent Systems, 17(2), 48-53. https://doi.org/10.54097/wzjhk661