1 Department of Information Technology, Washington University of Science and Technology, Virginia, USA.
2 Department of Computer Science, University of the Potomac, Washington, D.C, USA.
* Corresponding Author
ORCID Details
Mahfuz Islam Khan Jabed: https://orcid.org/0009-0001-2141-2894
Muhammad Imran: https://orcid.org/0009-0001-2069-0868
Clapher Ankur Gomes: https://orcid.org/0009-0004-9566-6784
Prudvi Saisaran Ponduru: https://orcid.org/0009-0009-5930-8265
World Journal of Advanced Research and Reviews, 2025, 27(01), 2817–2829
Article DOI: 10.30574/wjarr.2025.27.1.2649
Received on 07 June 2025; revised on 23 July 2025; accepted on 28 July 2025
Fair or poor self-rated health (SRH) is a concise population-health indicator, and scalable classification may support public-health surveillance. However, temporal robustness, calibration, complex survey design, and subgroup performance require careful evaluation. This study developed and temporally validated machine-learning models for identifying U.S. adults reporting fair or poor SRH using behavioral and socioeconomic factors. Official 2021 Behavioral Risk Factor Surveillance System data formed the development cohort, while 2022 data were reserved for temporal testing. Seventeen harmonized predictors were evaluated using logistic regression, decision tree, random forest, extra-trees, and histogram gradient-boosting models. State-grouped three-fold cross-validation guided model configuration and threshold selection. Survey weights informed model fitting and evaluation, and 200 stratified primary-sampling-unit bootstrap replicates produced 95% confidence intervals. The analysis also examined calibration, subgroup and state-level performance, predictor domains, parsimonious modeling, and complete cases. The analytical cohorts included 430,494 adults in 2021 and 434,659 in 2022, with weighted fair or poor SRH prevalence of 16.13% and 17.89%, respectively. Histogram gradient boosting performed best in the 2022 temporal test, achieving a weighted ROC-AUC of 0.801 (95% CI, 0.798–0.805), PR-AUC of 0.504 (0.497–0.513), balanced accuracy of 0.727 (0.724–0.731), and Brier score of 0.1167 (0.1154–0.1180). Its improvement over logistic regression was modest. Employment, physical activity, healthcare cost barriers, body mass index, income, and age were influential, with temporally stable importance rankings (Spearman ρ=0.978). Although behavioral and structural factors supported useful temporal classification, performance varied across population and geographic groups. Residual miscalibration, subgroup threshold disparities, self-reported data, and cross-sectional measurement preclude immediate clinical or individual-level deployment.
Self-Rated Health, Machine Learning, Health Informatics, Healthcare Analytics, Predictive Analytics, Population Health Management, Health Equity
Preview Article PDF
Mahfuz Islam Khan Jabed, Muhammad Imran, Clapher Ankur Gomes and Prudvi Saisaran Ponduru. MACHINE LEARNING-BASED PREDICTION OF POOR SELF-RATED HEALTH AMONG U.S. ADULTS USING BEHAVIORAL AND SOCIOECONOMIC FACTORS. World Journal of Advanced Research and Reviews, 2025, 27(01), 2817–2829. Article DOI: https://doi.org/10.30574/wjarr.2025.27.1.2649