Department

Industrial and Systems Engineering

Document Type

Article

Publication Date

5-9-2026

Embargo Period

8-3-2026

Abstract

Background

Although predictors of dental caries have been previously explored, a comprehensive understanding of factors influencing permanent‐molar decay in U.S. children and adolescents, especially with respect to racial and ethnic biases remains limited. This study aims to develop and evaluate machine‐learning (ML) models incorporating algorithmic fairness to predict caries in permanent molars.

Methods

Data from the National Health and Nutrition Examination Survey (NHANES) were analyzed, using the 2011–2014 cycles for training and validation and the 2015–2016 cycle for testing. The primary outcome was decayed, missing, and filled teeth (DMFT) in at least one permanent molar, dichotomized to represent the presence or absence of caries. Candidate predictors included demographic characteristics, dietary factors, laboratory and examination measures, behaviors, and lifestyle indicators. Three ML models (logistic regression, random forest, and XGBoost) were implemented and compared using area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, and specificity. Fairness was quantified by comparing group-specific true-positive rates and false-positive rates across race/ethnicity over 10 cross-validation folds, using Shapiro–Wilk tests plus paired tests to assess subgroup gaps and pre-/post-mitigation changes. To mitigate disparities, an in-processing equalized-odds constraint was applied through an exponentiated-gradient reduction framework. Sensitivity analyses using NHANES survey weights were also conducted.

Results

Among 5,710 participants aged 6–19 years, 37.30% had caries in at least one permanent molar. Logistic regression achieved the highest pre-mitigation AUC (78.50%). SHAP analyses identified age, family income, body mass index, and household education as key contributors. Fairness evaluation indicated minimal equal-opportunity gaps but evidence of equalized-odds violations across race/ethnicity. Post-mitigation, race/ethnicity-related disparities decreased across models, with varying performance trade-offs. XGBoost and random forest maintained stable AUC values in both unweighted (∼77%–78%) and survey-weighted evaluations (∼75%–77%). In contrast, logistic regression showed a substantial unweighted AUC decline (78.50% to 65.53%) and a marked reduction in recall and F1 score, whereas changes under weighted evaluation were more modest (77.91% to 76.54%).

Conclusions

ML models effectively predicted dental caries in children and adolescents using NHANES data. Applying equalized-odds constraints reduces disparities across models, though with varying fairness-performance trade-offs. These findings highlight the importance of tailoring bias mitigation strategies to clinical priorities, balancing equitable outcomes and predictive utility, and can be used to inform policy and better design targeted interventions to improve oral health.

Journal Title

AJPM Focus

Journal ISSN

2773-0654

Volume

In Press, Journal Pre-proof

Digital Object Identifier (DOI)

https://doi.org/10.1016/j.focus.2026.100506

Comments

This article received funding through Kennesaw State University's Faculty Open Access Publishing Fund, supported by the KSU Libraries and KSU Office of Research.

Share

COinS