Development of a machine learning model related to explore the association between heavy metal exposure and alveolar bone loss among US adults utilizing SHAP: a study based on NHANES 2015-2018

Published in Biomedical Research

Development of a machine learning model related to explore the association between heavy metal exposure and alveolar bone loss among US adults utilizing SHAP: a study based on NHANES 2015-2018
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Explore the Research

BioMed Central
BioMed Central BioMed Central

Development of a machine learning model related to explore the association between heavy metal exposure and alveolar bone loss among US adults utilizing SHAP: a study based on NHANES 2015–2018 - BMC Public Health

Background Alveolar bone loss (ABL) is common in modern society. Heavy metal exposure is usually considered to be a risk factor for ABL. Some studies revealed a positive trend found between urinary heavy metals and periodontitis using multiple logistic regression and Bayesian kernel machine regression. Overfitting using kernel function, long calculation period, the definition of prior distribution and lack of rank of heavy metal will affect the performance of the statistical model. Optimal model on this topic still remains controversy. This study aimed: (1) to develop an algorithm for exploring the association between heavy metal exposure and ABL; (2) filter the actual causal variables and investigate how heavy metals were associated with ABL; and (3) identify the potential risk factors for ABL. Methods Data were collected from National Health and Nutrition Examination Survey (NHANES) between 2015 and 2018 to develop a machine learning (ML) model. Feature selection was performed using the Least Absolute Shrinkage and Selection Operator (LASSO) regression with 10-fold cross-validation. The selected data were balanced using the Synthetic Minority Oversampling Technique (SMOTE) and divided into a training set and testing set at a 3:1 ratio. Logistic Regression (LR), Support Vector Machines (SVM), Random Forest (RF), K-Nearest Neighbor (KNN), Decision Tree (DT), and XGboost were used to construct the ML model. Accuracy, Area Under the Receiver Operating Characteristic Curve (AUC), Precision, Recall, and F1 score were used to select the optimal model for further analysis. The contribution of the variables to the ML model was explained using the Shapley Additive Explanations (SHAP) method. Results RF showed the best performance in exploring the association between heavy metal exposure and ABL, with an AUC (0.88), accuracy (0.78), precision (0.76), recall (0.83), and F1 score (0.79). Age was the most important factor in the ML model (mean| SHAP value| = 0.09), and Cd was the primary contributor. Sex had little effect on the ML model contribution. Conclusion In this study, RF showed superior performance compared with the other five algorithms. Among the 12 heavy metals, Cd was the most important factor in the ML model. The relationship of Co & Pb and ABL are weaker than that of Cd. Among all the independent variables, age was considered the most important factor for this model. As for PIR, low-income participants present association with ABL. Mexican American and Non-Hispanic White show low association with ABL compared to Non-Hispanic Black and other races. Gender feature demonstrates a weak association with ABL. In the future, more advanced algorithms should be developed to validate these results and related parameters can be tuned to improve the accuracy of the model. Clinical trial number not applicable.

Background

Alveolar bone loss (ABL) is common in modern society. Heavy metal exposure is usually considered to be a risk factor for ABL. Some studies revealed a positive trend found between urinary heavy metals and periodontitis using multiple logistic regression and Bayesian kernel machine regression. Overfitting using kernel function, long calculation period, the definition of prior distribution and lack of rank of heavy metal will affect the performance of the statistical model. Optimal model on this topic still remains controversy. This study aimed: (1) to develop an algorithm for exploring the association between heavy metal exposure and ABL; (2) filter the actual causal variables and investigate how heavy metals were associated with ABL; and (3) identify the potential risk factors for ABL.

Methods

 Data were collected from National Health and Nutrition Examination Survey (NHANES) between 2015 and 2018 to develop a machine learning (ML) model. Feature selection was performed using the Least Absolute Shrinkage and Selection Operator (LASSO) regression with 10-fold cross-validation. The selected data were balanced using the Synthetic Minority Oversampling Technique (SMOTE) and divided into a training set and testing set at a 3:1 ratio. Logistic Regression (LR), Support Vector Machines (SVM), Random Forest (RF), K-Nearest Neighbor (KNN), Decision Tree (DT), and XGboost were used to construct the ML model. Accuracy, Area Under the Receiver Operating Characteristic Curve (AUC), Precision, Recall, and F1 score were used to select the optimal model for further analysis. The contribution of the variables to the ML model was explained using the Shapley Additive Explanations (SHAP) method.

Results

 RF showed the best performance in exploring the association between heavy metal exposure and ABL, with an AUC (0.88), accuracy (0.78), precision (0.76), recall (0.83), and F1 score (0.79). Age was the most important factor in the ML model (mean| SHAP value| = 0.09), and Cd was the primary contributor. Sex had little effect on the ML model contribution.

Conclusion

 In this study, RF showed superior performance compared with the other five algorithms. Among the 12 heavy metals, Cd was the most important factor in the ML model. The relationship of Co & Pb and ABL are weaker than that of Cd. Among all the independent variables, age was considered the most important factor for this model. As for PIR, low-income participants present association with ABL. Mexican American and Non-Hispanic White show low association with ABL compared to Non-Hispanic Black and other races. Gender feature demonstrates a weak association with ABL. In the future, more advanced algorithms should be developed to validate these results and related parameters can be tuned to improve the accuracy of the model.

Follow the Topic

Biomedical Research
Life Sciences > Health Sciences > Biomedical Research

Related Collections

With Collections, you can get published faster and increase your visibility.

Appropriate use of antibiotics: public health strategies, knowledge, and practice gaps

BMC Public Health is calling for submissions to our Collection on Appropriate use of antibiotics: public health strategies, knowledge, and practice gaps.

Misuse of antibiotics contributes to antimicrobial resistance (AMR), posing a threat to the future management of bacterial diseases. However, it is a multi-faceted problem without a simple solution.

Antibiotic misuse can take various forms, each requiring different strategies to address. In healthcare systems, antibiotic overuse is often driven by the tension between clinical uncertainty and the desire to offer patients a treatment that may improve their symptoms or prevent them from developing complications. The prescription of antibiotics is often done empirically, driven by the difficulty of distinguishing between bacterial and viral infections at the point of care, or the worry that a lack of intervention could have consequences.

Added to this, people in the community can contribute to inappropriate use by reusing or sharing leftover antibiotics from prior prescriptions. Similarly, misplaced expectations around the benefits of antibiotics can drive misuse in the community, pointing to the need for community-focused and community-led initiatives to inform the public on the use of antibiotics and the collective impact of antimicrobial resistance.

However, social science reframes antibiotic overuse as more than an individual behaviour problem. Antibiotics frequently act as a social and structural “quick fix” that supports care, productivity, hygiene, and coping with inequality in everyday life. They are used to compensate for gaps in water, sanitation, social protection, and health-system capacity, suggesting that interventions focusing narrowly on individual knowledge and attitudes may be unlikely to succeed unless they also address these wider drivers.

This Collection aims to explore the various dimensions of the responsible use of antibiotics, examining the prevalence of misuse and strategies for reducing unnecessary use, covering interventions aimed at prescribers and pharmacists as well as 'bottom-up' strategies such as community education campaigns. We invite contributions that investigate the roles of healthcare providers, patients, professional guidance, communities, and policymakers in addressing this pressing issue.

Potential topics for submission include, but are not limited to:

Patterns of antibiotics overuse in various populations

The role of healthcare providers in preventing misuse

Public health campaigns and gross-roots initiatives to promote responsible use of antibiotics

Policy frameworks for improving the prescribing and dispensing of antibiotics

Social and structural drivers of antibiotic use (ethnographic, anthropological, and political-economy analyses).

Interventions that address upstream determinants (water, sanitation, social protection, labour conditions) alongside stewardship measures.

This Collection supports and amplifies research related to SDG 3 (Good Health and Well-being).

All manuscripts submitted to this journal, including those submitted to collections and special issues, are assessed in line with our editorial policies and the journal’s peer review process. Reviewers and editors are required to declare competing interests and can be excluded from the peer review process if a competing interest exists.

Publishing Model: Open Access

Deadline: Oct 03, 2026

Digital exclusion and health equity

Digital exclusion poses significant barriers to health equity, particularly in the context of an increasingly technology-driven healthcare landscape. Many individuals, especially from underserved communities face challenges in accessing digital health resources, including telehealth services, electronic health records, and health information online, alongside challenges from other social determinants of health. This Collection seeks to explore the intersection of digital health access and health equity, focusing on the disparities that arise from unequal access to digital tools and resources.

Addressing digital health exclusion is critical for promoting health equity and ensuring that all individuals can benefit from advancements in digital health. Recent strides in telehealth and digital health interventions have demonstrated the potential for technology to improve access to care, yet they also reveal significant disparities that must be acknowledged. By understanding the factors that contribute to eHealth disparities, we can develop targeted digital inclusion strategies that address the unique needs of diverse populations, ultimately fostering a more equitable health system.

Topics for submission include but are not limited to:

  • Equitable digital health systems: policy, infrastructure, and community engagement
  • Digital health access among underserved communities
  • Digital literacy and health outcomes across diverse populations
  • Innovations and implementation strategies for bridging digital health gaps in underserved communities
  • Measuring and monitoring digital health disparities: metrics and methodologies

This Collection supports and amplifies research related to SDG 3 (Good Health and Well Being) and SDG 10: (Reduced Inequalities).

All manuscripts submitted to this journal, including those submitted to collections and special issues, are assessed in line with our editorial policies and the journal’s peer-review process. Reviewers and editors are required to declare competing interests and can be excluded from the peer review process if a competing interest exists.

Publishing Model: Open Access

Deadline: Dec 11, 2026