Predicting Heart Disease Using Demographic, Behavioural, and Clinical Risk Factors: A Public Health Approach
Keywords:
heart disease, machine learning, public health, risk prediction, cardiovascular preventionAbstract
Identifying those at high risk of heart disease early is critical to population-level prevention and effective use of health care resources. In this study, machine learning models were created based on demographic, behavioral and clinical risk factors to assist public health screening. The data were used for a cross-sectional analysis. 5,510 participants were included after removing duplicate observations, with 527 having heart disease. The models were trained with Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, and XGBoost on 80% of the data and tested on the remaining 20% of the data which was a stratified split. The Synthetic Minority Over-sampling Technique was employed on the Training data to address class imbalance. Accuracy, precision, recall, F1 score, ROC–AUC, confusion matrix and five-fold cross-validation were used to assess the performance. The prevalence of heart disease was 9.56%. Participants with heart disease were older with higher rates of smoking history, hypertension, diabetes, high cholesterol, previous stroke, and family history of heart attack. Logistic Regression performed the best results in terms of overall performance with an accuracy of 0.790, a recall of 0.829, an F1 score of 0.430, an ROC–AUC of 0.876, and a cross-validated ROC–AUC of 0.865 ± 0.018. The most important predictors were age, followed by hypertension, smoking history, use of cholesterol-lowering drugs, history of stroke, sex, family history and diabetes. Interpretable machine learning plays a crucial role in population-based heart disease risk stratification and targeted preventive interventions.