Leveraging free-text clinical records for heart disease classification through structured feature mapping
Article excerpt
by Noha Alnazzawi Hypertension and diabetes are major risk factors for heart disease, which remains among the leading causes of morbidity and mortality worldwide. Heart disease includes heart failure, myocardial infarction, stroke, and atherosclerosis. The identification and monitoring of these…
by Noha Alnazzawi
Hypertension and diabetes are major risk factors for heart disease, which remains among the leading causes of morbidity and mortality worldwide. Heart disease includes heart failure, myocardial infarction, stroke, and atherosclerosis. The identification and monitoring of these risk factors are crucial for early intervention and effective management. Machine learning techniques have the potential to improve the management and prevention of heart disease by enabling the automatic detection of risk factors, which in turn can help doctors personalize treatment and facilitate preventive interventions. In this study, heart disease risk factors were automatically extracted using a combination of multimodal data: both unstructured clinical narratives (e.g., the PrevComp corpus) and structured datasets (e.g., the UCI heart disease dataset) were used to predict the presence or absence of heart disease. The classification model is based on the Light Gradient Boosting Machine (LightGBM), a state-of-the-art implementation of the gradient boosting framework that employs tree-based learning algorithms. The developed classification model demonstrated a robust predictive accuracy of 83% for the presence/absence of heart disease, supporting its potential to accurately identify high-risk patients and improve clinical outcomes.