تكييف نماذج اللغة المعتمدة على المحولات لاكتشاف أمراض القلب واستخراج عوامل الخطر
Adapting transformer-based language models for heart disease detection and risk factors extraction

المجلة: Journal Of Big Data، المجلد: 11، العدد: 1
DOI: https://doi.org/10.1186/s40537-024-00903-y
تاريخ النشر: 2024-04-04
المؤلف: Essam H. Houssein وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تبحث هذه الدراسة في تطبيق تقنيات التعلم بالنقل، وبالتحديد النماذج المعتمدة على المحولات، لتعزيز تحديد عوامل خطر مرض القلب من السجلات الصحية الإلكترونية (EHRs). تؤكد الدراسة على أهمية التنبؤ بدقة بمرض القلب من خلال الكشف عن عوامل الخطر المختلفة، بما في ذلك السكري وارتفاع ضغط الدم والتاريخ العائلي. نظرًا للتحديات التي تطرحها الطبيعة غير المنظمة للسجلات الصحية الإلكترونية، استخدم المؤلفون نماذج مثل BERT وRoBERTa وBioClinicalBERT وXLNet وBioBERT، التي تم تدريبها مسبقًا على أدبيات طبية واسعة. تم ضبط هذه النماذج باستخدام مجموعة بيانات i2b2، مما أدى إلى أداء متفوق مقارنة بالطرق التقليدية في استخراج المعلومات السريرية ذات الصلة.

تكشف النتائج أن نموذج RoBERTa حقق أعلى درجة F1 الدقيقة بنسبة 94.27%، بينما أظهرت النماذج الأخرى أيضًا نتائج تنافسية. أسفر نهج التجميع الذي يجمع بين جميع النماذج الخمسة عن درجة F1 الدقيقة بنسبة 94.26%، مما يوضح فعالية التعلم بالنقل في هذا المجال. تختتم الدراسة بأن النماذج المعتمدة على المحولات تحسن بشكل كبير من تحديد عوامل خطر مرض القلب وتقترح أن تركز الأعمال المستقبلية على معالجة القضايا المتعلقة بالتضمين وتطبيق الضبط الدقيق على مهام معالجة اللغة الطبيعية السريرية الأخرى.

مقدمة

تسلط المقدمة الضوء على التحدي الكبير الذي يشكله مرض القلب للصحة العامة، والذي لا يزال سببًا رئيسيًا للوفيات على مستوى العالم. مع تأثر حوالي 647,000 أمريكي سنويًا و17.8 مليون وفاة في جميع أنحاء العالم تُعزى إلى مرض القلب في عام 2017، يتم التأكيد على ضرورة وجود استراتيجيات فعالة للوقاية. يؤكد النص على ضرورة التعرف المبكر على عوامل الخطر مثل مرض الشريان التاجي (CAD) والسكري وارتفاع ضغط الدم والسمنة، خاصةً بالنظر إلى أن العديد من الأفراد الذين يتعرضون لنوبات قلبية لا يظهرون ملفات تعريف خطر تقليدية. يتم الإشارة إلى تعقيد استخراج عوامل الخطر ذات الصلة من السجلات الصحية الإلكترونية (EHRs)، خاصةً بسبب التباين والطبيعة الفردية للسرد السريري، مما يعقد جهود التنبؤ الآلي.

تقترح الورقة الاستفادة من تقنيات التعلم بالنقل، وخاصة من خلال النماذج المعتمدة على المحولات، لتعزيز تحديد عوامل خطر مرض القلب من بيانات EHR. من خلال نمذجة هذه المهمة كتحدي التعرف على الكيانات المسماة (NER)، تهدف الدراسة إلى تقييم فعالية هياكل المحولات المختلفة، بما في ذلك BERT وBioBERT وRoBERTa، في استخراج المعلومات السريرية ذات الصلة. توضح المقدمة المساهمات الرئيسية للبحث، بما في ذلك تطوير نموذج يستخدم التعلم بالنقل لتحسين تحديد عوامل الخطر، وتطبيق استراتيجيات التجميع لتعزيز الأداء عبر عوامل الخطر المتعددة. تم هيكلة الأقسام اللاحقة من الورقة لتوفير نظرة شاملة على الأعمال ذات الصلة، والمنهجيات، والنتائج التجريبية، والاتجاهات المستقبلية في هذا المجال من الدراسة.

الطرق

في هذا القسم، يقدم المؤلفون تحليلًا شاملاً للنتائج التجريبية والمحاكاة التي أجريت باستخدام نموذجهم المقترح المعتمد على المتوسط المرجح، والذي يستفيد من هياكل المحولات المعتمدة على الضبط الدقيق. تتم مقارنة النتائج مع أداء الأنظمة الرائدة من مهمة i2b2 المشتركة لعام 2014، كما هو موضح في الجدول 4. يبرز هذا الجدول فعالية نماذج المحولات المعتمدة على الضبط الدقيق بالنسبة للأنظمة الأعلى تصنيفًا التي تستخدم مزيجًا من الأساليب المعتمدة على المعرفة والبيانات، بالإضافة إلى تلك التي تعتمد فقط على الأساليب المعتمدة على المعرفة، بما في ذلك المصنفات المعتمدة على القاموس والقواعد.

تشير النتائج إلى أن نماذج المحولات المعتمدة على الضبط الدقيق تتفوق على كلا الفئتين من الأنظمة الموجودة، مما يظهر مزايا دمج تقنيات التعلم الآلي المتقدمة في سياق المهمة. يبرز هذا المقارنة إمكانية النموذج المقترح في تعزيز الأداء في التطبيقات ذات الصلة، مما يشير إلى تقدم كبير في هذا المجال.

النتائج

في هذه الدراسة، تم استخدام التعلم بالنقل لتطوير نموذج للكشف عن عوامل خطر مرض القلب من النصوص السريرية، باستخدام مجموعة بيانات تحدي معالجة اللغة الطبيعية السريرية i2b2 لعام 2014. تم ضبط خمسة نماذج معتمدة على المحولات—BERT وBioBERT وRoBERTa وBioClinicalBERT وXLNet—لمهمة التصنيف، مع التركيز على تحديد الأمراض وعوامل الخطر والأدوية والسمات الزمنية. تم تطبيق تقنيات زيادة البيانات لتعزيز مجموعة بيانات التدريب، مع معالجة التمثيل الناقص لفئة “الجلوكوز” في تصنيف مؤشر السكري. شمل ذلك تكرار حالات الفئة الأقل وتمييع مجموعة البيانات لضمان تمثيل متوازن، مما حسن من قدرات التعلم والتعميم للنماذج.

أشارت النتائج إلى أن RoBERTa حقق أعلى مقاييس الأداء، مع مقياس F بنسبة 93.94%، ودقة بنسبة 93.90%، واسترجاع بنسبة 94.27%. كما أن BERT وBioBERT أديا بشكل جيد، خاصةً في تحديد عوامل الخطر المحددة مثل ارتفاع ضغط الدم وحالة التدخين. أسفر نهج التجميع، الذي جمع بين مخرجات النماذج الخمسة المعتمدة على المحولات، عن درجة F1 إجمالية بنسبة 94.26%، مما يظهر فعالية تجميع النماذج في تعزيز الكشف عن عوامل الخطر. تسلط الدراسة الضوء على إمكانية النماذج المعتمدة على المحولات وتقنيات زيادة البيانات في تحسين دقة تطبيقات معالجة اللغة الطبيعية السريرية، خاصةً في سياق تقييم خطر مرض القلب.

المناقشة

في هذا القسم، يناقش المؤلفون الدافع والسياق لبحثهم حول تحديد عوامل خطر مرض القلب، خاصةً في ضوء التحديات التي واجهت خلال مهمة i2b2/UTHealth المشتركة لعام 2014. يبرزون قيود الأساليب السابقة في استخراج المعلومات السريرية (IE) وإمكانية تقنيات التعلم بالنقل، خاصةً مع النماذج المعتمدة على المحولات مثل BERT وBioBERT وRoBERTa. تهدف الدراسة إلى تطوير نموذج عالي الأداء لتحديد عوامل خطر مرض القلب من خلال الاستفادة من هذه النماذج المتقدمة واستكشاف طرق مختلفة خاصة بالتضمين لتعزيز كفاءة التصنيف.

يقدم المؤلفون فرضيتهم بأن أساليب التعلم بالنقل يمكن أن تحسن بشكل كبير من التنبؤ بعوامل خطر مرض القلب. يقومون بتقييم منهجي لخمس نماذج معتمدة على المحولات، مما يظهر أن نموذج RoBERTa حقق أعلى درجة F1 الدقيقة بنسبة 94.27%. تشير النتائج إلى أن التعلم بالنقل لا يتفوق فقط على النماذج التقليدية ولكن أيضًا يحدد عوامل خطر جديدة. يعترف المؤلفون بالتحديات مثل عدم توازن الفئات وتعقيد السمات الزمنية في الملاحظات السريرية، والتي أثرت على أداء النموذج في بعض عوامل الخطر. يختتمون بالتأكيد على فعالية نهجهم ويحددون الأعمال المستقبلية التي تركز على تحسين تقنيات التضمين وتطبيق نموذجهم على مهام معالجة اللغة الطبيعية السريرية الأخرى.

Journal: Journal Of Big Data, Volume: 11, Issue: 1
DOI: https://doi.org/10.1186/s40537-024-00903-y
Publication Date: 2024-04-04
Author(s): Essam H. Houssein et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

This research investigates the application of transfer learning techniques, specifically transformer-based models, to enhance the identification of heart disease risk factors from Electronic Health Records (EHRs). The study emphasizes the importance of accurately predicting heart disease by detecting various risk factors, including diabetes, hypertension, and familial history. Given the challenges posed by the unstructured nature of EHRs, the authors employed models such as BERT, RoBERTa, BioClinicalBERT, XLNet, and BioBERT, which are pre-trained on extensive medical literature. These models were fine-tuned using the i2b2 dataset, resulting in superior performance compared to traditional methods in extracting relevant clinical information.

The findings reveal that the RoBERTa model achieved the highest micro F1-score of 94.27%, while the other models also demonstrated competitive results. An ensemble approach combining all five models yielded a micro F1-score of 94.26%, further illustrating the effectiveness of transfer learning in this domain. The study concludes that transformer-based models significantly improve the identification of heart disease risk factors and suggests future work will focus on addressing embedding-specific issues and applying fine-tuning to other clinical NLP tasks.

Introduction

The introduction highlights the significant public health challenge posed by heart disease, which remains a leading cause of mortality globally. With approximately 647,000 Americans affected annually and 17.8 million deaths worldwide attributed to heart disease in 2017, the urgency for effective prevention strategies is underscored. The text emphasizes the necessity for early identification of risk factors such as Coronary Artery Disease (CAD), diabetes, hypertension, and obesity, particularly given that many individuals who experience heart attacks do not exhibit traditional risk profiles. The complexity of extracting relevant risk factors from electronic health records (EHRs) is noted, particularly due to the variability and idiosyncratic nature of clinical narratives, which complicates automated prediction efforts.

The paper proposes leveraging transfer learning techniques, particularly through transformer-based models, to enhance the identification of heart disease risk factors from EHR data. By modeling this task as a Named Entity Recognition (NER) challenge, the study aims to evaluate the effectiveness of various transformer architectures, including BERT, BioBERT, and RoBERTa, in extracting relevant clinical information. The introduction outlines the primary contributions of the research, including the development of a model that utilizes transfer learning to improve risk factor identification, and the application of ensemble strategies to enhance performance across multiple risk factors. The subsequent sections of the paper are structured to provide a comprehensive overview of related work, methodologies, experimental results, and future directions in this area of study.

Methods

In this section, the authors present a comprehensive analysis of the experimental results and simulations conducted using their proposed weighted-averaged model, which leverages fine-tuned transformer-based architectures. The findings are compared against the performance of leading systems from the 2014 i2b2 shared task, as detailed in Table 4. This table highlights the effectiveness of the fine-tuned transformer models in relation to top-ranked systems that employ a combination of knowledge- and data-driven approaches, as well as those relying solely on knowledge-driven methods, including lexicon and rule-based classifiers.

The results indicate that the fine-tuned transformer-based models outperform both categories of existing systems, demonstrating the advantages of integrating advanced machine learning techniques in the context of the task. This comparison underscores the potential of the proposed model to enhance performance in relevant applications, suggesting a significant advancement in the field.

Results

In this study, transfer learning was employed to develop a model for detecting heart disease risk factors from clinical texts, utilizing the 2014 i2b2 clinical NLP challenge dataset. Five transformer-based models—BERT, BioBERT, RoBERTa, BioClinicalBERT, and XLNet—were fine-tuned for the classification task, focusing on identifying diseases, risk factors, medications, and temporal attributes. Data augmentation techniques were applied to enhance the training dataset, particularly addressing the under-representation of the ‘glucose’ class in the Diabetes-Indicator classification. This involved duplicating instances of the minority class and shuffling the dataset to ensure balanced representation, which improved the models’ learning and generalization capabilities.

The results indicated that RoBERTa achieved the highest performance metrics, with an F-measure of 93.94%, precision of 93.90%, and recall of 94.27%. BERT and BioBERT also performed well, particularly in identifying specific risk factors such as hypertension and smoking status. The ensemble approach, which combined the outputs of the five transformer models, yielded an overall F1-score of 94.26%, demonstrating the effectiveness of model ensembling in enhancing risk factor detection. The study highlights the potential of transformer-based models and data augmentation techniques in improving the accuracy of clinical NLP applications, particularly in the context of heart disease risk assessment.

Discussion

In this section, the authors discuss the motivation and context for their research on heart disease risk factor identification, particularly in light of the challenges faced during the 2014 i2b2/UTHealth shared task. They highlight the limitations of previous approaches in clinical Information Extraction (IE) and the potential of transfer learning techniques, particularly with transformer-based models such as BERT, BioBERT, and RoBERTa. The study aims to develop a high-performance model for identifying heart disease risk factors by leveraging these advanced models and exploring various embedding-specific methods to enhance classification efficiency.

The authors present their hypothesis that transfer learning methods can significantly improve the prediction of heart disease risk factors. They systematically evaluate five transformer-based models, demonstrating that the RoBERTa model achieved the highest micro F1-score of 94.27%. The findings indicate that transfer learning not only outperforms traditional models but also identifies novel risk factors. The authors acknowledge challenges such as class imbalance and the complexity of temporal attributes in clinical notes, which affected the model’s performance on certain risk factors. They conclude by emphasizing the effectiveness of their approach and outline future work focused on refining embedding techniques and applying their model to other clinical NLP tasks.