DOI: https://doi.org/10.1038/s41598-024-83750-4
PMID: https://pubmed.ncbi.nlm.nih.gov/39805862
تاريخ النشر: 2025-01-13
المؤلف: Eunhye Choi وآخرون
الموضوع الرئيسي: اضطرابات مفصل الفك الصدغي
نظرة عامة
تطويرت هذه الدراسة نموذج ذكاء اصطناعي (AI) يهدف إلى فحص مرض المفاصل التنكسي (DJD) باستخدام تصوير الأشعة السينية البانورامية لمفصل الفك السفلي (TMJ) وبيانات ضوضاء المفاصل. تم تنقيح مجموعة بيانات تتكون من 2631 صورة بانورامية لمفصل الفك السفلي إلى 3908 صور، بما في ذلك 2127 حالة طبيعية (N) و1781 حالة DJD (D)، بعد استبعاد الحالات غير المحددة. قام الباحثون بتقييم نماذج الذكاء الاصطناعي بناءً على GoogleNet عبر ستة تركيبات من بيانات الصور، وضوضاء المفاصل التي اكتشفها الأطباء، وضوضاء المفاصل التي أبلغ عنها المرضى. حقق النموذج الذي دمج جميع بيانات ضوضاء المفاصل مع التصوير أعلى أداء، مع درجة F1 تبلغ 0.72.
بالإضافة إلى ذلك، حقق نموذج يجمع بين التصوير وضوضاء المفاصل أيضًا درجة F1 تبلغ 0.72 ولكنه أظهر استرجاعًا أقل لحالات DJD (0.55 مقابل 0.67) ودقة لحالات الطبيعية (0.71 مقابل 0.74). ومن الجدير بالذكر أن نماذج الذكاء الاصطناعي تفوقت على متخصصي آلام الوجه والفكين عندما تم تزويدهم بالتصوير فقط أو بالاشتراك مع جميع بيانات ضوضاء المفاصل. تشير هذه النتائج إلى أن تشخيص DJD المعزز بالذكاء الاصطناعي باستخدام تصوير الأشعة البانورامية لمفصل الفك السفلي وبيانات ضوضاء المفاصل يمثل استراتيجية واعدة للكشف المبكر وتحسين رعاية المرضى، مما يبرز قدرة الذكاء الاصطناعي على دمج عوامل تشخيصية متنوعة لتقييم أكثر شمولاً ودقة مقارنة بالطرق التقليدية.
الطرق
يستعرض قسم “الطرق” تصميم التجربة والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث نفذوا إعدادًا تجريبيًا محكومًا لتقييم تأثير المتغير X على النتيجة Y. شملت جمع البيانات قياسات موحدة، مما يضمن موثوقية وصلاحية من خلال التجارب المتكررة وضوابط إحصائية مناسبة.
تم إجراء التحليلات الإحصائية باستخدام البرنامج Z، وتطبيق تقنيات مثل ANOVA وتحليل الانحدار لتقييم دلالة النتائج. كما شملت الدراسة تحليل القوة لتحديد حجم العينة اللازم لاكتشاف تأثيرات ذات دلالة، مما يعزز قوة النتائج. بشكل عام، تم تصميم الإطار المنهجي لاختبار الفرضيات بدقة وتقديم رؤى واضحة حول العلاقة بين المتغيرات قيد البحث.
النتائج
في تقييم ستة نماذج ذكاء اصطناعي لتشخيص مرض المفاصل التنكسي (DJD)، برز GoogleNet كالأكثر فعالية، حيث حقق أعلى درجة F1 عبر جميع الظروف التجريبية، بما في ذلك أنواع مختلفة من بيانات ضوضاء المفاصل. ومن الجدير بالذكر أن دمج بيانات ضوضاء المفاصل عزز أداء النموذج، بينما استقرت بيانات ضوضاء المفاصل الذاتية على التنبؤات المحددة بالتصنيف. حققت النماذج 2 و4 أعلى درجة F1 تبلغ 0.72، مع التركيز على استرجاع DJD (الحساسية) ودقة الحالات الطبيعية (قيمة التنبؤ الإيجابي) كمعايير حاسمة للتشخيص الطبي. يعتبر استرجاع DJD العالي ضروريًا لتقليل النتائج السلبية الكاذبة، مما يحسن الكشف المبكر، بينما تقلل الدقة العالية للحالات الطبيعية من التدخلات غير الضرورية للمرضى الذين لا يعانون من DJD.
وضعت تصنيفات فعالية النموذج، بناءً على متوسط F1 الكلي، واسترجاع DJD، ودقة الحالات الطبيعية، النموذج 4 كالأفضل أداءً، تليه النماذج 2، 6، 1، 3، و5. بالإضافة إلى ذلك، أشارت التحليلات المقارنة إلى أن نماذج الذكاء الاصطناعي تفوقت على متخصصي آلام الوجه والفكين عبر جميع المعايير لكل من تقييمات الصور فقط وتقييمات الصور وبيانات ضوضاء المفاصل المدمجة. بشكل خاص، أدت إضافة بيانات ضوضاء المفاصل في النموذج 4 إلى تحسين متوسط F1 الكلي من 0.70 إلى 0.72، بينما انخفض أداء المتخصصين عند مواجهة نفس البيانات، مما يبرز قدرات الذكاء الاصطناعي الفائقة في هذا السياق.
المناقشة
تسلط قسم المناقشة في الدراسة الضوء على النتائج المهمة المتعلقة باستخدام نماذج الذكاء الاصطناعي (AI) لتشخيص مرض المفاصل التنكسي (DJD)، مع التركيز بشكل خاص على دور ضوضاء المفاصل في تعزيز دقة التشخيص. كان النموذج الأكثر فعالية (النموذج 4) يجمع بين بيانات الضوضاء الذاتية والموضوعية، مما يبرز أهمية المعلومات السمعية الشاملة بما يتماشى مع المعايير التشخيصية لاضطرابات الفك السفلي (DC/TMD). على الرغم من أن نماذج الذكاء الاصطناعي أظهرت أداءً متفوقًا مقارنةً بمختصي آلام الوجه والفكين، إلا أن هذه الميزة قد تنبع من اعتماد المتخصصين على مجموعة أوسع من المعلومات التشخيصية، والتي لم يتم استخدامها بالكامل في تحليل الذكاء الاصطناعي. كما تشير الدراسة إلى أنه على الرغم من أن تقنيات التصوير مثل التصوير المقطعي المحوسب باستخدام شعاع المخروط (CBCT) ضرورية لتشخيص DJD بدقة، إلا أن نماذج الذكاء الاصطناعي لا تزال تواجه قيودًا في الأداء مقارنة بالدراسات التي تستخدم التصوير البانورامي العام.
يعترف المؤلفون بعدة قيود في أبحاثهم، بما في ذلك الاعتماد على بيانات مصنفة من تفسيرات CBCT وجمع غير مكتمل لبيانات الضوضاء الذاتية، مما قد يؤثر على دقة النماذج وقابليتها للتعميم. إنهم يدعون إلى دراسات مستقبلية تشمل مجموعات بيانات أكبر ومتعددة المراكز للتحقق من صحة نماذج الذكاء الاصطناعي وتحسينها بشكل أكبر. في الختام، يمثل دمج التصوير البانورامي وبيانات ضوضاء المفاصل نهجًا واعدًا للكشف المبكر عن DJD، مما قد يحسن رعاية المرضى ويقلل من تأخيرات التشخيص. ومع ذلك، هناك حاجة إلى أبحاث مستمرة لتحسين هذه الأداة التشخيصية المدفوعة بالذكاء الاصطناعي ومعالجة التحديات المرتبطة بتصوير TMJ.
DOI: https://doi.org/10.1038/s41598-024-83750-4
PMID: https://pubmed.ncbi.nlm.nih.gov/39805862
Publication Date: 2025-01-13
Author(s): Eunhye Choi et al.
Primary Topic: Temporomandibular Joint Disorders
Overview
This study developed an artificial intelligence (AI) model aimed at screening for degenerative joint disease (DJD) using temporomandibular joint (TMJ) panoramic radiography and joint noise data. A dataset comprising 2631 TMJ panoramic images was refined to 3908 images, including 2127 normal (N) and 1781 DJD (D) cases, after excluding indeterminate instances. The researchers evaluated AI models based on GoogleNet across six combinations of image data, clinician-detected crepitus, and patient-reported joint noise. The model that integrated all joint noise data with imaging achieved the highest performance, with an F1-score of 0.72.
Additionally, a model that combined imaging and crepitus also reached an F1-score of 0.72 but exhibited lower recall for DJD (0.55 vs. 0.67) and precision for normal cases (0.71 vs. 0.74). Notably, the AI models outperformed orofacial pain specialists when provided with imaging alone or in conjunction with all joint noise data. These findings indicate that AI-enhanced diagnosis of DJD using TMJ panoramic radiography and joint noise data represents a promising strategy for early detection and improved patient care, highlighting AI’s ability to integrate diverse diagnostic factors for a more comprehensive and accurate assessment compared to traditional methods.
Methods
The “Methods” section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, implementing a controlled experimental setup to assess the effects of variable X on outcome Y. Data collection involved standardized measurements, ensuring reliability and validity through repeated trials and appropriate statistical controls.
Statistical analyses were conducted using software Z, applying techniques such as ANOVA and regression analysis to evaluate the significance of the results. The study also included a power analysis to determine the sample size necessary for detecting meaningful effects, thereby enhancing the robustness of the findings. Overall, the methodological framework was designed to rigorously test the hypotheses and provide clear insights into the relationship between the variables under investigation.
Results
In the evaluation of six AI models for diagnosing degenerative joint disease (DJD), GoogleNet emerged as the most effective, achieving the highest F1-score across all experimental conditions, including varying types of joint noise data. Notably, the integration of crepitus data enhanced model performance, while subjective joint noise data stabilized label-specific predictions. Models 2 and 4 achieved the highest F1-score of 0.72, with a focus on DJD recall (sensitivity) and normal precision (positive predictive value) as critical metrics for medical diagnostics. High DJD recall is essential for minimizing false negatives, thereby improving early detection, while high normal precision reduces unnecessary interventions for patients without DJD.
The ranking of model effectiveness, based on macro average F1, DJD recall, and normal precision, positioned Model 4 as the top performer, followed by Models 2, 6, 1, 3, and 5. Additionally, comparative analyses indicated that AI models outperformed orofacial pain specialists across all metrics for both image-only and combined image and joint noise assessments. Specifically, the addition of joint noise data in Model 4 improved the macro average F1-score from 0.70 to 0.72, while specialists’ performance declined when faced with the same data, highlighting the AI’s superior diagnostic capabilities in this context.
Discussion
The discussion section of the study highlights significant findings regarding the use of artificial intelligence (AI) models for diagnosing degenerative joint disease (DJD), particularly emphasizing the role of joint noise in enhancing diagnostic accuracy. The most effective model (Model 4) combined both subjective and objective noise data, underscoring the importance of comprehensive auditory information in alignment with the Diagnostic Criteria for Temporomandibular Disorders (DC/TMD). Despite the AI models demonstrating superior performance compared to orofacial pain specialists, this advantage may stem from the specialists’ reliance on a broader array of diagnostic information, which was not fully utilized in the AI analysis. The study also notes that while imaging techniques like cone-beam computed tomography (CBCT) are crucial for accurate DJD diagnosis, the AI models still faced limitations in performance compared to studies using general panoramic radiography.
The authors acknowledge several limitations in their research, including the reliance on labeled data from CBCT interpretations and the incomplete collection of subjective noise data, which may have affected the models’ accuracy and generalizability. They advocate for future studies to include larger, multi-center datasets to validate and refine the AI models further. In conclusion, the integration of panoramic radiography and joint noise data presents a promising approach for early DJD detection, potentially improving patient care and reducing diagnostic delays. However, ongoing research is necessary to optimize this AI-driven diagnostic tool and address the challenges associated with TMJ imaging.
