DOI: https://doi.org/10.1038/s41598-025-93783-y
PMID: https://pubmed.ncbi.nlm.nih.gov/40603531
تاريخ النشر: 2025-07-03
المؤلف: Yunus Balel وآخرون
الموضوع الرئيسي: الأشعة السينية السنية والتصوير
نظرة عامة
تقدم هذه الدراسة نموذج تعلم عميق مصمم للكشف التلقائي عن الأضراس الثالثة المتأثرة، مع دمج التصنيفات السريرية مثل تصنيف بيل وغريغوري، وتصنيف وينتر، ومؤشر صعوبة بيدرسون. أظهر النموذج دقة عالية ومقاييس أداء، محققًا نتائج قابلة للمقارنة مع طرق التقييم اليدوية التقليدية.
يمكن أن يؤدي تنفيذ هذا النظام المدعوم بالذكاء الاصطناعي في الإعدادات السريرية إلى توفير الوقت، وتبسيط سير العمل، وتقليل التباين في التقييمات. تشير الأداء القوي للنموذج عبر مجموعات بيانات متنوعة إلى قابليته للاستخدام على نطاق واسع. يمكن تحقيق تحسينات مستقبلية من خلال معالجة عدم التوازن في التصنيفات وتوسيع مجموعة البيانات، مما سيزيد من فعالية النموذج.
النتائج
في هذه الدراسة، تم تدريب نموذج تعلم عميق، والتحقق منه، واختباره باستخدام مجموعة بيانات من الأشعة السينية البانورامية من عدة مرضى، مع التركيز بشكل خاص على الأضراس الثالثة المتأثرة. تضمنت مجموعة البيانات 2,300 صورة للتدريب، و765 للتحقق، و765 للاختبار، مع إجمالي 7,624 ضرسًا متأثرًا مصنفًا في مجموعة التدريب، و2,580 في مجموعة التحقق، و2,493 في مجموعة الاختبار. تم تصنيف كل ضرس باستخدام نظام متعدد التسميات مع 98 تسمية فريدة بناءً على معايير موضعية وتصنيفية مختلفة، بما في ذلك تصنيف بيل وغريغوري ومؤشر صعوبة بيدرسون. استخدم النموذج، YOLOv11، آلية إيقاف مبكر تلقائية، وانتهى التدريب عند العصر 153، وحقق مقاييس أداء مثيرة للإعجاب: دقة متوسطة تبلغ 0.980، واسترجاع 0.948، ودرجة F1 تبلغ 0.974، وmAP@50 تبلغ 0.990، وmAP@50:95 تبلغ 0.974.
على الرغم من هذه المقاييس العالية، تم ملاحظة تحيزات محتملة في مجموعة البيانات، خاصة بسبب هيمنة التصنيفات الشائعة، مما قد يكون قد أثر على النتائج. أظهرت التسميات ذات التمثيل المنخفض، مثل 48-Distoangular-C-III و38-Horizontal-C-I، أداءً ضعيفًا، مما يشير إلى الحاجة إلى زيادة مجموعة البيانات لتحسين دقة التصنيف للفئات الممثلة بشكل غير كاف. كشفت مصفوفة الارتباك العادية عن تباينات كبيرة في أداء التنبؤ عبر التسميات، حيث حقق بعضها، مثل 48-Distoangular-C-II، درجة F1 تبلغ 0.000. وهذا يبرز القيود في شمولية مجموعة البيانات ويثير القلق بشأن قابلية تعميم النموذج على مجموعات بيانات أخرى. بشكل عام، بينما أظهر النموذج أداءً قويًا، تسلط النتائج الضوء على ضرورة معالجة عدم التوازن في مجموعة البيانات وتعزيز تمثيل الميزات لتحسين نتائج التصنيف.
المناقشة
تسلط قسم المناقشة في هذه الدراسة الضوء على التكامل المبتكر لعدة أنظمة تصنيف—بيل وغريغوري، وينتر، ومؤشر صعوبة بيدرسون—في إطار مدفوع بالذكاء الاصطناعي لتقييم الأضراس الثالثة المتأثرة. تتناول هذه الطريقة قيود التصنيف اليدوي التقليدي، الذي غالبًا ما يكون مستهلكًا للوقت وعرضة للتباين بسبب خبرة الممارس. من خلال استخدام نموذج التعلم العميق YOLOv11، تظهر الدراسة معدلات دقة عالية تتجاوز 95%، قابلة للمقارنة مع التقييمات البشرية، مما يعزز اتخاذ القرارات السريرية وقد يحسن سلامة المرضى.
تؤكد الدراسة على أهمية مجموعة بيانات كبيرة، تتكون من 3,830 صورة بانورامية، والتي تتجاوز تلك المستخدمة في الأبحاث السابقة، مما يساهم في قوة النموذج وقابلية تعميمه. كما أن استخدام تقنيات تعزيز البيانات المتقدمة أثناء التدريب يزيد من تنوع مجموعة البيانات، مما يحسن أداء النموذج على البيانات غير المرئية. ومع ذلك، يعترف المؤلفون بالقيود، بما في ذلك التحيزات المحتملة بسبب عدم توازن مجموعة البيانات والحاجة إلى التحقق الخارجي عبر مجموعات سكانية متنوعة. يجب أن تركز النسخ المستقبلية من النموذج على تعزيز قابلية التعميم، ومعالجة عدم التوازن في التصنيفات، ودمج آليات الذكاء الاصطناعي القابلة للتفسير لتعزيز الثقة في التطبيقات السريرية. بشكل عام، تشير النتائج إلى أن التصنيف المدعوم بالذكاء الاصطناعي يمكن أن يبسط سير العمل ويدعم الأطباء، خاصة في الإعدادات التي تعاني من نقص في الوصول إلى جراحي الفم ذوي الخبرة.
DOI: https://doi.org/10.1038/s41598-025-93783-y
PMID: https://pubmed.ncbi.nlm.nih.gov/40603531
Publication Date: 2025-07-03
Author(s): Yunus Balel et al.
Primary Topic: Dental Radiography and Imaging
Overview
This study presents a deep learning model designed for the automated detection of impacted third molars, incorporating clinical classifications such as the Pell and Gregory classification, Winter classification, and Pederson Difficulty Index. The model demonstrated high accuracy and performance metrics, achieving results comparable to traditional manual assessment methods.
The implementation of this AI-based decision support system in clinical settings has the potential to save time, streamline workflows, and minimize variability in assessments. The model’s robust performance across diverse datasets indicates its applicability for large-scale use. Future enhancements could be achieved by addressing classification imbalances and expanding the dataset, which would further improve the model’s effectiveness.
Results
In this study, a deep learning model was trained, validated, and tested using a dataset of panoramic radiographs from multiple patients, specifically focusing on impacted third molars. The dataset comprised 2,300 images for training, 765 for validation, and 765 for testing, with a total of 7,624 labeled impacted third molars in the training set, 2,580 in the validation set, and 2,493 in the test set. Each molar was classified using a multi-label system with 98 unique labels based on various positional and classification parameters, including the Pell and Gregory classification and the Pederson Difficulty Index. The model, YOLOv11, utilized an automatic early stopping mechanism, concluding training at the 153rd epoch, and achieved impressive performance metrics: an average precision of 0.980, recall of 0.948, F1 score of 0.974, mAP@50 of 0.990, and mAP@50:95 of 0.974.
Despite these high metrics, potential biases in the dataset were noted, particularly due to the predominance of common classifications, which may have skewed results. Labels with lower representation, such as 48-Distoangular-C-III and 38-Horizontal-C-I, demonstrated poorer performance, indicating a need for dataset augmentation to improve classification accuracy for underrepresented classes. The normalized Confusion Matrix revealed significant discrepancies in prediction performance across labels, with some, like 48-Distoangular-C-II, achieving an F1 score of 0.000. This underscores limitations in the dataset’s comprehensiveness and raises concerns about the model’s generalizability to other datasets. Overall, while the model showed strong performance, the findings highlight the necessity for addressing dataset imbalances and enhancing feature representation to improve classification outcomes.
Discussion
The discussion section of this study highlights the innovative integration of multiple classification systems—Pell and Gregory, Winter, and the Pederson Difficulty Index—into a single AI-driven framework for assessing impacted third molars. This approach addresses the limitations of traditional manual classification, which is often time-consuming and subject to variability due to practitioner experience. By employing the YOLOv11 deep learning model, the study demonstrates high accuracy rates exceeding 95%, comparable to human assessments, thereby enhancing clinical decision-making and potentially improving patient safety.
The study emphasizes the importance of a large dataset, comprising 3,830 panoramic radiographs, which surpasses those used in previous research, thereby contributing to the model’s robustness and generalizability. The use of advanced data augmentation techniques during training further diversifies the dataset, improving the model’s performance on unseen data. However, the authors acknowledge limitations, including potential biases due to dataset imbalances and the need for external validation across diverse populations. Future iterations of the model should focus on enhancing generalizability, addressing classification imbalances, and integrating explainable AI mechanisms to foster trust in clinical applications. Overall, the findings suggest that AI-assisted classification can streamline workflows and support clinicians, particularly in settings with limited access to experienced oral surgeons.
