آلية الانتباه الهرمية المدمجة مع الشبكات العصبية العميقة للتجزئة الدلالية الدقيقة للهياكل السنية في الأشعة السينية البانورامية
Hierarchical attention mechanism combined with deep neural networks for accurate semantic segmentation of dental structures in panoramic radiographs

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-22560-8
PMID: https://pubmed.ncbi.nlm.nih.gov/41193542
تاريخ النشر: 2025-11-05
المؤلف: Mehrdad Esmaeili وآخرون
الموضوع الرئيسي: الأشعة السينية السنية والتصوير

نظرة عامة

تقدم ورقة البحث نموذجًا جديدًا قائمًا على التعلم العميق لتقسيم الصور الدلالي يهدف إلى تعزيز تفسير الأشعة السينية البانورامية في طب الأسنان. تم تصميم هذا النموذج لتحديد وتصنيف 24 هيكلًا تشريحيًا ومرضياً بدقة، مما يعالج تحديات التفسيرات اليدوية التي تستغرق وقتًا طويلاً وتكون عرضة للأخطاء. باستخدام مجموعة بيانات تتكون من 844 صورة موضحة، يستخدم النموذج آلية انتباه متعددة المقاييس هرمية، مما يسمح بتحسين الدقة من خلال تحليل الصور بدقات مختلفة. تظهر النتائج مقاييس أداء مثيرة للإعجاب، بما في ذلك دقة عامة تبلغ 98.73%، وخصوصية 98.86%، وقيمة تقاطع على اتحاد (IoU) تبلغ 78.76%، ودقة واسترجاع كلاهما عند 86.97%، ودرجة F1 تبلغ 84.54%. بينما يتفوق النموذج في تحديد هياكل مثل الغرسات والترميمات، يواجه صعوبات مع الميزات الدقيقة مثل لب الأسنان والتسوس بسبب تداخل الهياكل.

تخلص الدراسة إلى أن دمج الذكاء الاصطناعي في تشخيصات الأسنان يمكن أن يعزز بشكل كبير الدقة والكفاءة. إن النهج الهرمي متعدد المقاييس للنموذج المقترح يوازن بشكل فعال بين التفاصيل الدقيقة والسياق الأوسع، مما يجعله قابلاً للتكيف مع السيناريوهات السريرية الواقعية. على الرغم من نجاحاته، فإن قيود النموذج في اكتشاف بعض الهياكل المعقدة تبرز ضرورة استخدام تقنيات تصوير مكملة. ستركز الأبحاث المستقبلية على الاستفادة من مخرجات التقسيم لتطوير وحدات ذكاء اصطناعي إضافية تعالج تحديات سريرية محددة، مثل تقييم عمق التسوس وتقييم حشوات قنوات الجذر. بالإضافة إلى ذلك، سيكون استكشاف الهياكل المعتمدة على المحولات وتوسيع مجموعة البيانات أمرًا حاسمًا لتحسين قابلية تعميم النموذج وأدائه في سياقات سريرية متنوعة، مما يدعم في النهاية الأطباء في تخطيط العلاج وتحسين نتائج المرضى.

طرق

تستخدم المنهجية المعتمدة في هذا البحث آلية انتباه متعددة المقاييس هرمية مدمجة مع هيكلين شبكيين متميزين: ResNet-50 و HRNet-OCR. تم هيكلة المعمارية لتشمل شبكة أساسية مخصصة لاستخراج الميزات، مكملة برؤوس توقع دلالية ورؤوس انتباه. يسهل هذا التصميم معالجة فعالة للبيانات المدخلة لتعزيز دقة التقسيم.

تم تنفيذ خطوات ما بعد المعالجة اللاحقة لتنقيح المخرجات، مما ينتج في النهاية خريطة تقسيم عالية الدقة. يتم توضيح مزيد من التفاصيل حول هذه المكونات وتفاعلاتها في الأقسام التالية من الورقة.

نتائج

أظهر النموذج المقترح القائم على آلية انتباه متعددة المقاييس هرمية لتقسيم الصور الأشعة السينية البانورامية أداءً استثنائيًا في تحديد وتحديد مختلف الهياكل والآفات السنية. تم تقييم النموذج على مجموعة اختبار، حيث حقق دقة متوسطة تبلغ 98.73%، وخصوصية 98.86%، وقيمة تقاطع على اتحاد (IoU) تبلغ 78.76%، ودقة واسترجاع تبلغ 86.97%، ودرجة F1 تبلغ 84.54% عبر التحقق المتقاطع ذو 5 طيات. تشير فترات الثقة بنسبة 95% لهذه المقاييس إلى أداء قوي ومتسق. ومن الجدير بالذكر أن بعض الفئات، مثل الأنياب السفلية والأضراس العلوية، أظهرت دقة تتجاوز 98%، وذلك بسبب وضوحها التشريحي المتميز. ومع ذلك، تم ملاحظة تحديات في تقسيم الهياكل الدقيقة مثل لب الأسنان واكتشاف الميزات الدقيقة مثل التسوس والترميمات المركبة، مما أدى إلى قيم دقة واسترجاع أقل.

أظهرت الأمثلة المرئية لأداء النموذج كفاءته في تقسيم الهياكل المحددة جيدًا، بينما سلطت الضوء أيضًا على الصعوبات في التقاط التفاصيل الدقيقة. تضمنت معلمات النموذج معدل تعلم قدره 0.005، وحجم دفعة قدره 1، و150 دورة، مع إجراء التدريب على 492 صورة. على الرغم من تعقيد التعامل مع 24 فئة متميزة – أكثر من معظم الدراسات الحالية – حقق النموذج نتائج تنافسية، مع درجة F1 قابلة للمقارنة مع نماذج أخرى مثل Mask R-CNN و U-Net، التي تم تدريبها على عدد أقل من الفئات. وهذا يشير إلى إمكانية النموذج للتطبيقات السريرية في العالم الحقيقي، بينما يشير أيضًا إلى الحاجة إلى تقنيات تصوير مكملة لتحسين اكتشاف الميزات الدقيقة.

مناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على أهمية نموذج الانتباه الهرمي متعدد المقاييس الذي تم تطويره لتقسيم الصور الأشعة السينية البانورامية في التطبيقات السنية. تتكون مجموعة البيانات، التي تضم 844 صورة تم جمعها من مراكز متعددة، من توضيحات دقيقة من قبل خبراء الأسنان، مما يضمن بيانات حقيقة عالية الجودة. تدمج معمارية النموذج شبكة أساسية (في البداية ResNet-50، ثم تم ترقيتها إلى HRNet-OCR) مع رؤوس توقع متقدمة لمخرجات دلالية ورؤوس انتباه. تعزز آلية الانتباه الهرمي دقة التقسيم من خلال تعلم ديناميكي للاهتمام النسبي بين المقاييس المجاورة، مما يجمع بشكل فعال بين السياق العالمي والتفاصيل المحلية دون زيادة الحمل الحسابي.

تظهر النتائج أن النموذج المقترح يتفوق على المعماريات التقليدية، مثل U-Net، حيث حقق درجة F1 تبلغ 84.54%. تؤكد دراسة الإزالة أن آلية الانتباه الهرمي ضرورية لتحسين دقة التقسيم، خاصة في تحديد الهياكل السنية المعقدة. تعزز قدرة النموذج على التعميم عبر خصائص تصوير متنوعة من أجهزة الأشعة السينية المتعددة قوته للتطبيقات السريرية. ومع ذلك، لا تزال هناك تحديات في تقسيم الهياكل المتداخلة والميزات الدقيقة مثل لب الأسنان، مما يشير إلى مجالات للبحث المستقبلي. بشكل عام، تؤسس الدراسة نموذج تقسيم أساسي يمكن أن يدعم تحليلات تشخيصية أكثر تقدمًا في الأشعة السينية السنية.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-22560-8
PMID: https://pubmed.ncbi.nlm.nih.gov/41193542
Publication Date: 2025-11-05
Author(s): Mehrdad Esmaeili et al.
Primary Topic: Dental Radiography and Imaging

Overview

The research paper presents a novel deep learning-based semantic segmentation model aimed at enhancing the interpretation of panoramic radiographs in dentistry. This model is designed to accurately identify and classify 24 anatomical and pathological structures, addressing the challenges of time-consuming and error-prone manual interpretations. Utilizing a dataset of 844 annotated images, the model employs a hierarchical multi-scale attention mechanism, which allows for improved accuracy by analyzing images at various resolutions. The results demonstrate impressive performance metrics, including an overall accuracy of 98.73%, specificity of 98.86%, Intersection over Union (IoU) of 78.76%, precision and recall both at 86.97%, and an F1 score of 84.54%. While the model excels in identifying structures like implants and restorations, it faces difficulties with subtle features such as dental pulp and caries due to overlapping structures.

The study concludes that integrating AI into dental diagnostics can significantly enhance accuracy and efficiency. The proposed model’s hierarchical multi-scale approach effectively balances fine details with broader context, making it adaptable to real-world clinical scenarios. Despite its successes, the model’s limitations in detecting certain intricate structures highlight the necessity for complementary imaging techniques. Future research will focus on leveraging the segmentation outputs to develop additional AI modules that address specific clinical challenges, such as assessing caries depth and evaluating root canal fillings. Additionally, exploring transformer-based architectures and expanding the dataset will be crucial for improving the model’s generalizability and performance in diverse clinical contexts, ultimately supporting clinicians in treatment planning and improving patient outcomes.

Methods

The methodology employed in this research utilizes a hierarchical multi-scale attention mechanism integrated with two distinct network architectures: ResNet-50 and HRNet-OCR. The architecture is structured to include a backbone network dedicated to feature extraction, complemented by semantic and attention prediction heads. This design facilitates the effective processing of input data to enhance segmentation accuracy.

Subsequent post-processing steps are implemented to refine the output, ultimately producing a high-resolution segmentation map. Further details regarding these components and their interactions are elaborated in the subsequent sections of the paper.

Results

The proposed hierarchical multi-scale attention model for semantic segmentation of panoramic radiographic images demonstrated exceptional performance in identifying and delineating various dental structures and pathologies. Evaluated on a test set, the model achieved an average accuracy of 98.73%, specificity of 98.86%, an Intersection over Union (IoU) value of 78.76%, precision and recall of 86.97%, and an F1 score of 84.54% across 5-fold cross-validation. The 95% confidence intervals for these metrics indicate robust and consistent performance. Notably, certain classes, such as lower canines and upper molars, exhibited accuracy exceeding 98%, attributed to their distinct anatomical visibility. However, challenges were noted in segmenting fine structures like dental pulp and detecting subtle features such as caries and composite restorations, which yielded lower precision and recall values.

Visual examples of the model’s performance illustrated its proficiency in segmenting well-defined structures, while also highlighting difficulties in capturing finer details. The model’s hyperparameters included a learning rate of 0.005, a batch size of 1, and 150 epochs, with training conducted on 492 images. Despite the complexity of handling 24 distinct classes—more than most existing studies—the model achieved competitive results, with an F1 score comparable to other models like Mask R-CNN and U-Net, which were trained on fewer classes. This suggests the model’s potential for real-world clinical applications, while also indicating the need for complementary imaging techniques for improved detection of subtle features.

Discussion

The discussion section of the research paper highlights the significance of the hierarchical multi-scale attention model developed for semantic segmentation of panoramic radiographic images in dental applications. The dataset, comprising 844 images collected from multiple centers, was meticulously annotated by dental experts, ensuring high-quality ground truth data. The model’s architecture integrates a backbone network (initially ResNet-50, later upgraded to HRNet-OCR) with advanced prediction heads for semantic and attention outputs. The hierarchical attention mechanism enhances segmentation accuracy by dynamically learning relative attention between adjacent scales, effectively combining global context and local detail without increasing computational load.

The results demonstrate that the proposed model outperforms traditional architectures, such as U-Net, achieving an F1-score of 84.54%. The ablation study confirms that the hierarchical attention mechanism is crucial for improving segmentation precision, particularly in delineating complex dental structures. The model’s ability to generalize across diverse imaging characteristics from multiple radiographic devices enhances its robustness for clinical applications. However, challenges remain in accurately segmenting overlapping structures and subtle features like dental pulp, indicating areas for future research. Overall, the study establishes a foundational segmentation model that can support more advanced diagnostic analyses in dental radiology.

شارك: