DOI: https://doi.org/10.1038/s41598-024-60375-1
PMID: https://pubmed.ncbi.nlm.nih.gov/38678060
تاريخ النشر: 2024-04-27
المؤلف: Yan Wang وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة
نظرة عامة
تقدم الورقة نموذجًا جديدًا، MST-DeepLabv3+، مصممًا لتعزيز التقسيم الدلالي لصور الاستشعار عن بعد عالية الدقة مع معالجة التحديات المتعلقة بدقة النموذج وكفاءة التدريب. يبني هذا النموذج على بنية DeepLabv3+ الحالية، ويطبق ثلاث تحسينات رئيسية: استبدال MobileNetV2 بـ Xception في العمود الفقري لتقليل معلمات النموذج، ودمج آلية الانتباه SENet لتعزيز دقة التقسيم، واستخدام التعلم المنقول لتعزيز قدرات التعرف على الميزات. تم تقييم النموذج على مجموعات بيانات متعددة، بما في ذلك مجموعات بيانات صور ISPRS وGaofen، محققًا درجات متوسط التقاطع على الاتحاد (MIoU) تبلغ 82.47% و73.44%، على التوالي، وأظهر أداءً أعلى على مجموعة بيانات الأراضي المزروعة Taikang مع MIoU قدره 90.77%.
تشير النتائج إلى أن MST-DeepLabv3+ يحسن بشكل كبير من دقة التقسيم الدلالي مع تقليل حجم المعلمات، مما يسهل أوقات التدريب الأسرع. يبرز المؤلفون أهمية التقسيم الدلالي في الاستشعار عن بعد لمجموعة متنوعة من التطبيقات، بما في ذلك أبحاث التربة وإدارة الموارد الأرضية. يُقترح أن تركز الأعمال المستقبلية على معالجة قيود النموذج في استخراج معلومات الحدود وتعزيز قدراته على التعميم عبر التضاريس المتنوعة. بالإضافة إلى ذلك، يُقترح دمج المعلومات متعددة الطيف كمسار محتمل لتحسين دقة التقسيم بشكل أكبر.
الطرق
في هذا القسم، يوضح المؤلفون الطرق التجريبية والنتائج لتقييم أداء نموذج MST-DeepLabv3+ مقابل نماذج التقسيم الأخرى، بما في ذلك DeepLabv3+ وPSPNet وUNet، عبر مجموعات بيانات متنوعة. تم إجراء التجارب على نظام يعمل بنظام CentOS7.9 مع معالج AMD EPYC 7402 و عدة وحدات معالجة رسومات NVIDIA GeForce RTX 3090، باستخدام PyTorch 3.6. كانت المعلمات الرئيسية تشمل حجم دفعة قدره 8، و100 تكرار، ومعدل تعلم قدره 0.0005، ومُحسِّن Adam لضبط معدل التعلم الديناميكي. تشير النتائج إلى أن MST-DeepLabv3+ تفوقت باستمرار على النماذج الأخرى في مقاييس مثل متوسط التقاطع على الاتحاد (MIoU)، والدقة العامة (OA)، والدقة، والاسترجاع، ودرجة F1 عبر مجموعات بيانات ISPRS وGID وTaikang المزروعة.
على وجه التحديد، حقق MST-DeepLabv3+ MIoU قدره 82.47% على مجموعة بيانات ISPRS، متجاوزًا PSPNet وUNet وDeepLabv3+ بفارق كبير. كما أظهر قدرات تقسيم متفوقة، حيث حدد بدقة الأجسام الصغيرة وحافظ على حدود واضحة في المشاهد المعقدة. على مجموعة بيانات GID، سجل MST-DeepLabv3+ MIoU قدره 73.44%، متجاوزًا مرة أخرى النماذج الأخرى. في مجموعة بيانات الأراضي المزروعة Taikang، حقق MIoU قدره 90.77%، متجاوزًا أداء UNet بشكل كبير، مما أدى إلى تحسين كبير في التعرف على الحدود ودقة التقسيم. بشكل عام، تشير النتائج إلى أن MST-DeepLabv3+ يعالج بفعالية التحديات الشائعة في التقسيم، مثل التصنيف الخاطئ والحدود الخشنة، مما يؤدي إلى تحسين دقة التعرف عبر فئات متنوعة.
المناقشة
في هذا القسم، يناقش المؤلفون تطوير وتقييم نموذج MST-DeepLabv3+، المصمم لتعزيز التقسيم الدلالي لصور الاستشعار عن بعد عالية الدقة. يدمج النموذج شبكة عمود فقري خفيفة الوزن (MobileNetV2)، وآلية انتباه (SENet)، والتعلم المنقول لمعالجة التحديات مثل كفاءة التدريب المنخفضة والتعرف غير الدقيق على الأهداف. يقلل استخدام MobileNetV2 بشكل كبير من معلمات النموذج إلى 22.19 ميغابايت، مما يحسن سرعة التدريب مع الحفاظ على أداء التقسيم. تعوض إدخال SENet عن خسائر الدقة المرتبطة بالهيكل الخفيف الوزن، مما يؤدي إلى تحسين متوسط التقاطع على الاتحاد (MIoU) من 64.26% إلى 69.61% على مجموعة بيانات ISPRS.
يتم التحقق من أداء النموذج بشكل أكبر عبر مجموعات بيانات متعددة، محققًا درجات MIoU تبلغ 82.47% و73.44% و90.77% لمجموعات بيانات ISPRS وGID وTaikang المزروعة، على التوالي. تشير هذه النتائج إلى أن MST-DeepLabv3+ يعزز بفعالية دقة التقسيم مع الحفاظ على حجم نموذج مضغوط قدره 22.96 ميغابايت. يقترح المؤلفون أن تركز الأعمال المستقبلية على تحسين استخراج معلومات الحدود ودمج البيانات متعددة الطيف لتعزيز دقة التقسيم بشكل أكبر. بشكل عام، يظهر نموذج MST-DeepLabv3+ أداءً متفوقًا مقارنة بالنماذج الحالية، مع معالجة القضايا الرئيسية في تصنيف صور الاستشعار عن بعد.
DOI: https://doi.org/10.1038/s41598-024-60375-1
PMID: https://pubmed.ncbi.nlm.nih.gov/38678060
Publication Date: 2024-04-27
Author(s): Yan Wang et al.
Primary Topic: Advanced Neural Network Applications
Overview
The paper presents a novel model, MST-DeepLabv3+, designed to enhance the semantic segmentation of high-resolution remote sensing images while addressing challenges related to model precision and training efficiency. This model builds upon the existing DeepLabv3+ architecture, implementing three key improvements: substituting MobileNetV2 for Xception in the backbone to reduce model parameters, incorporating the SENet attention mechanism to enhance segmentation accuracy, and utilizing transfer learning to bolster feature recognition capabilities. The model was evaluated on multiple datasets, including the ISPRS and Gaofen image datasets, achieving mean intersection over union (MIoU) scores of 82.47% and 73.44%, respectively, and demonstrating even higher performance on the Taikang cultivated land dataset with a MIoU of 90.77%.
The findings indicate that MST-DeepLabv3+ significantly improves the accuracy of semantic segmentation while minimizing parameter size, thus facilitating faster training times. The authors highlight the importance of semantic segmentation in remote sensing for various applications, including soil research and land resource management. Future work is suggested to address the model’s limitations in boundary information extraction and to enhance its generalization capabilities across diverse terrains. Additionally, the incorporation of multispectral information is proposed as a potential avenue for further improving segmentation precision.
Methods
In this section, the authors detail the experimental methods and results for evaluating the performance of the MST-DeepLabv3+ model against other segmentation models, including DeepLabv3+, PSPNet, and UNet, across various datasets. The experiments were conducted on a system running CentOS7.9 with an AMD EPYC 7402 CPU and multiple NVIDIA GeForce RTX 3090 GPUs, utilizing PyTorch 3.6. Key parameters included a batch size of 8, 100 iterations, a learning rate of 0.0005, and the Adam optimizer for dynamic learning rate adjustment. The results indicate that MST-DeepLabv3+ consistently outperformed the other models in metrics such as Mean Intersection over Union (MIoU), Overall Accuracy (OA), Precision, Recall, and F1-score across the ISPRS, GID, and Taikang cultivated land datasets.
Specifically, MST-DeepLabv3+ achieved a MIoU of 82.47% on the ISPRS dataset, surpassing PSPNet, UNet, and DeepLabv3+ by significant margins. It also demonstrated superior segmentation capabilities, accurately identifying small objects and maintaining clear boundaries in complex scenes. On the GID dataset, MST-DeepLabv3+ recorded a MIoU of 73.44%, again outperforming the other models. In the Taikang cultivated land dataset, it achieved a MIoU of 90.77%, exceeding UNet’s performance and significantly improving boundary recognition and segmentation accuracy. Overall, the findings suggest that MST-DeepLabv3+ effectively addresses common segmentation challenges, such as misclassification and rough boundaries, leading to enhanced recognition accuracy across various categories.
Discussion
In this section, the authors discuss the development and evaluation of the MST-DeepLabv3+ model, designed to enhance semantic segmentation of high-resolution remote sensing images. The model integrates a lightweight backbone network (MobileNetV2), an attention mechanism (SENet), and transfer learning to address challenges such as low training efficiency and inaccurate target recognition. The use of MobileNetV2 significantly reduces model parameters to 22.19 MB, improving training speed while maintaining segmentation performance. The introduction of SENet compensates for accuracy losses associated with the lightweight architecture, resulting in a mean Intersection over Union (MIoU) improvement from 64.26% to 69.61% on the ISPRS dataset.
The model’s performance is further validated across multiple datasets, achieving MIoU scores of 82.47%, 73.44%, and 90.77% for the ISPRS, GID, and Taikang cultivated land datasets, respectively. These results indicate that MST-DeepLabv3+ effectively enhances segmentation accuracy while maintaining a compact model size of 22.96 MB. The authors suggest future work could focus on improving boundary information extraction and incorporating multispectral data to further enhance segmentation precision. Overall, the MST-DeepLabv3+ model demonstrates superior performance compared to existing models, addressing key issues in remote sensing image classification.
