MINet: شبكة تفاعلية متعددة المقاييس للكشف عن الكائنات البارزة في الوقت الحقيقي لعيوب سطح الفولاذ المدلفن
MINet: Multiscale Interactive Network for Real-Time Salient Object Detection of Strip Steel Surface Defects

المجلة: IEEE Transactions on Industrial Informatics، المجلد: 20، العدد: 5
DOI: https://doi.org/10.1109/tii.2024.3366221
تاريخ النشر: 2024-02-29
المؤلف: Kunye Shen وآخرون
الموضوع الرئيسي: الانتباه البصري واكتشاف الأهمية

نظرة عامة

تقدم ورقة البحث نهجًا جديدًا للكشف الآلي عن عيوب السطح في الإنتاج الصناعي، مع معالجة قيود الطرق الحالية المعتمدة على البروز، والتي غالبًا ما تتضمن أحجام معلمات كبيرة وتكاليف حسابية عالية. لتعزيز كفاءة الكشف، يقدم المؤلفون وحدة تفاعلية متعددة المقاييس (MI) تستخدم الالتفاف العمقي (DWConv) والالتفاف النقطي (PWConv) لاستخراج الميزات المستقلة والدمج التفاعلي عبر مقاييس مختلفة. تعتبر هذه الوحدة MI جزءًا أساسيًا من الشبكة التفاعلية متعددة المقاييس الخفيفة (MINet)، التي تحقق كشفًا في الوقت الحقيقي عن عيوب سطح الفولاذ المدلفن مع الحفاظ على دقة تنافسية.

تظهر النتائج التجريبية على مجموعة بيانات SD-Saliency-900، التي تتضمن صورًا لمختلف عيوب سطح الفولاذ المدلفن (الشوائب، البقع، والخدوش)، أن MINet تعمل بسرعات مثيرة للإعجاب تصل إلى 721 إطارًا في الثانية على وحدة معالجة الرسوميات و6.3 إطارًا في الثانية على وحدة المعالجة المركزية لصور بحجم 368×368، مع وجود 0.28 مليون معلمة فقط. تشير النتائج إلى أن MINet لا تقلل من التكاليف الحسابية فحسب، بل تعزز أيضًا الأداء في الوقت الحقيقي مقارنة بالطرق الحديثة. يخطط المؤلفون لتوسيع MINet لاستيعاب البيانات متعددة الأنماط واستكشاف تقنيات خفيفة مثل التقليم وإعادة المعلمات لتحقيق مزيد من التقدم في نماذج الكشف.

مقدمة

في مقدمة هذه الورقة البحثية، يؤكد المؤلفون على الدور الحاسم للكشف عن عيوب سطح الفولاذ في الإنتاج الصناعي، مشيرين إلى التأثير السلبي للعيوب مثل الشوائب، البقع، والخدوش على جمالية المنتج وموثوقيته. يلاحظون أن دقة الكشف عن العيوب تتأثر بعوامل متعددة، بما في ذلك بيئة الإنتاج المعقدة والحاجة إلى أداء نموذج فعال من حيث الحجم وسرعة الاستدلال. أدت التطورات الأخيرة في التعلم العميق، خاصة من خلال الشبكات العصبية الالتفافية (CNNs)، إلى تحسينات كبيرة في الكشف عن العيوب، حيث أظهرت النماذج المعتمدة على البروز نتائج واعدة. ومع ذلك، لا تزال التحديات قائمة، مثل التباين المنخفض بين العيوب والخلفيات، وأشكال العيوب وأحجامها المتنوعة، وظروف الإضاءة المنخفضة.

لمعالجة هذه التحديات، يقترح المؤلفون وحدة تفاعلية متعددة المقاييس (MI) جديدة تعزز الالتفاف القابل للفصل التقليدي (DSConv) من خلال دمج استراتيجية متعددة المقاييس. تستخدم هذه الوحدة الالتفاف العمقي والنقطي لاستخراج وتجميع الميزات متعددة المقاييس بشكل فعال مع الحفاظ على عبء حسابي منخفض. كما يقدم المؤلفون نموذج بروز خفيف الوزن، وهو الشبكة التفاعلية متعددة المقاييس (MINet)، المصممة للكشف عن العيوب في الوقت الحقيقي. تحقق MINet توازنًا بين دقة الكشف وكفاءة الحساب، مما يظهر أداءً مقارنًا مع النماذج الحديثة بينما تعمل بمعدلات إطارات عالية وتتطلب عددًا أقل من المعلمات. توضح الورقة مساهمات وحدة MI، وهيكل العمود الفقري في الوقت الحقيقي، والتحقق التجريبي من فعالية MINet في الكشف عن العيوب.

الطرق

في هذا القسم، يتم توضيح منهجية الدراسة، مع التركيز على التركيب والتكلفة الحسابية لهندسة DSConv المقترحة. يقدم المؤلفون ثلاث طرق لدمج الميزات متعددة المقاييس، مع تسليط الضوء على التحديات المرتبطة بالجمع والدمج أثناء عملية الدمج. ثم يوضحون وحدة MI والعمود الفقري في الوقت الحقيقي المعتمد على MI، مما يؤدي إلى تطوير هندسة MINet. يختتم القسم بوصف دالة الخسارة المستخدمة في عملية التدريب.

يتضمن الإعداد التجريبي تقييمًا شاملاً باستخدام مجموعة بيانات SD-Saliency-900، التي تتكون من 810 صور تدريب و900 صورة اختبار، كل منها تصور ثلاثة أنواع متميزة من العيوب في إنتاج الفولاذ المدلفن. لتعزيز مجموعة بيانات التدريب وتقليل الإفراط في التكيف، يقوم المؤلفون بتطبيق الانعكاس الأفقي، مما يؤدي إلى إجمالي 1620 صورة تدريب. يتم تغيير حجم الصور إلى 368 × 368 بكسل للاختبار، مع قص عشوائي إلى 336 × 336 بكسل أثناء التدريب. يتم تدريب MINet على جهاز كمبيوتر عالي الأداء، مع تفاصيل محددة عن تكوينات الأجهزة والبرامج، بما في ذلك استخدام مُحسِّن آدم ومعدل تعلم قدره $4 \times 10^{-3}$. تشمل معايير التقييم ثمانية مقاييس دقة (مثل متوسط الخطأ المطلق، والتقاطع على الاتحاد) وثلاث مقاييس كفاءة حسابية (حجم المعلمات، والعمليات العائمة، وسرعة الاستدلال)، مما يوفر تقييمًا شاملاً لأداء النموذج.

المناقشة

في مناقشة الكشف عن عيوب السطح، تبرز التطورات الأخيرة منهجيات متنوعة تهدف إلى تعزيز دقة وكفاءة الكشف. تشمل الأساليب البارزة استخدام الشبكات العصبية الانتباهية مع مجالات استقبالية متغيرة لتقليل ضوضاء الخلفية، بالإضافة إلى نماذج Faster R-CNN المحسنة التي تتضمن الالتفافات القابلة للتشويه ومحاذاة منطقة الاهتمام (RoI). بالإضافة إلى ذلك، ظهرت طرق تستخدم الكشف على مستوى البكسل وإطارات التكيف غير الخاضعة للإشراف، مما يعالج التحديات مثل أحجام العينات المحدودة في المجالات المستهدفة. كما اكتسب استكشاف تقنيات الكشف عن الأجسام البارزة زخمًا، مع اقتراح هياكل مثل الشبكات المشفرة-المفككة من طرف إلى طرف وشبكات إعادة استخدام الحواف لتحسين عمليات الكشف عن العيوب.

تؤكد الورقة أيضًا على أهمية الشبكات العصبية في الوقت الحقيقي، خاصة في السيناريوهات التي تعاني من قيود في الموارد الحسابية. تم تقديم ابتكارات مثل الشبكات الثنائية للتجزئة وشبكات هرمية انتباهية لتحقيق توازن بين أداء التجزئة وسرعة الاستدلال. تبني الوحدة التفاعلية متعددة المقاييس (MI) المقترحة على الالتفافات العمقية والنقطية لتعزيز استخراج الميزات والدمج، مما يحسن من توصيف العيوب مع الحفاظ على الكفاءة الحسابية. تم تصميم هيكل وحدة MI لتسهيل دمج الميزات متعددة المقاييس، مما يساهم في تطوير عمود فقري خفيف الوزن في الوقت الحقيقي للكشف عن العيوب يتفوق على الطرق الحالية من حيث السرعة والدقة.

Journal: IEEE Transactions on Industrial Informatics, Volume: 20, Issue: 5
DOI: https://doi.org/10.1109/tii.2024.3366221
Publication Date: 2024-02-29
Author(s): Kunye Shen et al.
Primary Topic: Visual Attention and Saliency Detection

Overview

The research paper presents a novel approach to automated surface defect detection in industrial production, addressing the limitations of existing saliency-based methods, which often involve large parameter sizes and high computational costs. To enhance detection efficiency, the authors introduce a multi-scale interactive (MI) module that utilizes depthwise convolution (DWConv) and pointwise convolution (PWConv) for independent feature extraction and interactive fusion across different scales. This MI module is integral to the proposed lightweight Multi-scale Interactive Network (MINet), which achieves real-time salient object detection of strip steel surface defects while maintaining competitive accuracy.

Experimental results on the SD-Saliency-900 dataset, which includes images of various strip steel surface defects (inclusions, patches, and scratches), demonstrate that MINet operates at impressive speeds of 721 FPS on GPU and 6.3 FPS on CPU for 368×368 images, with only 0.28 million parameters. The findings indicate that MINet not only reduces computational costs but also enhances real-time performance compared to state-of-the-art methods. The authors plan to extend MINet to accommodate multi-modal data and explore lightweight techniques such as pruning and re-parameterization for further advancements in detection models.

Introduction

In the introduction of this research paper, the authors emphasize the critical role of steel surface defect detection in industrial production, highlighting the negative impact of defects such as inclusions, patches, and scratches on product aesthetics and reliability. They note that the accuracy of defect detection is influenced by various factors, including the complex production environment and the need for efficient model performance in terms of size and inference speed. Recent advancements in deep learning, particularly through convolutional neural networks (CNNs), have led to significant improvements in defect detection, with saliency-based models showing promising results. However, challenges remain, such as low contrast between defects and backgrounds, varying defect shapes and scales, and low illumination conditions.

To address these challenges, the authors propose a novel Multi-scale Interactive (MI) module that enhances traditional depthwise separable convolution (DSConv) by integrating a multi-scale strategy. This module employs depthwise and pointwise convolutions to effectively extract and aggregate multi-scale features while maintaining low computational overhead. The authors also introduce a lightweight saliency model, the Multi-scale Interactive Network (MINet), designed for real-time defect detection. MINet achieves a balance between detection accuracy and computational efficiency, demonstrating comparable performance to state-of-the-art models while operating at high frame rates and requiring fewer parameters. The paper outlines the contributions of the MI module, the real-time backbone architecture, and the experimental validation of MINet’s effectiveness in defect detection.

Methods

In this section, the methodology of the study is outlined, focusing on the composition and computational cost of the proposed DSConv architecture. The authors introduce three multi-scale feature fusion methods, highlighting the challenges associated with summation and concatenation during the fusion process. They then detail the MI module and the MI-based real-time backbone, culminating in the development of the MINet architecture. The section concludes with a description of the loss function utilized in the training process.

The experimental setup involves a comprehensive evaluation using the SD-Saliency-900 dataset, which comprises 810 training images and 900 testing images, each depicting three distinct types of defects in strip steel production. To enhance the training dataset and mitigate overfitting, the authors apply horizontal flipping, resulting in a total of 1620 training images. The images are resized to 368 × 368 pixels for testing, with random cropping to 336 × 336 pixels during training. The MINet is trained on a high-performance PC, with specific hardware and software configurations detailed, including the use of the Adam optimizer and a learning rate of $4 \times 10^{-3}$. The evaluation criteria encompass eight accuracy metrics (e.g., Mean Absolute Error, Intersection over Union) and three computational efficiency metrics (parameter size, floating-point operations, and inference speed), providing a comprehensive assessment of the model’s performance.

Discussion

In the discussion of surface defect detection, recent advancements highlight various methodologies aimed at enhancing detection accuracy and efficiency. Notable approaches include the use of attention neural networks with variable receptive fields to suppress background noise, as well as improved Faster R-CNN models that incorporate deformable convolutions and Region-of-Interest (RoI) alignment. Additionally, methods employing pixel-wise detection and unsupervised domain adaptation frameworks have emerged, addressing challenges such as limited sample sizes in target domains. The exploration of salient object detection techniques has also gained traction, with architectures like end-to-end encoder-decoder networks and edge reuse networks being proposed to refine defect detection processes.

The paper further emphasizes the significance of real-time neural networks, particularly in scenarios with constrained computational resources. Innovations such as bilateral segmentation networks and attention pyramid networks have been introduced to balance segmentation performance with inference speed. The proposed multi-scale interactive (MI) module builds on depthwise and pointwise convolutions to enhance feature extraction and fusion, thereby improving the characterization of defects while maintaining computational efficiency. The MI module’s architecture is designed to facilitate multi-scale feature integration, ultimately contributing to the development of a lightweight, real-time backbone for defect detection that outperforms existing methods in terms of speed and accuracy.