DOI: https://doi.org/10.32604/cmes.2024.051083
تاريخ النشر: 2024-01-01
المؤلف: Lirong Yin وآخرون
الموضوع الرئيسي: الشبكات العصبية والتطبيقات
نظرة عامة
تقدم هذه الدراسة شبكة استخراج ميزات الصور القائمة على المحولات الفعالة (CEFormer)، والتي تعالج قيود نماذج المحولات التقليدية في استخراج ميزات الصور، وخاصة نقص التحيز الاستقرائي للهياكل البصرية. يدمج CEFormer الانتباه الإلكتروني، والتفافات قابلة للفصل حسب العمق، والتفافات موسعة لتقديم تحيزات استقرائية أساسية مثل عدم التغير تحت الترجمة، والمحلية، وعدم التغير تحت المقياس. تعزز هذه البنية المبتكرة الاستقرار وتسارع التقارب، مما يؤدي إلى تحسين الأداء في مهام تصنيف الصور. تظهر النتائج التجريبية على مجموعة بيانات ImageNet1k Top-1 أن CEFormer يحقق دقة تصل إلى 85.0%، متجاوزًا نماذج الأساس المختلفة ويظهر أداءً متفوقًا عند دمجه في إطار عمل شبكة الأعصاب الالتفافية الإقليمية (R-CNN).
تشير النتائج إلى أن CEFormer يتفوق باستمرار على كل من النماذج القائمة على الالتفاف والنماذج النقية للمحول، محققًا دقة أعلى تصل إلى 4.5% مقارنة بنماذج مثل ResNet وRegNetY بينما يتطلب عددًا أقل من العمليات العائمة (FLOPs) مقارنة بأحدث المحولات البصرية. تكشف دراسة الإزالة أنه بينما تساهم جميع الخصائص الأربع—عدم التغير تحت الترجمة، والمحلية، وعدم التغير تحت المقياس، والاستقرار—في أداء النموذج، فإن الاستقرار هو العامل الأكثر أهمية الذي يؤثر على الدقة. كما تلاحظ الدراسة أن دمج وحدات التفاف متعددة يزيد من عدد معلمات النموذج وتعقيد الوقت، مما يشير إلى أن الأبحاث المستقبلية يجب أن تستكشف تقنيات مثل التقليم وتقطير المعرفة لتعزيز كفاءة النموذج دون التضحية بالأداء.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الدور المهم لرؤية الكمبيوتر ضمن مجال التعلم العميق، مع التركيز على تطبيقاتها في مهام مثل التعرف على الوجه، وتتبع الأجسام، واكتشاف المشاة. يبرز النمو السريع لبيانات الصور، المدعوم بالتقدم في البيانات الضخمة وأداء الكمبيوتر، أهمية طرق استخراج الميزات الفعالة لتعزيز دقة وكفاءة المهام البصرية. بينما كانت الشبكات العصبية الالتفافية (CNNs) هي الأداة الرئيسية لاستخراج الميزات، فإنها تواجه قيودًا في التقاط السياق العالمي ونمذجة الاعتماد بين الميزات، مما دفع لاستكشاف هياكل المحولات التي تم تطويرها في الأصل لمعالجة اللغة الطبيعية.
تقدم المحولات مزايا في نمذجة الاعتماد على المدى الطويل والمعالجة المتوازية من خلال آلية الانتباه الذاتي، مما يسمح بتحليل أكثر شمولية للصور مقارنةً بالشبكات العصبية الالتفافية. ومع ذلك، تحدد الورقة التحديات الحرجة المرتبطة بالنماذج القائمة على المحولات، بما في ذلك تعقيد الوقت والمساحة التربيعيين المرتبطين بتسلسلات رموز الإدخال، مما يقيد تطبيقها على الصور عالية الدقة والأجهزة الطرفية. بالإضافة إلى ذلك، تتطلب المحولات مجموعات بيانات تدريب مسبق واسعة وتظهر حساسية لاختيار المعلمات الفائقة، مما يؤدي غالبًا إلى تقارب أبطأ وأداء أقل قليلاً مقارنةً بالشبكات العصبية الالتفافية. لمعالجة هذه القضايا، يقترح المؤلفون شبكة استخراج ميزات الصور القائمة على المحولات الفعالة (CEFormer)، التي تدمج استراتيجيات الالتفاف لتعزيز عدم التغير تحت الترجمة، والمحلية، وعدم التغير تحت المقياس، مما يحسن سرعة التقارب واستقرار النموذج في مهام معالجة الصور.
طرق
في قسم الطرق، يوضح المؤلفون البيئة التجريبية وإعدادات المعلمات الفائقة المستخدمة في دراستهم. يتم تلخيص تكوينات البرمجيات والأجهزة في الجدول 1، بينما يقدم الجدول 2 إعدادات النموذج لأحجام مختلفة. يتم تعريف المعلمات الفائقة بدقة: تمتد عملية التدريب على 300 حقبة مع حجم دفعة قدره 1024 ومعدل تعلم أساسي قدره 0.0005. يتم استخدام مُحسِّن AdamW، مدعومًا باستراتيجية تدهور معدل التعلم الكوسيني. تشمل المعلمات الإضافية تدهور الوزن بنسبة 0.05، ومعدل تسرب قدره 0.1، وفترة تسخين قدرها 5 حقب. يتم تحديد الطبقات الالتفافية على أنها تستخدم كل من الالتفافات القابلة للفصل حسب العمق والالتفافات الموسعة بحجم نواة قدره $3 \times 3$، بينما تستخدم وحدة الالتفاف الخفيفة حجم نواة قدره $7 \times 7$.
نتائج
يقدم قسم “النتائج” من الورقة البحثية النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح النتائج التي توصلت إليها الدراسة، مع تسليط الضوء على نقاط البيانات والاتجاهات المهمة التي تم ملاحظتها. عادةً ما تدعم النتائج تحليلات إحصائية، والتي قد تشمل قيم p، وفترات الثقة، أو مقاييس أخرى ذات صلة للتحقق من النتائج.
بالإضافة إلى ذلك، من المحتمل أن تتضمن أي تمثيلات رسومية، مثل المخططات أو الجداول، لتوضيح البيانات بشكل أكثر فعالية. قد يناقش القسم أيضًا تداعيات النتائج فيما يتعلق بالفرضيات المطروحة في بداية الدراسة، مما يوفر ارتباطًا واضحًا بين الأدلة التجريبية والإطار النظري. بشكل عام، يخدم هذا القسم لنقل المساهمات الأساسية للبحث إلى الجسم المعرفي القائم في هذا المجال.
نقاش
في هذا القسم، يناقش المؤلفون دمج تقنيات الالتفاف في بنية المحولات لتعزيز أدائها في مهام الرؤية. يتضمن النموذج المقترح، CEFormer، الالتفافات القابلة للفصل حسب العمق والالتفافات الموسعة، بهدف تحقيق عدم التغير تحت الترجمة، وعدم التغير تحت المقياس، والمحلية. يقلل الالتفاف القابل للفصل حسب العمق من الحمل الحسابي من خلال استخدام عمليات التفاف منفصلة للأبعاد المكانية والقناة، بينما توسع الالتفافات الموسعة المجال الاستقبالي دون تقليل الدقة، مما يحافظ على دقة الصورة والتفاصيل.
يبرز المؤلفون مزايا دمج عمليات الالتفاف مع آلية الانتباه الذاتي للمحولات. يسمح هذا الدمج بالاحتفاظ بعدم التغير تحت الترجمة من الطبقات الالتفافية مع الاستفادة من الطبيعة الديناميكية المعتمدة على الإدخال لآليات الانتباه. تؤكد الدراسة على أن وحدة الالتفاف الخفيفة المقترحة تعزز استقرار النموذج وسرعة التقارب، مما يعالج التحديات التي تواجه عادةً أثناء تدريب المحولات. تظهر النتائج التجريبية أن CEFormer يتفوق على النماذج التقليدية القائمة على الالتفاف والمحول من حيث الدقة وكفاءة الحساب، مما يشير إلى إمكانيته في تطبيقات الرؤية المختلفة.
DOI: https://doi.org/10.32604/cmes.2024.051083
Publication Date: 2024-01-01
Author(s): Lirong Yin et al.
Primary Topic: Neural Networks and Applications
Overview
This study presents the Convolution-based Efficient Transformer Image Feature Extraction Network (CEFormer), which addresses the limitations of traditional Transformer models in image feature extraction, particularly their lack of inductive bias for visual structures. The CEFormer integrates E-Attention, depthwise separable convolution, and dilated convolution to introduce essential inductive biases such as translation invariance, locality, and scale invariance. This innovative architecture enhances stability and accelerates convergence, resulting in improved performance on image classification tasks. Experimental results on the ImageNet1k Top-1 dataset demonstrate that CEFormer achieves up to 85.0% accuracy, surpassing various baseline models and showing superior performance when integrated into the Mask Region-Convolutional Neural Network (R-CNN) framework.
The findings indicate that CEFormer consistently outperforms both convolution-based and pure Transformer models, achieving up to 4.5% higher accuracy than models like ResNet and RegNetY while requiring fewer floating-point operations (FLOPs) than state-of-the-art vision Transformers. An ablation study reveals that while all four characteristics—translation invariance, locality, scale invariance, and stability—contribute to the model’s performance, stability is the most critical factor affecting accuracy. The study also notes that the incorporation of multiple convolutional modules increases the model’s parameter count and time complexity, suggesting that future research should explore techniques such as pruning and knowledge distillation to enhance model efficiency without sacrificing performance.
Introduction
The introduction of this research paper highlights the significant role of computer vision within the realm of deep learning, emphasizing its applications in tasks such as face recognition, object tracking, and pedestrian detection. The rapid growth of image data, fueled by advancements in big data and computer performance, underscores the importance of effective feature extraction methods to enhance the accuracy and efficiency of visual tasks. While convolutional neural networks (CNNs) have been the primary tool for feature extraction, they face limitations in capturing global context and modeling inter-feature dependencies, which has prompted the exploration of Transformer architectures originally developed for natural language processing.
Transformers offer advantages in modeling long-range dependencies and parallel processing through their self-attention mechanism, allowing for a more comprehensive analysis of images compared to CNNs. However, the paper identifies critical challenges associated with Transformer-based models, including quadratic time and space complexity related to input token sequences, which restrict their application to high-resolution images and edge devices. Additionally, Transformers require extensive pre-training datasets and exhibit sensitivity to hyperparameter selection, often resulting in slower convergence and marginally inferior performance compared to CNNs. To address these issues, the authors propose the Convolution-based Efficient Transformer Image Feature Extraction Network (CEFormer), which integrates convolutional strategies to enhance translation invariance, spatial locality, and scale invariance, thereby improving both convergence speed and model stability in image processing tasks.
Methods
In the Methods section, the authors detail the experimental environment and hyperparameter settings utilized in their study. The software and hardware configurations are summarized in Table 1, while Table 2 presents the model settings for various sizes. The hyperparameters are meticulously defined: the training process spans 300 epochs with a batch size of 1024 and a base learning rate of 0.0005. The AdamW optimizer is employed, complemented by a cosine learning rate decay strategy. Additional parameters include a weight decay of 0.05, a dropout rate of 0.1, and a warmup period of 5 epochs. The convolutional layers are specified as using both depthwise and dilated convolutions with a kernel size of $3 \times 3$, while the lightweight convolution module employs a kernel size of $7 \times 7$.
Results
The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It details the outcomes of the study, highlighting significant data points and trends observed. The results are typically supported by statistical analyses, which may include p-values, confidence intervals, or other relevant metrics to validate the findings.
Additionally, any graphical representations, such as charts or tables, are likely included to illustrate the data more effectively. The section may also discuss the implications of the results in relation to the hypotheses posed at the outset of the study, providing a clear connection between the empirical evidence and the theoretical framework. Overall, this section serves to convey the core contributions of the research to the existing body of knowledge in the field.
Discussion
In this section, the authors discuss the integration of convolutional techniques into the Transformer architecture to enhance its performance in vision tasks. The proposed model, CEFormer, incorporates depthwise separable convolution and dilated convolution, aiming to achieve translation invariance, scale invariance, and locality. Depthwise separable convolution reduces computational overhead by using separate convolution operations for spatial and channel dimensions, while dilated convolution expands the receptive field without downsampling, thus preserving image resolution and detail.
The authors highlight the advantages of combining convolutional operations with the self-attention mechanism of Transformers. This integration allows for the retention of translation invariance from convolutional layers while leveraging the dynamic, input-dependent nature of attention mechanisms. The study emphasizes that the proposed lightweight convolution module enhances model stability and convergence speed, addressing challenges typically faced during Transformer training. The experimental results demonstrate that CEFormer outperforms traditional convolutional and Transformer-based models in terms of accuracy and computational efficiency, indicating its potential for various vision applications.
