DOI: https://doi.org/10.1038/s41598-026-42748-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41974758
تاريخ النشر: 2026-04-13
المؤلف: Vineet Kumar Rakesh وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة
نظرة عامة
تبحث هذه الدراسة بشكل منهجي في تأثير خيارات الهايبر بارامتر على أداء الشبكات الخفيفة القائمة على الالتفاف والتحويل لتصنيف الصور في الوقت الحقيقي على الأجهزة ذات الموارد المحدودة. تركز الدراسة على سبع معماريات خفيفة حديثة—ConvNeXt-Tiny وEfficientNetV2-S وMobileNetV3-L وMobileViT v2 (S/XS) وRepVGG-A2 وTinyViT-21M—تم تدريبها من الصفر على مجموعة متوازنة من ImageNet-1K على بروتوكول موحد من 300 دورة. تشمل الهايبر بارامترات الرئيسية التي تم تحليلها مقدار معدل التعلم، واختيار المحسن (SGD مقابل AdamW)، وتقنيات تنظيم مختلفة مثل RandAugment وMixup وCutMix وlabel smoothing، مع استخدام عمليات البحث الآلي للت优化.
تشير النتائج إلى أن ضبط الهايبر بارامترات بعناية يمكن أن يؤدي إلى تحسينات كبيرة في الدقة (حوالي 1.5-3.5% Top-1) دون تغيير بنية الشبكة، بينما تحدد أيضًا مناطق الاستقرار المعتمدة على البنية. تكشف تقييمات الأداء أن MobileNetV3-L وRepVGG-A2 تحققان نقاط تشغيل مثالية في الوقت الحقيقي مع زمن استجابة منخفض وإنتاجية عالية على GPU. بالإضافة إلى ذلك، تؤكد اختبارات نشر الحافة على الفوائد المحدودة للتجميع على وحدات المعالجة المركزية ذات الطاقة المنخفضة وتبرز ضرورة اختيار النماذج بناءً على اعتبارات زمن الاستجابة. الكود الكامل وسجلات التجارب متاحة في مستودع GitHub المقدم.
مقدمة
تناقش مقدمة ورقة البحث التحديات المتعلقة بتصنيف الصور في الوقت الحقيقي على الأجهزة الطرفية والأطراف، مع التأكيد على الحاجة إلى نماذج التعلم العميق التي تحافظ على أداء تنبؤي قوي مع الالتزام بقيود صارمة على الحوسبة والذاكرة وزمن الاستجابة. أدت هذه الحاجة إلى تطوير معماريات خفيفة، تتكون عادةً من حوالي 30 مليون معلمة، والتي تهدف إلى تحقيق توازن بين الدقة وقابلية النشر. تسلط الورقة الضوء على تنوع عائلات النماذج الخفيفة المعاصرة، بما في ذلك الشبكات العصبية الالتفافية الفعالة وتصميمات CNN-Transformer الهجينة، حيث يقدم كل منها تنازلات فريدة من حيث القدرة التمثيلية وسلوك التحسين وكفاءة الاستدلال.
على الرغم من التقدم في تصميم المعمارية، فإن الأداء العملي لهذه النماذج في التطبيقات الواقعية يتأثر بشكل كبير بالهايبر بارامترات التدريب، مما يعقد المقارنات المباشرة عبر المعماريات المختلفة. لمعالجة ذلك، تستكشف الدراسة تقنيات التعلم الآلي الآلي (AutoML) لتقليل الضبط اليدوي وتوحيد عمليات التدريب. يقوم المؤلفون بالتحقيق بشكل منهجي في سبع هياكل خفيفة معتمدة على نطاق واسع—EfficientNetV2-S وConvNeXt-Tiny وMobileNetV3-L وMobileViT v2 وTinyViT-21M وRepVGG-A2—باستخدام مجموعة متوازنة من ImageNet-1K تحتوي على 90,000 صورة تدريب. يسهل هذا التصميم التجريبي المنضبط المقارنات السريعة والقابلة للتكرار تحت قيود موحدة، على الرغم من أن المؤلفين يحذرون من أن قيم الدقة المطلقة لا ينبغي مقارنتها مباشرة بمعايير ImageNet الكاملة. الفرضية المركزية المطروحة هي أن تحسين الهايبر بارامترات بشكل منهجي يمكن أن يؤدي إلى تحسينات متسقة في الأداء عبر هذه النماذج.
الطرق
في هذا القسم، يناقش المؤلفون المنهجية المستخدمة لتحقيق دقة عالية مع الحد الأدنى من التأخير الحاسوبي في الشبكات العصبية الالتفافية (CNNs). يسلطون الضوء على تطوير عدة معماريات CNN فعالة مصممة خصيصًا لتطبيقات الرؤية المحمولة والمضمنة. تشمل النماذج البارزة المذكورة SqueezeNet وMobileNets وShuffleNet وEfficientNetV2، والتي تم تصميمها لتحسين الأداء مع الحفاظ على استهلاك منخفض للموارد. تعكس هذه التقدمات اتجاهًا كبيرًا في هذا المجال نحو إنشاء نماذج خفيفة مناسبة للمعالجة في الوقت الحقيقي في البيئات المقيدة.
النتائج
في هذا القسم، يقدم المؤلفون تحليلًا شاملاً لتحسينات الأداء التي تم الحصول عليها من خلال تحسين الهايبر بارامترات بشكل منهجي المطبق على سبع نماذج تعلم عميق خفيفة. تم تدريب كل نموذج لمدة 300 دورة متسقة، باستخدام مجموعة موحدة من مجموعة بيانات ImageNet-1K، التي تتكون من 90,000 صورة تدريب و10,000 صورة تحقق. يضمن هذا النهج مقارنة عادلة ومراقبة بين النماذج.
كان القرار بالتدريب لمدة 300 دورة استراتيجيًا، حيث كانت خمس من المعماريات السبع قد أبلغت سابقًا عن جداول تدريب مماثلة في دراساتهم الأصلية، بينما تم تدريب النموذجين الآخرين لفترات طويلة. يسمح هذا التوافق في مدة التدريب بتقييم أكثر دقة لتأثير تحسين الهايبر بارامترات على أداء النموذج.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على التقدم في نماذج تصنيف الصور الخفيفة، مع التأكيد على التوازن بين الدقة وزمن الاستجابة. تظهر المعماريات البارزة مثل MobileNetV3 وEfficientNetV2 وRepVGG-A2 تحسينات كبيرة في الدقة مع الحفاظ على متطلبات موارد منخفضة. على سبيل المثال، تحقق MobileNetV3-L دقة 75.2% في Top-1 مع 5.4 مليون معلمة، متفوقة على سابقتها بنسبة 3.2% مع تقليل زمن الاستجابة. يصل EfficientNetV2-S، مع 22 مليون معلمة، إلى دقة 83.9% من خلال التوسع المشترك في العمق والعرض والدقة. يناقش القسم أيضًا دمج تحسين الهايبر بارامترات واستراتيجيات التدريب، كاشفًا أن تقنيات مثل التعزيز (مثل RandAugment وMixup وCutMix) وجدولة معدل التعلم ضرورية لتحسين أداء النموذج. على سبيل المثال، حسنت CutMix دقة ResNet-50 من 76.3% إلى 78.6%، مما يوضح تأثير منهجيات التدريب الفعالة.
تقوم الورقة بتقييم منهجي لسبع معماريات متطورة، بما في ذلك ConvNeXt-Tiny وTinyViT-21M، تحت أنظمة تدريب محسّنة لتقييم تأثير الهايبر بارامترات. تشير النتائج إلى أن ضبط الهايبر بارامترات، لا سيما فيما يتعلق بمعدلات التعلم وأحجام الدفعات، يؤثر بشكل كبير على دقة النموذج واستقرار التقارب. تؤكد الدراسة على أهمية إعداد تجريبي منضبط، باستخدام مجموعة تمثيلية من مجموعة بيانات ImageNet-1K لتسهيل المقارنات العادلة عبر النماذج. بشكل عام، تؤكد النتائج أنه بينما يمكن للنماذج الخفيفة تحقيق دقة تنافسية، فإن أدائها حساس للغاية لتغيرات الهايبر بارامترات، مما يتطلب مزيدًا من الاستكشاف في التطبيقات الواقعية.
القيود
تقدم الدراسة قيدًا كبيرًا في اعتمادها على مجموعة متوازنة من 90,000 صورة من ImageNet-1K، وذلك أساسًا بسبب قيود GPU. يؤثر هذا القيد على قابلية المقارنة المباشرة لقيم الدقة التي تم الحصول عليها مع تلك من المعايير الكاملة. تؤكد البحث على تحليل حساسية الهايبر بارامترات النسبية وأنماط التفاعل—مثل استقرار التقارب وتنازلات حجم الدفعة—بدلاً من التركيز على تحقيق مقاييس أداء مثالية محددة لمجموعة البيانات.
على الرغم من إمكانية تغير قيم الهايبر بارامترات المثلى مع أحجام مجموعات البيانات المختلفة، تتماشى الدراسة مع الأبحاث السابقة التي تشير إلى أن الاتجاهات النوعية في حساسية الهايبر بارامترات تظل متسقة عبر مجموعات فرعية مختلفة. وبالتالي، فإن النتائج تهدف إلى توجيه اختيار الهايبر بارامترات بشكل فعال وتأسيس أنظمة تدريب قوية، على الرغم من أن التعديلات النهائية ستكون ضرورية عند الانتقال إلى تدريب مجموعة البيانات الكاملة.
DOI: https://doi.org/10.1038/s41598-026-42748-w
PMID: https://pubmed.ncbi.nlm.nih.gov/41974758
Publication Date: 2026-04-13
Author(s): Vineet Kumar Rakesh et al.
Primary Topic: Advanced Neural Network Applications
Overview
This research systematically investigates the impact of hyperparameter choices on the performance of lightweight convolutional and transformer-based networks for real-time image classification on resource-constrained hardware. The study focuses on seven modern lightweight architectures—ConvNeXt-Tiny, EfficientNetV2-S, MobileNetV3-L, MobileViT v2 (S/XS), RepVGG-A2, and TinyViT-21M—trained from scratch on a balanced subset of ImageNet-1K over a standardized 300-epoch protocol. Key hyperparameters analyzed include learning-rate magnitude, optimizer selection (SGD vs. AdamW), and various regularization techniques such as RandAugment, Mixup, CutMix, and label smoothing, with automated searches employed for optimization.
The findings indicate that careful hyperparameter tuning can yield significant accuracy improvements (approximately 1.5-3.5% Top-1) without altering the network architecture, while also identifying architecture-dependent stability regions. Performance evaluations reveal that MobileNetV3-L and RepVGG-A2 achieve optimal real-time operating points with low latency and high throughput on GPU. Additionally, edge deployment tests underscore the limited advantages of batching on low-power CPUs and emphasize the necessity of selecting models based on latency considerations. The complete code and experimental logs are available at the provided GitHub repository.
Introduction
The introduction of the research paper addresses the challenges of real-time image categorization on edge and peripheral devices, emphasizing the need for deep learning models that maintain strong predictive performance while adhering to strict constraints on compute, memory, and latency. This necessity has led to the development of lightweight architectures, typically comprising around 30 million parameters, which aim to strike a balance between accuracy and deployability. The paper highlights the diversity of contemporary lightweight model families, including efficient convolutional neural networks and hybrid CNN-transformer designs, each presenting unique trade-offs in terms of representational capacity, optimization behavior, and inference efficiency.
Despite advancements in architectural design, the practical performance of these models in real-world applications is significantly influenced by training hyperparameters, complicating direct comparisons across different architectures. To address this, the study explores automated machine learning (AutoML) techniques to minimize manual tuning and standardize training processes. The authors systematically investigate seven widely adopted lightweight backbones—EfficientNetV2-S, ConvNeXt-Tiny, MobileNetV3-L, MobileViT v2, TinyViT-21M, and RepVGG-A2—using a class-balanced subset of ImageNet-1K with 90,000 training images. This controlled experimental design facilitates rapid and reproducible comparisons under uniform constraints, although the authors caution that absolute accuracy values should not be directly compared to full ImageNet benchmarks. The central hypothesis posited is that systematic hyperparameter optimization can lead to consistent performance improvements across these models.
Methods
In this section, the authors discuss the methodology employed to achieve high accuracy with minimal computational delay in convolutional neural networks (CNNs). They highlight the development of several efficient CNN architectures specifically tailored for mobile and embedded vision applications. Notable models mentioned include SqueezeNet, MobileNets, ShuffleNet, and EfficientNetV2, which have been designed to optimize performance while maintaining low resource consumption. These advancements reflect a significant trend in the field towards creating lightweight models suitable for real-time processing in constrained environments.
Results
In this section, the authors provide a comprehensive analysis of the performance enhancements obtained through systematic hyperparameter optimization applied to seven lightweight deep learning models. Each model was trained for a consistent duration of 300 epochs, utilizing a uniform subset of the ImageNet-1K dataset, comprising 90,000 training images and 10,000 validation images. This approach ensured a fair and controlled comparison among the models.
The decision to train for 300 epochs was strategic, as five of the seven architectures had previously reported similar training schedules in their original studies, while the other two models were trained for extended periods. This alignment in training duration allows for a more accurate assessment of the impact of hyperparameter optimization on model performance.
Discussion
The discussion section of the research paper highlights advancements in lightweight image classification models, emphasizing the balance between accuracy and latency. Notable architectures such as MobileNetV3, EfficientNetV2, and RepVGG-A2 demonstrate significant improvements in accuracy while maintaining low resource requirements. For instance, MobileNetV3-L achieves 75.2% top-1 accuracy with 5.4 million parameters, outperforming its predecessor by 3.2% with reduced latency. EfficientNetV2-S, with 22 million parameters, reaches 83.9% accuracy through joint scaling of depth, width, and resolution. The section also discusses the integration of hyperparameter optimization and training strategies, revealing that techniques like augmentation (e.g., RandAugment, Mixup, CutMix) and learning rate scheduling are crucial for enhancing model performance. For example, CutMix improved ResNet-50 accuracy from 76.3% to 78.6%, illustrating the impact of effective training methodologies.
The paper systematically evaluates seven cutting-edge architectures, including ConvNeXt-Tiny and TinyViT-21M, under optimized training regimes to assess the effects of hyperparameters. The findings indicate that hyperparameter tuning, particularly concerning learning rates and batch sizes, significantly influences model accuracy and convergence stability. The study emphasizes the importance of a controlled experimental setup, utilizing a representative subset of the ImageNet-1K dataset to facilitate fair comparisons across models. Overall, the results underscore that while lightweight models can achieve competitive accuracy, their performance is highly sensitive to hyperparameter variations, necessitating further exploration in real-world applications.
Limitations
The study presents a significant limitation in its reliance on a class-balanced subset of 90,000 images from ImageNet-1K, primarily due to GPU constraints. This restriction impacts the direct comparability of the accuracy values obtained with those from full-scale benchmarks. The research emphasizes the analysis of relative hyperparameter sensitivity and interaction patterns—such as convergence stability and batch size trade-offs—rather than focusing on achieving optimal performance metrics specific to the dataset.
Despite the potential variability of optimal hyperparameter values with different dataset sizes, the study aligns with previous research indicating that qualitative trends in hyperparameter sensitivity remain consistent across various subsets. Consequently, the findings are intended to guide efficient hyperparameter selection and establish robust training regimes, although final adjustments will be necessary when transitioning to full dataset training.
