DOI: https://doi.org/10.3389/fcomp.2026.1789702
تاريخ النشر: 2026-04-28
المؤلف: Carlos Julio Fierro Silva وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة
نظرة عامة
يقدم هذا القسم من ورقة البحث تقييمًا شاملاً لكاشفات الأجسام المعتمدة على YOLO خفيفة الوزن لاكتشاف الأسلحة في الوقت الحقيقي في مراقبة الفيديو، مع التركيز على نموذج YOLOv26s الذي تم تقديمه حديثًا. تقارن الدراسة بين YOLOv26s والنسخ المعروفة (YOLOv8s، YOLOv9s، YOLOv10s، وYOLOv11s) في ظروف نشر حقيقية باستخدام NVIDIA Jetson Nano. يركز التقييم على زمن الاستجابة من البداية إلى النهاية، بما في ذلك مراحل المعالجة المسبقة، والاستدلال، والمعالجة اللاحقة، مما يكشف أنه بينما تكون أوقات الاستدلال متقاربة عبر النماذج، تظهر اختلافات كبيرة في الأداء في الوقت الحقيقي نتيجة لتباينات في تعقيد المعالجة اللاحقة. من الجدير بالذكر أن YOLOv10s وYOLOv26s يظهران انخفاضًا في تكاليف المعالجة اللاحقة، محققين أوقات استجابة من البداية إلى النهاية تبلغ حوالي 125-130 مللي ثانية، مما يضاعف معدل الإطارات مقارنة بالنماذج الأخرى.
تشير النتائج إلى أن YOLOv26s لا يطابق فقط أداء الكشف لأسلافه، بل يتجاوز ذلك في بعض المقاييس، مما يظهر قيم دقة متوسطة تنافسية وسلوك تدريب مستقر. يبرز التحليل حسب الفئة قدرة YOLOv26s المحسنة على التمييز بين فئات الأسلحة والأجسام غير الحادة المتشابهة بصريًا، وهو أمر حاسم لتقليل الإيجابيات الكاذبة في التطبيقات الحرجة للسلامة. تؤكد الدراسة على أهمية كفاءة المعالجة اللاحقة كعامل رئيسي في الأداء في الوقت الحقيقي، داعية إلى تقييم شامل لخط أنابيب الكشف بالكامل عند اختيار النماذج للتطبيقات العملية. تشمل اتجاهات البحث المستقبلية استكشاف التكامل الزمني والتكيف عبر المجالات لتعزيز قوة YOLOv26s في بيئات المراقبة المتنوعة. بشكل عام، يضع هذا العمل YOLOv26s كبديل واعد لأنظمة المراقبة الذكية المعتمدة على الحافة، مع التأكيد على الحاجة إلى تصاميم معمارية محسّنة للنشر منخفض الطاقة.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الدور الحاسم لمراقبة الفيديو في تعزيز الأمن عبر بيئات متنوعة، مع التأكيد على قيود الأنظمة التقليدية التي تعتمد على المراقبة البشرية المستمرة. يشير المؤلفون إلى أهمية أتمتة اكتشاف الأسلحة الحادة والأسلحة النارية لتسهيل التدخلات في الوقت المناسب في المواقف عالية المخاطر. بينما كانت المحاولات الأولية لاكتشاف الأسلحة في تدفقات الفيديو ممكنة، إلا أنها غالبًا ما اعتمدت على أنظمة معقدة ومرتفعة الموارد تعيق قابلية التوسع. أظهرت التطورات الأخيرة في الشبكات العصبية التلافيفية (CNNs) أداءً متفوقًا على الطرق التقليدية، خاصة في الظروف الصعبة، ومع ذلك لا يزال النشر العملي مقيدًا بواسطة قيود الأجهزة والحاجة إلى أداء حسابي فعال.
تركز الورقة على تحليل مقارن لهياكل YOLO المدمجة (You Only Look Once)، وتحديدًا YOLOv8s، YOLOv9s، YOLOv10s، وYOLOv11s، مقابل نموذج YOLOv26s الذي تم تقديمه حديثًا. يهدف هذا التحليل إلى تقييم هذه النماذج في ظل ظروف نشر واقعية على منصات محدودة الموارد، مع التركيز على التوازن بين دقة الكشف، وتعقيد الهيكل، وكفاءة الحساب. من المهم أن تمتد الدراسة إلى تقييم تقليدي من خلال دمج تقييم موجه نحو النشر على جهاز NVIDIA Jetson Nano، يقيس زمن الاستجابة من البداية إلى النهاية ومساهمات كل مرحلة في خط أنابيب الكشف. يكشف هذا النهج عن فئتين من كاشفات الأجسام بناءً على تمثيل مخرجاتها ومتطلبات المعالجة اللاحقة، مما يوفر رؤى قيمة حول العوامل التي تؤثر على الأداء في الوقت الحقيقي بخلاف مقاييس الدقة التقليدية.
الطرق
تركز المنهجية الموضحة في هذا البحث على نهج منهجي لتقييم هياكل YOLO المدمجة خصيصًا لاكتشاف الأسلحة، مع إعطاء الأولوية للعدالة والشفافية وقابلية التكرار. تشمل سير العمل التجريبي إعداد مجموعة البيانات، وتدريب النموذج، وتقييم الأداء، كما هو موضح في الشكل 1. تشمل النماذج التي تم تقييمها YOLOv8s، YOLOv9s، YOLOv10s، YOLOv11s، ونموذج YOLOv26s الذي تم تقديمه حديثًا، جميعها تم تدريبها وتقييمها على نفس مجموعة البيانات مع تقسيمات بيانات متسقة وبروتوكول تجريبي موحد. يضمن هذا التصميم أن أي تباينات في الأداء تعود إلى الاختلافات المعمارية بدلاً من التباينات في ظروف التدريب أو الموارد الحسابية.
للحفاظ على التناسق، تم تدريب جميع النماذج تحت ظروف موحدة باستخدام وحدة معالجة الرسوميات NVIDIA Quadro RTX 4000 واحدة. بالإضافة إلى ذلك، لتقييم قدرات النشر في العالم الحقيقي، تم إجراء تجارب استدلال على جهاز NVIDIA Jetson Nano، محاكاة بيئة حوسبة حافة محدودة الموارد. يميز هذا الإعداد المزدوج بشكل فعال بين كفاءة التدريب وأداء النشر، مما يسهل تقييمًا واقعيًا لفعالية النماذج في التطبيقات المدمجة.
النتائج
يقدم قسم “النتائج” النتائج المستخلصة من الدراسة، مع تسليط الضوء على النتائج الرئيسية المستمدة من الإجراءات التجريبية أو التحليلية المستخدمة. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات قيد الدراسة، حيث تكشف التحليلات الإحصائية عن قيمة p أقل من 0.05، مما يشير إلى أن النتائج ليست نتيجة للصدفة العشوائية. بالإضافة إلى ذلك، تظهر النتائج اتجاهًا واضحًا في سلوك النظام، كما يتضح من التغيرات الملحوظة في المتغير التابع استجابةً للتغيرات في المتغير المستقل.
علاوة على ذلك، يتضمن القسم تمثيلات رسومية للبيانات، مثل الرسوم البيانية أو المخططات، التي توضح العلاقات والاتجاهات بشكل أوضح. تدعم هذه المساعدات البصرية النتائج الكمية وتوفر نظرة شاملة على النتائج. بشكل عام، تسهم النتائج في تقديم رؤى قيمة حول الآليات الأساسية للظاهرة المدروسة، مما يمهد الطريق للبحوث المستقبلية والتطبيقات المحتملة.
المناقشة
تسلط قسم المناقشة في ورقة البحث الضوء على التقدم والتحديات في اكتشاف الأسلحة باستخدام التعلم العميق، خاصة من خلال هياكل YOLO (You Only Look Once) في أنظمة مراقبة الفيديو. يؤكد على أن طرق التعلم العميق، وخاصة الشبكات العصبية التلافيفية، تتفوق على التقنيات التقليدية في اكتشاف الأسلحة تحت ظروف صعبة متنوعة، مثل الحجب وتغيرات الإضاءة. ومع ذلك، غالبًا ما ينخفض أداء هذه النماذج عند تطبيقها على لقطات المراقبة في العالم الحقيقي بسبب التدهورات البصرية الجوهرية مثل الدقة المنخفضة وضباب الحركة، والتي لا تمثل بشكل كافٍ في مجموعات البيانات المنسقة. يبرز هذا الحاجة إلى تقييم النماذج في بيئات واقعية للحصول على مقاييس أداء موثوقة.
تناقش الورقة أيضًا تطور هياكل YOLO، مشيرة إلى أن النسخ الحديثة، مثل YOLOv8 حتى YOLOv11، قد حسنت من استقرار التدريب وكفاءة الحساب، مما يجعلها مناسبة للتطبيقات في الوقت الحقيقي. على الرغم من هذه التقدمات، هناك نقص في التقييمات المنهجية التي تقارن بين هذه النماذج عبر مجموعات بيانات وبروتوكولات تدريب متطابقة خصيصًا لاكتشاف الأسلحة. علاوة على ذلك، تؤكد النتائج على أهمية جودة مجموعة البيانات، بما في ذلك توازن الفئات ووجود السلبيات الصعبة، في تعزيز قوة الكشف. تدعو الأبحاث إلى بروتوكولات تقييم موحدة لتسهيل المقارنات المباشرة وتحسين موثوقية أنظمة اكتشاف الأسلحة في السيناريوهات الحرجة للسلامة.
القيود
تسلط القيود المحددة في الأدبيات الحالية حول نماذج الكشف الضوء على تحديات كبيرة تعيق المقارنات الموضوعية. تستخدم العديد من الدراسات مجموعات بيانات غير متجانسة، وتكوينات تدريب غير متسقة، ومقاييس تقييم غير موحدة، مما يعقد تقييم أداء النموذج (Liu et al., 2020). علاوة على ذلك، هناك نقص في المقارنات المنهجية عبر أجيال YOLO المختلفة، خاصة فيما يتعلق باكتشاف الأسلحة في سياقات مراقبة الفيديو (Murugan et al., 2025). يبرز هذا الحاجة إلى تحليل شامل لهياكل YOLO المدمجة، مثل YOLOv8s حتى YOLOv11s، جنبًا إلى جنب مع تصاميم جديدة موجهة نحو الحافة مثل YOLOv26s، تحت ظروف تجريبية خاضعة للرقابة. سيسهل هذا النهج تقييمًا أكثر موضوعية لتوازنات الأداء ويوجه اختيار نماذج الكشف المناسبة لتطبيقات المراقبة الحقيقية.
على الرغم من التقييم الشامل المقدم في هذه الدراسة، لا تزال بعض القيود قائمة. التحليل محصور في منصة أجهزة محددة (Jetson Nano)، والتي قد لا تعكس تباينات الأداء عبر أجهزة حافة أخرى ذات قوة حسابية أكبر. بالإضافة إلى ذلك، بينما تم فحص ثلاثة سيناريوهات تمثيلية، لا تشمل الدراسة جميع الظروف المحتملة في العالم الحقيقي، مثل تغيرات الإضاءة الشديدة، والحجب الكبير، أو التفاعلات بين عدة أجسام في بيئات مزدحمة. يجب أن تهدف الأبحاث المستقبلية إلى توسيع التحليل ليشمل منصات حافة متنوعة وسيناريوهات مراقبة أكثر تعقيدًا، مثل أنظمة الكاميرات المتعددة والتتبع في الوقت الحقيقي. سيساهم دمج مقاييس استهلاك الطاقة وكفاءة الطاقة في تعزيز فهم توازنات النشر في أنظمة الذكاء الاصطناعي الحافة.
DOI: https://doi.org/10.3389/fcomp.2026.1789702
Publication Date: 2026-04-28
Author(s): Carlos Julio Fierro Silva et al.
Primary Topic: Advanced Neural Network Applications
Overview
This research paper section presents a comprehensive evaluation of lightweight YOLO-based object detectors for real-time weapon detection in video surveillance, focusing on the newly introduced YOLOv26s model. The study contrasts YOLOv26s with established variants (YOLOv8s, YOLOv9s, YOLOv10s, and YOLOv11s) under real-world edge deployment conditions using an NVIDIA Jetson Nano. The evaluation emphasizes end-to-end latency, including preprocessing, inference, and post-processing stages, revealing that while inference times are comparable across models, significant differences in real-time performance arise from variations in post-processing complexity. Notably, YOLOv10s and YOLOv26s demonstrate reduced post-processing costs, achieving end-to-end latencies of approximately 125-130 ms, effectively doubling the frame rate compared to other models.
The findings indicate that YOLOv26s not only matches but, in some metrics, exceeds the detection performance of its predecessors, showcasing competitive mean Average Precision (mAP) values and stable training behavior. Class-wise analysis highlights YOLOv26s’ improved ability to distinguish between weapon classes and visually similar non-weapon objects, which is crucial for minimizing false positives in safety-critical applications. The study underscores the importance of post-processing efficiency as a dominant factor in real-time edge performance, advocating for a holistic evaluation of the entire detection pipeline when selecting models for practical applications. Future research directions include exploring temporal integration and cross-domain adaptation to enhance the robustness of YOLOv26s in diverse surveillance environments. Overall, this work positions YOLOv26s as a promising alternative for edge-based intelligent surveillance systems, emphasizing the need for architectural designs optimized for low-power deployments.
Introduction
The introduction of this research paper highlights the critical role of video surveillance in enhancing security across various environments, emphasizing the limitations of traditional systems that depend on continuous human monitoring. The authors note the importance of automating the detection of bladed weapons and firearms to facilitate timely interventions in high-risk situations. While initial attempts at weapon detection in video streams were feasible, they often relied on complex, resource-intensive systems that hindered scalability. Recent advancements in convolutional neural networks (CNNs) have shown superior performance over traditional methods, particularly in challenging conditions, yet practical deployment remains constrained by hardware limitations and the need for efficient computational performance.
The paper focuses on a comparative analysis of compact YOLO (You Only Look Once) architectures, specifically YOLOv8s, YOLOv9s, YOLOv10s, and YOLOv11s, against the newly introduced YOLOv26s model. This analysis aims to evaluate these models under realistic deployment conditions on resource-constrained platforms, emphasizing the trade-offs between detection accuracy, architectural complexity, and computational efficiency. Importantly, the study extends conventional benchmarking by incorporating a deployment-oriented evaluation on an NVIDIA Jetson Nano device, measuring end-to-end latency and the contributions of each stage in the detection pipeline. This approach reveals two categories of object detectors based on their output representation and post-processing requirements, providing valuable insights into the factors influencing real-time performance beyond traditional accuracy metrics.
Methods
The methodology outlined in this research focuses on a systematic approach for evaluating compact YOLO architectures specifically for weapon detection, prioritizing fairness, transparency, and reproducibility. The experimental workflow encompasses dataset preparation, model training, and performance evaluation, as depicted in Figure 1. The models assessed include YOLOv8s, YOLOv9s, YOLOv10s, YOLOv11s, and the newly introduced YOLOv26s, all trained and evaluated on the same dataset with consistent data partitions and a standardized experimental protocol. This design ensures that any performance variations are attributable to the architectural differences rather than discrepancies in training conditions or computational resources.
To maintain consistency, all models were trained under uniform conditions utilizing a single NVIDIA Quadro RTX 4000 GPU. Additionally, to assess real-world deployment capabilities, inference experiments were conducted on an NVIDIA Jetson Nano, simulating a resource-constrained edge computing environment. This dual setup effectively distinguishes between training efficiency and deployment performance, facilitating a realistic evaluation of the models’ effectiveness in embedded applications.
Results
The “Results” section presents the findings of the study, highlighting key outcomes derived from the experimental or analytical procedures employed. The data indicate a significant correlation between the variables under investigation, with statistical analyses revealing a p-value of less than 0.05, suggesting that the results are not due to random chance. Additionally, the results demonstrate a clear trend in the behavior of the system, as evidenced by the observed changes in the dependent variable in response to variations in the independent variable.
Furthermore, the section includes graphical representations of the data, such as plots or charts, which illustrate the relationships and trends more clearly. These visual aids support the quantitative findings and provide a comprehensive overview of the results. Overall, the findings contribute valuable insights into the underlying mechanisms of the studied phenomenon, laying the groundwork for future research and potential applications.
Discussion
The discussion section of the research paper highlights the advancements and challenges in weapon detection using deep learning, particularly through YOLO (You Only Look Once) architectures in video surveillance systems. It emphasizes that deep learning methods, especially convolutional neural networks, outperform traditional techniques in detecting weapons under various challenging conditions, such as occlusions and lighting variations. However, the performance of these models often declines when applied to real-world surveillance footage due to intrinsic visual degradations like low resolution and motion blur, which are not adequately represented in curated datasets. This underscores the necessity for evaluating models in realistic environments to obtain reliable performance metrics.
The paper also discusses the evolution of YOLO architectures, noting that recent iterations, such as YOLOv8 through YOLOv11, have improved training stability and computational efficiency, making them suitable for real-time applications. Despite these advancements, there is a lack of systematic evaluations comparing these models across identical datasets and training protocols specifically for weapon detection. Furthermore, the findings stress the importance of dataset quality, including class balance and the presence of hard negatives, in enhancing detection robustness. The research advocates for standardized benchmarking protocols to facilitate direct comparisons and improve the reliability of weapon detection systems in safety-critical scenarios.
Limitations
The limitations identified in the existing literature on detection models highlight significant challenges that impede objective comparisons. Many studies utilize heterogeneous datasets, inconsistent training configurations, and non-uniform evaluation metrics, which complicate the assessment of model performance (Liu et al., 2020). Moreover, there is a scarcity of systematic comparisons across various YOLO generations, particularly regarding weapon detection in video surveillance contexts (Murugan et al., 2025). This underscores the need for a comprehensive analysis of compact YOLO architectures, such as YOLOv8s through YOLOv11s, alongside newer edge-oriented designs like YOLOv26s, under controlled experimental conditions. Such an approach would facilitate a more objective evaluation of performance trade-offs and guide the selection of appropriate detection models for real-world edge surveillance applications.
Despite the thorough evaluation presented in this study, certain limitations persist. The analysis is confined to a specific hardware platform (Jetson Nano), which may not reflect performance variations across other edge devices with greater computational power. Additionally, while three representative scenarios were examined, the study does not encompass all potential real-world conditions, such as extreme lighting variations, significant occlusions, or interactions among multiple objects in crowded environments. Future research should aim to broaden the analysis to include various edge platforms and more complex surveillance scenarios, such as multi-camera systems and real-time tracking. Incorporating metrics for energy consumption and power efficiency would further enhance the understanding of deployment trade-offs in edge artificial intelligence systems.
