تصميم نموذج متكامل مع انتباه الرسم البياني الزمني وRNNs المعززة بالمحول لتحسين الكشف عن الشذوذ
Design of an integrated model with temporal graph attention and transformer-augmented RNNs for enhanced anomaly detection

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-85822-5
PMID: https://pubmed.ncbi.nlm.nih.gov/39837915
تاريخ النشر: 2025-01-21
المؤلف: Sai Babu Veesam وآخرون
الموضوع الرئيسي: تقنيات الكشف عن الشذوذ وتطبيقاتها

نظرة عامة

تقدم ورقة البحث إطارًا جديدًا لاكتشاف الشذوذ في أنظمة المراقبة متعددة الكاميرات، مع معالجة القيود التي تواجهها الطرق الحالية التي تكافح مع الاعتماديات الزمنية طويلة الأمد والارتباطات المكانية. يدمج الإطار المقترح نماذج التعلم العميق المتقدمة، بما في ذلك شبكات الانتباه الزمنية الرسومية (TGAT)، والشبكات العصبية المتكررة المعززة بالمحول (TARNN)، ومشفّر التباين متعدد الأنماط (MVAE)، لتعزيز نمذجة السياقات الزمنية والمكانية. بالإضافة إلى ذلك، يستخدم الشبكات النموذجية للتعلم من عدد قليل من الأمثلة للتعميم من الشذوذ المعلّم المحدود ويستخدم مشفّر زمني مكاني لاكتشاف الشذوذ غير المراقب من خلال تعلم أنماط السلوك الطبيعي. يظهر الإطار تحسينات كبيرة في الدقة والاسترجاع ودرجات F1، محققًا تحسينًا يصل إلى 20% في التعميم على الشذوذ غير المرئي مقارنة بالنماذج التقليدية.

تشير النتائج إلى أن مجموعة TGAT + TARNN تحقق درجة F1 تبلغ 90.6% وAUC تبلغ 95.2%، متجاوزة الطرق التقليدية بنسبة 12%. يدمج MVAE بشكل فعال بيانات الفيديو والصوت وبيانات مستشعر الحركة، مما يظهر قوة ضد الضوضاء مع تقليل بنسبة 25% في خطأ إعادة البناء للأحداث الشاذة. تحقق الشبكات النموذجية دقة تبلغ 87.9% في سيناريوهات التعلم من عدد قليل من الأمثلة، بينما يقلل المشفّر الزمني المكاني من الإيجابيات الكاذبة إلى 6.5%. على الرغم من هذه التقدمات، يواجه الإطار تحديات في بيئات البيانات النادرة ومدخلات منخفضة الدقة، مما قد يؤدي إلى زيادة الإيجابيات الكاذبة وتقليل الدقة. يقترح المؤلفون أن تركز الأعمال المستقبلية على تقنيات التكيف المتقدمة مع المجال وتحسينات الحوسبة الطرفية لتعزيز الأداء في النشر على نطاق واسع وسيناريوهات البيانات منخفضة الجودة.

النتائج

يقدم قسم النتائج تقييمًا شاملاً لنموذج TGAT + TARNN المقترح مقابل طرق مرجعية مختلفة لاكتشاف الشذوذ، باستخدام مجموعة بيانات UCSD للمشاة. حقق النموذج قيم دقة واسترجاع مثيرة للإعجاب تبلغ 92.5% و90.1%، على التوالي، مما أدى إلى درجة F1 تبلغ 91.3%. بالمقابل، حققت أفضل طريقة مرجعية أداءً قياسات أقل تبلغ 83.6% دقة، و81.9% استرجاع، ودرجة F1 تبلغ 82.7%. بالإضافة إلى ذلك، وصلت مجموعة MVAE والمشفّر الزمني المكاني إلى منطقة تحت المنحنى (AUC) تبلغ 96.4% لمجموعة بيانات UCSD Ped2، متفوقة بشكل كبير على القيمة الثانية الأفضل البالغة 88.7%. تؤكد هذه النتائج فعالية النماذج المقترحة، لا سيما في اكتشاف الشذوذات النادرة والدقيقة بدقة.

علاوة على ذلك، تم تقييم قوة النماذج تحت ظروف البيانات المزعجة وغير المكتملة. من خلال إدخال الضوضاء من خلال تدهور 20% من إطارات الفيديو وتغيير جودة الصوت، حافظ المشفّر الزمني المكاني على درجة F1 تبلغ 89.1% ودقة اكتشاف الشذوذ تبلغ 87.8%، مقارنة بـ 76.5% فقط من المشفرات التقليدية. يُعزى هذا الصمود إلى بنية النموذج، التي تدمج طبقات الالتفاف لاستخراج الميزات المكانية وطبقات LSTM لنمذجة المعلومات الزمنية، مما يمكّن من اكتشاف الشذوذ بشكل موثوق حتى مع بيانات الإدخال الم compromised.

المناقشة

يسلط قسم المناقشة في الورقة الضوء على التعقيد المتزايد لأنظمة المراقبة الحديثة وعدم كفاية طرق اكتشاف الشذوذ الحالية في بيئات متعددة الكاميرات. تكافح النماذج التقليدية، التي تعتمد غالبًا على الميزات المصنوعة يدويًا أو تقنيات التعلم العميق المبكرة، لالتقاط الديناميات المكانية والزمنية الدقيقة اللازمة لاكتشاف الشذوذ بشكل فعال، خاصة بالنظر إلى ندرة ودقة الشذوذ في مجموعات البيانات الضخمة. يقترح المؤلفون نموذجًا متكاملًا جديدًا، وهو شبكة الانتباه الزمنية الرسومية (TGAT)، التي تستخدم آليات الانتباه للتركيز ديناميكيًا على تدفقات الكاميرا ذات الصلة، مما يعزز مراقبة البيئات المعقدة. يجمع هذا النموذج بين قوة الشبكات العصبية المتكررة (RNNs) للاعتماديات قصيرة المدى مع المحولات للاعتماديات طويلة المدى، مما يسهل نهجًا شاملاً لاكتشاف الشذوذ.

بالإضافة إلى ذلك، تقدم الورقة مشفّر التباين متعدد الأنماط (MVAE) الذي يدمج البيانات من أوضاع مستشعرات مختلفة، مما يحسن مساحة كامنة مشتركة لتعزيز القوة ضد البيانات المزعجة أو غير المكتملة. تعالج دمج التعلم من عدد قليل من الأمثلة عبر الشبكات النموذجية تحدي مجموعات البيانات المعلّمة المحدودة، مما يسمح للنموذج بالتعميم بشكل فعال على الشذوذ غير المرئي. يوفر المشفّر الزمني المكاني إطارًا غير مراقب لتعلم أنماط السلوك الطبيعي، مما يقلل بشكل كبير من معدلات الإيجابيات الكاذبة. تظهر النتائج التجريبية أن الطرق المقترحة تتفوق على النماذج الحالية الرائدة من حيث الدقة والتعميم، مما يجعلها قابلة للتطبيق بشكل كبير في تطبيقات الأمن والمراقبة المتنوعة. تستعرض المناقشة أيضًا النماذج الحالية لاكتشاف شذوذ الحشود، مع التأكيد على قيود أساليب التعلم المراقب وإمكانية النماذج الهجينة التي تدمج تقنيات التعلم المراقب وغير المراقب والتعلم المعزز لتعزيز القابلية للتوسع والقدرة على التكيف في البيئات الحضرية المعقدة.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-85822-5
PMID: https://pubmed.ncbi.nlm.nih.gov/39837915
Publication Date: 2025-01-21
Author(s): Sai Babu Veesam et al.
Primary Topic: Anomaly Detection Techniques and Applications

Overview

The research paper presents a novel framework for anomaly detection in multi-camera surveillance systems, addressing the limitations of existing methods that struggle with long-term temporal dependencies and spatial correlations. The proposed framework integrates advanced deep learning models, including Temporal Graph Attention Networks (TGAT), Transformer-Augmented Recurrent Neural Networks (TARNN), and a Multimodal Variational Autoencoder (MVAE), to enhance the modeling of temporal and spatial contexts. Additionally, it employs Prototypical Networks for few-shot learning to generalize from limited labeled anomalies and utilizes a Spatiotemporal Autoencoder for unsupervised anomaly detection by learning normal behavior patterns. The framework demonstrates significant improvements in precision, recall, and F1-scores, achieving up to 20% better generalization to unseen anomalies compared to traditional models.

The results indicate that the TGAT + TARNN combination yields an F1-score of 90.6% and an AUC of 95.2%, surpassing traditional methods by 12%. The MVAE effectively fuses video, audio, and motion sensor data, showing robustness against noise with a 25% reduction in reconstruction error for anomalous events. The Prototypical Networks achieve an accuracy of 87.9% in few-shot learning scenarios, while the Spatiotemporal Autoencoder reduces false positives to 6.5%. Despite these advancements, the framework faces challenges in sparse data environments and low-resolution inputs, which can lead to increased false negatives and reduced accuracy. The authors suggest that future work should focus on advanced domain adaptation techniques and edge-computing optimizations to enhance performance in large-scale deployments and low-quality data scenarios.

Results

The results section presents a comprehensive evaluation of the proposed TGAT + TARNN model against various benchmark methods for anomaly detection, utilizing the UCSD Pedestrian Dataset. The model achieved impressive precision and recall values of 92.5% and 90.1%, respectively, culminating in an F1-score of 91.3%. In contrast, the best-performing benchmark method yielded lower metrics of 83.6% precision, 81.9% recall, and an F1-score of 82.7%. Additionally, the combination of MVAE and Spatiotemporal Autoencoder reached an Area Under the Curve (AUC) of 96.4% for the UCSD Ped2 dataset, significantly outperforming the second-best value of 88.7%. These findings underscore the effectiveness of the proposed models, particularly in accurately detecting rare and subtle anomalies.

Furthermore, the robustness of the models was assessed under conditions of noisy and incomplete data. By introducing noise through the degradation of 20% of video frames and varying audio quality, the Spatiotemporal Autoencoder maintained an F1-score of 89.1% and an anomaly detection accuracy of 87.8%, compared to just 76.5% from conventional autoencoders. This resilience is attributed to the model’s architecture, which integrates convolutional layers for spatial feature extraction and LSTM layers for temporal information modeling, enabling reliable anomaly detection even with compromised input data.

Discussion

The discussion section of the paper highlights the increasing complexity of modern surveillance systems and the inadequacies of existing anomaly detection methods in multi-camera environments. Traditional models, often reliant on hand-crafted features or early deep learning techniques, struggle to capture the nuanced spatial and temporal dynamics necessary for effective anomaly detection, particularly given the rarity and subtlety of anomalies in vast datasets. The authors propose a novel integrated model, the Temporal Graph Attention Network (TGAT), which utilizes attention mechanisms to dynamically focus on relevant camera feeds, thereby enhancing the monitoring of complex environments. This model combines the strengths of Recurrent Neural Networks (RNNs) for short-term correlations with transformers for long-range dependencies, facilitating a comprehensive approach to anomaly detection.

Additionally, the paper introduces a Multimodal Variational Autoencoder (MVAE) that fuses data from various sensor modalities, optimizing a joint latent space to improve robustness against noisy or incomplete data. The incorporation of few-shot learning via Prototypical Networks addresses the challenge of limited labeled datasets, allowing the model to generalize to unseen anomalies effectively. The Spatiotemporal Autoencoder offers an unsupervised framework for learning normal behavior patterns, significantly reducing false positive rates. Experimental results demonstrate that the proposed methods outperform existing state-of-the-art models in terms of accuracy and generalization, making them highly applicable for diverse security and monitoring applications. The discussion also reviews existing models for crowd anomaly detection, emphasizing the limitations of supervised learning approaches and the potential of hybrid models that integrate supervised, unsupervised, and reinforcement learning techniques to enhance scalability and adaptability in complex urban environments.

شارك: