DOI: https://doi.org/10.1016/j.inffus.2026.104513
تاريخ النشر: 2026-06-01
المؤلف: Jian Song وآخرون
الموضوع الرئيسي: طرق المراقبة بالفيديو والتتبع
نظرة عامة
في هذا القسم، يقدم المؤلفون نهجًا جديدًا لتقدير الحركة في تتبع الأجسام المتعددة (MOT) من خلال تطوير شبكة كالمان المستقلة عن الدلالة (SIKNet). يعالج هذا الأسلوب قيود مرشح كالمان التقليدي (KF)، خاصة في السيناريوهات التي تتضمن حركات الأجسام غير الثابتة والمعلمات غير المتطابقة. تستخدم SIKNet مشفرًا مستقلًا عن الدلالة (SIE) الذي يستفيد من عملية من خطوتين: أولاً، يتم استخراج معلومات دلالية مستقلة من العناصر المتجانسة من خلال تلافيف أحادية البعد بحجم نواة 1، تليها طبقة متصلة بالكامل تلتقط العلاقات غير الخطية بين العناصر غير المتجانسة.
لتقييم فعالية SIKNet، قام المؤلفون بإنشاء مجموعة بيانات شبه محاكاة على نطاق واسع مشتقة من مجموعات بيانات MOT الموجودة. تشير النتائج التجريبية إلى أن SIKNet تتفوق بشكل كبير على KF التقليدي ومرشحات التعلم المدعومة الأخرى، حيث تحقق تحسينًا بنسبة 6% في متوسط الاسترجاع (mAR) مقارنة بالنماذج الموجودة وتحسينًا ملحوظًا بنسبة 40% مقارنة بـ KF. يعترف المؤلفون بأنه بينما تظهر SIKNet أداءً قويًا، لا تزال هناك تحديات في دمجها في أطر MOT الكاملة للتدريب من البداية إلى النهاية مع مكونات الكشف. ستركز الأبحاث المستقبلية على تحقيق هذا الدمج لتعزيز أداء التتبع العام.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على أهمية تتبع الأجسام المتعددة عبر الإنترنت وفي الوقت الحقيقي (MOT) في تطبيقات مثل القيادة الذاتية وتحليل الحركة. تؤكد على الاهتمام المتزايد بأساليب التتبع من خلال الكشف (TBD)، التي تستخدم كاشفات قائمة على التعلم العميق لتتبع صناديق الكشف. تتكون هذه الأساليب عادةً من مكونين رئيسيين: وحدة تقدير الحركة (MEM) ووحدة ارتباط البيانات (DAM). يعتمد أداء هذه المتعقبين بشكل كبير على توقعات MEM، التي غالبًا ما تستخدم مرشح كالمان القائم على نموذج سرعة ثابتة خطية (CV). ومع ذلك، فإن هذا الافتراض يفشل في كثير من الأحيان في السيناريوهات الواقعية بسبب الضوضاء غير القابلة للتنبؤ وأنماط حركة الأجسام المعقدة، مما يؤدي إلى فشل التتبع.
لمعالجة قيود أساليب KF التقليدية، يقترح المؤلفون KalmanNet، وهو نموذج تصفية كالمان المدعوم بالتعلم (LAKF)، ويقدمون نسخة محسنة تسمى شبكة كالمان المستقلة عن الدلالة (SIKNet). تستفيد هذه النماذج من الشبكات العصبية المتكررة (RNNs) لتتعلم بشكل تكيفي مكاسب كالمان من ميزات الإدخال، مما يحسن دقة تقدير الحركة في الظروف غير الخطية وغير الثابتة. يعزز دمج مشفر مستقل عن الدلالة (SIE) من القوة من خلال إدارة المعلومات الدلالية غير المتجانسة بشكل فعال. كما يقوم المؤلفون بإنشاء مجموعة بيانات شبه محاكاة للتقييم المستقل لـ MEM، مما يوضح أن SIKNet تتفوق على KF التقليدي ومرشحات التعلم المدعومة الموجودة من حيث الاستقرار والدقة. تشمل المساهمات الرئيسية تطوير SIKNet، مجموعة البيانات شبه المحاكاة، وإطار عمل مفتوح المصدر، FilterNet، للمقارنة السهلة بين أساليب التصفية المختلفة.
النتائج
تظهر النتائج التجريبية المقدمة في هذا القسم فعالية نموذج SIKNet المقترح مقارنة بمرشحات كالمان القائمة على النموذج التقليدي (KF) ومرشحات التعلم المدعومة الأخرى عبر مجموعات اختبار مختلفة. كما هو موضح في الجدول 1، تحقق SIKNet تحسينًا في متوسط الاسترجاع (mAR) بنسبة تقارب 40% مقارنة بـ KF القائم على النموذج، مع ملاحظات أكبر تحت ظروف ضوضاء أعلى. من الجدير بالذكر أن SIKNet تتفوق على مرشحات التعلم المدعومة الموجودة، حيث تحقق تحسينًا بنسبة 6% في mAR وتحسينًا بنسبة 10% في متوسط الاسترجاع (mRe)، مما يشير إلى قدرتها الفائقة في تقديم تقديرات دقيقة لصناديق الحدود. يتم توضيح قوة النموذج بشكل أكبر في الشكل 5، الذي يوضح تقلبات mAR عبر مستويات ضوضاء مختلفة، مما يبرز أداء SIKNet المتسق.
يكشف التحليل الإضافي في الجدول 2 أن الأساليب المدعومة بالتعلم، بما في ذلك SIKNet، تتفوق بشكل كبير على KF القائم على النموذج عبر فئات مختلفة، حيث تظهر تحسينًا في الأداء بنسبة 50% لتتبع “المشاة”، بينما التحسينات لتتبع “الراقص” و”اللاعب” تبلغ حوالي 35%. يُعزى هذا التباين إلى أنماط الحركة الأكثر دراماتيكية للفئات الأخيرة، مما يقدم تحديات أكبر في عدم تطابق النموذج. تؤكد النتائج المرئية في الشكل 6 على اتساق SIKNet الفائق عبر أوضاع الحالة. بالإضافة إلى ذلك، في السيناريوهات التي تتضمن عدم تطابق معلمات الضوضاء في التدريب والاختبار، تتفوق SIKNet باستمرار على KF القائم على النموذج، كما هو موضح في الشكل 7. توضح النتائج المرئية المختارة من مجموعة بيانات DanceTrack فعالية SIKNet، حيث تحقق أعلى متوسط تقاطع على اتحاد (IoU) يبلغ 0.86، مع تباين ضئيل، مما يبرز استقرارها في معالجة التسلسلات الطويلة. بشكل عام، تؤكد النتائج أن SIKNet ومرشحات التعلم المدعومة الأخرى تتجاوز بشكل كبير نماذج KF التقليدية عبر مقاييس وظروف متعددة.
نقاش
في هذا القسم، يناقش المؤلفون تطوير وتقييم مرشح جديد مدعوم بالتعلم، SIKNet، مصمم لتقدير الحركة بشكل قوي في تتبع الأجسام المتعددة (MOT). تتضمن SIKNet مشفرًا مستقلًا عن الدلالة (SIE) يستخدم طبقة تلافيف أحادية البعد وطبقة متصلة بالكامل لفصل عمليات تضمين الميزات بشكل فعال بين العناصر الدلالية المتجانسة وغير المتجانسة داخل متجه الحالة. يعالج هذا النهج مشكلات عدم استقرار التدريب التي تواجه عادةً في مرشحات التعلم المدعومة الموجودة، مثل KalmanNet (KNet)، والتي تنشأ من المعلومات الدلالية المتنوعة والمقاييس العددية لعناصر الحالة.
يؤكد المؤلفون صحة SIKNet من خلال التجارب على مجموعات بيانات شبه محاكاة مشتقة من مجموعات بيانات MOT المعروفة، مما يوضح أن SIKNet تتفوق بشكل كبير على كل من مرشح كالمان القائم على النموذج التقليدي ومرشحات التعلم المدعومة الموجودة. على وجه التحديد، تحقق SIKNet تحسينًا بنسبة 6% في متوسط الاسترجاع (mAR) مقارنة بمرشحات التعلم المدعومة الأخرى وتحسينًا ملحوظًا بنسبة 40% مقارنة بمرشح كالمان القائم على النموذج. على الرغم من هذه التقدمات، يعترف المؤلفون بأن SIKNet لم يتم دمجه بعد بشكل كامل في أطر MOT الشاملة للتدريب من البداية إلى النهاية، مما يشير إلى اتجاه للأبحاث المستقبلية يركز على تحسين دمج مرشحات التعلم المدعومة ضمن أنظمة التتبع الكاملة لتعزيز الأداء العام.
DOI: https://doi.org/10.1016/j.inffus.2026.104513
Publication Date: 2026-06-01
Author(s): Jian Song et al.
Primary Topic: Video Surveillance and Tracking Methods
Overview
In this section, the authors present a novel approach to motion estimation in multi-object tracking (MOT) through the development of the Semantic-Independent KalmanNet (SIKNet). This method addresses the limitations of the traditional Kalman filter (KF), particularly in scenarios involving nonstationary object movements and mismatched parameters. SIKNet employs a Semantic-Independent Encoder (SIE) that utilizes a two-step process: first, a 1D convolution with a kernel size of 1 extracts independent semantic information from homogeneous elements, followed by a fully-connected layer that captures nonlinear relationships among heterogeneous elements.
To evaluate the effectiveness of SIKNet, the authors constructed a large-scale semi-simulated dataset derived from existing MOT datasets. Experimental results indicate that SIKNet significantly outperforms the traditional KF and other learning-aided filters, achieving a 6% improvement in mean Average Recall (mAR) over existing models and a remarkable 40% improvement over the KF. The authors acknowledge that while SIKNet demonstrates robust performance, challenges remain in integrating it into complete MOT frameworks for end-to-end training with detection components. Future research will focus on achieving this integration to enhance overall tracking performance.
Introduction
The introduction of this research paper highlights the significance of online and real-time multi-object tracking (MOT) in applications such as autonomous driving and motion analysis. It emphasizes the growing interest in tracking-by-detection (TBD) methods, which utilize deep learning-based detectors to track detection bounding boxes. These methods typically consist of two main components: the Motion Estimation Module (MEM) and the Data Association Module (DAM). The performance of these trackers is heavily reliant on the MEM’s predictions, which often employ a model-based Kalman filter (KF) based on a linear constant-velocity (CV) model. However, this assumption frequently fails in real-world scenarios due to unpredictable noise and complex object motion patterns, leading to tracking failures.
To address the limitations of traditional KF approaches, the authors propose KalmanNet, a learning-aided Kalman filtering (LAKF) model, and introduce an enhanced version called Semantic-Independent KalmanNet (SIKNet). These models leverage recurrent neural networks (RNNs) to adaptively learn Kalman gain from input features, improving motion estimation accuracy in non-linear and non-stationary conditions. The integration of a Semantic-Independent Encoder (SIE) further enhances robustness by effectively managing heterogeneous semantic information. The authors also construct a semi-simulated dataset for independent evaluation of the MEM, demonstrating that SIKNet outperforms traditional KF and existing learning-aided filters in terms of stability and accuracy. Key contributions include the development of SIKNet, the semi-simulated dataset, and an open-source framework, FilterNet, for easy comparison of various filtering methods.
Results
The experimental results presented in this section demonstrate the effectiveness of the proposed SIKNet model compared to traditional model-based Kalman Filters (KF) and other learning-aided filters across various test sets. As shown in Table 1, SIKNet achieves a mean average recall (mAR) improvement of approximately 40% over the model-based KF, with even greater advantages observed under higher noise conditions. Notably, SIKNet outperforms existing learning-aided filters, achieving a 6% improvement in mAR and a 10% improvement in mean recall (mRe), indicating its superior capability in providing accurate bounding-box estimates. The model’s robustness is further illustrated in Fig. 5, which visualizes mAR fluctuations across different noise levels, highlighting SIKNet’s consistent performance.
Further analysis in Table 2 reveals that learning-aided methods, including SIKNet, significantly outperform the model-based KF across various categories, particularly showing a 50% performance improvement for “Pedestrian” tracking, while improvements for “Dancer” and “Player” are around 35%. This discrepancy is attributed to the more dramatic motion patterns of the latter categories, which introduce greater model-mismatch challenges. Visual results in Fig. 6 confirm SIKNet’s superior consistency across state modes. Additionally, in scenarios with mismatched training and testing noise parameters, SIKNet consistently outperforms the model-based KF, as depicted in Fig. 7. Selected visual results from the DanceTrack dataset further illustrate SIKNet’s efficacy, achieving the highest average Intersection over Union (IoU) of 0.86, with minimal variance, underscoring its stability in processing long-term sequences. Overall, the results affirm that SIKNet and other learning-aided filters significantly surpass traditional KF models across multiple metrics and conditions.
Discussion
In this section, the authors discuss the development and evaluation of a novel learning-aided filter, SIKNet, designed for robust motion estimation in multi-object tracking (MOT). SIKNet incorporates a Semantic-Independent Encoder (SIE) that utilizes a 1D convolutional layer and a fully connected layer to effectively separate the feature-embedding processes of homogeneous and heterogeneous semantic elements within the state vector. This approach addresses the training instability issues commonly encountered in existing learning-aided filters, such as KalmanNet (KNet), which arise from the diverse semantic information and numerical scales of state elements.
The authors validate SIKNet through experiments on semi-simulated datasets derived from well-known MOT datasets, demonstrating that SIKNet significantly outperforms both the traditional model-based Kalman filter and existing learning-aided filters. Specifically, SIKNet achieves a 6% improvement in mean Average Recall (mAR) over other learning-aided filters and a remarkable 40% improvement over the model-based Kalman filter. Despite these advancements, the authors acknowledge that SIKNet has not yet been fully integrated into comprehensive MOT frameworks for end-to-end training, indicating a direction for future research focused on optimizing the integration of learning-aided filters within complete tracking systems to enhance overall performance.
