البحث في التعرف على وجه الأغنام بناءً على تقنية دمج معلومات الميزات متعددة الأبعاد في بيئة معقدة
Research on herd sheep facial recognition based on multi-dimensional feature information fusion technology in complex environment

شارك:
المجلة: Frontiers in Veterinary Science، المجلد: 12
DOI: https://doi.org/10.3389/fvets.2025.1404564
PMID: https://pubmed.ncbi.nlm.nih.gov/40151568
تاريخ النشر: 2025-03-13
المؤلف: Fu Zhang وآخرون
الموضوع الرئيسي: دراسات سلوك الحيوان ورفاهيته

نظرة عامة

تقدم البحث نظام مراقبة ذكي ثلاثي المراحل لإدارة الماشية على نطاق واسع، مستفيدًا من تقنيات التعلم العميق لتعزيز مراقبة قطعان الأغنام. يتكون النظام من شبكة كشف متعددة الأجزاء لعد قطعان الأغنام، ونموذج تصنيف الوجه للتعرف على الهوية، وشبكة تحليل تعبير الوجه لتقييم الصحة. تشمل الابتكارات الرئيسية تعديلات على بنية YOLOv5s، مثل دمج كتلة دمج الالتفاف متعددة الروابط (MCFB) لتحسين استخراج الميزات وإدخال هيكل الالتفاف القابل لإعادة المعلمة (RepConv) لتعزيز كشف الأهداف الصغيرة والكثيفة. تم تحسين نموذج تصنيف الوجه من خلال استبدال الهيكل في المرحلة السادسة من GhostNet بآلية انتباه ذاتي قابلة للفصل مكانيًا (SSSA) مكونة من أربع طبقات، مما أدى إلى تحسينات كبيرة في الدقة.

تشير النتائج التجريبية إلى أن شبكة التعرف متعددة الأجزاء المحسنة حققت دقة كشف متوسطة قدرها 95.84%، مع زيادة قدرها 2.55% في متوسط الدقة (mAP) مقارنة بـ YOLOv5s الأصلية. وصلت شبكة تصنيف الوجه إلى دقة مجموعة الاختبار قدرها 98.9%، متجاوزة GhostNet بنسبة 3.1%، بينما حققت شبكة تحليل تعبير الوجه دقة مجموعة اختبار قدرها 99.2%، مما يمثل تحسينًا بنسبة 3.6% مقارنة بـ EfficientNet. لا يعزز النظام المقترح دقة وكفاءة مراقبة قطعان الأغنام فحسب، بل يسهل أيضًا التقييمات الصحية في الوقت الحقيقي، مما يساهم في إدارة الماشية الذكية. تؤكد النتائج على فعالية تقنيات استخراج الميزات المتقدمة وتحسين النماذج في مواجهة تحديات مراقبة الماشية.

مقدمة

تسلط المقدمة الضوء على أهمية عد قطعان الأغنام، والتعرف على الوجه، وتحليل الصحة في إدارة المزارع على نطاق واسع، مشددة على أدوارها في تعزيز خطط التربية، وتتبع الأغنام الفردية، وتقييم حالة الصحة. تعتبر هذه التقنيات محورية لتحسين كفاءة المزرعة ورفاهية الحيوانات. تواجه طرق التعرف التقليدية على الماشية، سواء كانت تلامسية أو غير تلامسية، قيودًا مثل قيود المسافة وأخطاء التسجيل اليدوي. بينما تم تطوير طرق غير تلامسية، وخاصة تلك التي تستخدم الخصائص الفسيولوجية، إلا أنها غالبًا ما تواجه صعوبات في البيئات المعقدة وتتطلب عمليات جمع بيانات معقدة.

أظهرت التطورات الأخيرة في الشبكات العصبية الالتفافية (CNNs) وعدًا في التعرف على الماشية، خاصة من خلال تقنية التعرف على الوجه، التي تعتبر غير تدخليه وبديهية. أظهرت دراسات متنوعة فعالية الشبكات العصبية الالتفافية في التعرف على الخنازير والأبقار، محققة معدلات دقة عالية. بالمقابل، تقدم أبحاث التعرف على وجه الأغنام أيضًا، مع دراسات ملحوظة تحقق دقة كشف تتجاوز 99%. ومع ذلك، لا تدمج التقنيات الحالية العد الفردي أو تحليل تعبير الوجه، وهي أمور حاسمة للتغذية الدقيقة ومراقبة الصحة. لمعالجة هذه الفجوات، يقترح البحث نظامًا ثلاثي المراحل للتعرف على الأغنام يتضمن التعرف متعدد الأجزاء لعد القطعان، والتعرف الفردي، وتقييم حالة الصحة من خلال تحليل تعبير الوجه.

طرق

في هذه الدراسة، استخدم المؤلفون لقطات فيديو لوجوه الأغنام تم التقاطها في بيئة داخلية خاضعة للرقابة في شركة لويوانغ شيانغشون للتكنولوجيا الزراعية وتربية الحيوانات في مقاطعة خنان، الصين. تضمنت الموضوعات التجريبية 30 خروفًا بالغًا مصنفة إلى أربع حالات متميزة: غير حامل، مخصب، حامل، وما بعد الولادة. لتسهيل التحليل، تم استخدام كاميرا كانون لتتبع الفيديو نصف الشهري للأغنام داخل قطيعها، حيث استمرت كل جلسة تسجيل لمدة لا تقل عن 30 دقيقة وتعمل بمعدل إطار يبلغ 30 إطارًا في الثانية. يتم توضيح مشاهد التصوير المحددة في الشكل 1.

نتائج

تقدم نتائج هذه الدراسة تقييمًا شاملاً لمختلف خوارزميات التعرف على الأغنام، مع التركيز على أداء نموذج التعرف متعدد الأجزاء المحسن المستند إلى YOLOv5. استخدم التحليل مجموعة بيانات شائعة لمقارنة عدة خوارزميات كشف، بما في ذلك YOLOv4-tiny وYOLOv5-x وYOLOv5-s، عبر مقاييس مثل الإطارات في الثانية (FPS) ومتوسط الدقة (mAP) والمعلمات واستخدام الذاكرة. من الجدير بالذكر أن YOLOv5-s تم اختياره كنموذج أساسي بسبب أدائه المتوازن، حيث حقق mAP قدره 93.29% وFPS قدره 64. أدت التحسينات التي تم إجراؤها على YOLOv5s، مما أسفر عن نماذج YOLOv5s-MCFB وYOLOv5s-MCFB+، إلى قيم mAP قدرها 95.77% و95.84%، على التوالي، مع الحفاظ على معدلات FPS معقولة قدرها 57 و53. أظهر نموذج YOLOv5s-MCFB+ دقة عد استثنائية، حيث تم التعرف على 1,161 من أصل 1,173 خروفًا بدقة كشف قدرها 98.97%.

بالإضافة إلى نموذج التعرف متعدد الأجزاء، قامت الدراسة أيضًا بتقييم شبكة تصنيف الوجه المحسنة باستخدام هياكل متنوعة، بما في ذلك GhostNet ومتغيراته. حقق نموذج GhostNetVitm، الذي دمج هيكل LSA-GSA ذو الطبقتين وآلية الانتباه متعددة الرؤوس المكونة من أربع طبقات، أعلى مقاييس الأداء، مع معدلات دقة واسترجاع ودقة قدرها 99.4% و99.6% و98.7%، على التوالي. علاوة على ذلك، أظهرت شبكة تصنيف تعبير الوجه المستندة إلى EfficientNet تحسينات كبيرة في دقة الكشف والسرعة، خاصة مع EfficientNet0.5، الذي حقق دقة قدرها 99.5% وسرعة كشف قدرها 140 FPS. أشارت تصورات خريطة الحرارة إلى أن النموذج ركز بشكل فعال على المناطق الحيوية للوجه، مثل جسر الأنف ومناطق العين، والتي تعتبر حيوية للتفريق بين الأغنام الصحية والمريضة. بشكل عام، أدت التحسينات المقترحة في هذا البحث إلى تحسين كبير في دقة وكفاءة مهام التعرف على الأغنام وتصنيفها.

مناقشة

في هذه الدراسة، طور المؤلفون نظامًا شاملاً للتعرف على هوية وجه الأغنام وتحليل التعبير، معالجين التحديات مثل تشابه الصور والاحتجاب في بيئات التربية المعقدة. تم تنسيق مجموعة البيانات بعناية، باستخدام خوارزمية تجزئة إدراكية (d-Hash) لتصفية الصور المتشابهة، مما أسفر عن 3,078 صورة فعالة. شملت بنية النظام شبكة تعرف متعددة الأجزاء تعتمد على نموذج YOLOv5s المحسن، الذي تم تحسينه لاكتشاف وجوه الأغنام والظهور والأغنام الصغيرة. تضمنت التحسينات هيكل دمج ميزات الالتفاف متعددة الروابط وطبقة الالتفاف Repconv لتحسين دقة الكشف لأحجام الأهداف المتنوعة، خاصة في الظروف المحجوبة.

استخدمت شبكة تصنيف الوجه بنية GhostNet المعدلة، مع دمج آلية انتباه ذاتي قابلة للفصل مكانيًا (SSSA) لتعزيز استخراج الميزات من وجوه الأغنام، التي غالبًا ما تظهر تباينًا داخليًا ضئيلًا. تم بناء شبكة تحليل التعبير على إطار عمل EfficientNet، مع تطبيق تقنيات ضغط النموذج لتحسين سرعة الكشف ودقته. تم استخدام مقاييس التقييم مثل الدقة المتوسطة (AP) ومتوسط الدقة (mAP) والدقة لتقييم أداء النظام. تشير النتائج إلى أنه بينما تحدد شبكة التعرف متعددة الأجزاء ميزات الأغنام بشكل فعال، هناك حاجة إلى مزيد من التحسين لتعزيز دقة الكشف عن الوجه. بالإضافة إلى ذلك، تسلط الدراسة الضوء على الإمكانية لاستخدام ميزات منطقة العين لتقييم صحة الأغنام، مما يقترح نهجًا جديدًا لإدارة الصحة في تربية الأغنام.

Journal: Frontiers in Veterinary Science, Volume: 12
DOI: https://doi.org/10.3389/fvets.2025.1404564
PMID: https://pubmed.ncbi.nlm.nih.gov/40151568
Publication Date: 2025-03-13
Author(s): Fu Zhang et al.
Primary Topic: Animal Behavior and Welfare Studies

Overview

The research presents a three-phase intelligent monitoring system for large-scale livestock management, leveraging deep learning techniques to enhance the monitoring of sheep flocks. The system comprises a multi-part detection network for flock inventory counting, a facial classification model for identity recognition, and a facial expression analysis network for health assessment. Key innovations include modifications to the YOLOv5s architecture, such as the integration of a multi-link convolution fusion block (MCFB) for improved feature extraction and the introduction of a Re-Parameterizable Convolution (RepConv) structure to enhance the detection of small, dense targets. The facial classification model was optimized by replacing the sixth-stage structure of GhostNet with a four-layer spatially separable self-attention mechanism (SSSA), leading to significant improvements in accuracy.

Experimental results indicate that the enhanced multi-part identification network achieved an average detection accuracy of 95.84%, with a 2.55% increase in mean average precision (mAP) compared to the original YOLOv5s. The facial classification network reached a test set accuracy of 98.9%, surpassing GhostNet by 3.1%, while the facial expression analysis network attained a test set accuracy of 99.2%, representing a 3.6% improvement over EfficientNet. The proposed system not only enhances the accuracy and efficiency of sheep flock monitoring but also facilitates real-time health assessments, thereby contributing to intelligent livestock management. The findings underscore the effectiveness of advanced feature extraction and model optimization techniques in addressing the challenges of livestock monitoring.

Introduction

The introduction highlights the importance of sheep inventory counting, facial recognition, and health analysis in large-scale farm management, emphasizing their roles in enhancing breeding plans, tracking individual sheep, and assessing health status. These technologies are pivotal for improving farm efficiency and animal welfare. Traditional livestock identification methods, both contact and non-contact, face limitations such as distance constraints and manual registration errors. While non-contact methods, particularly those utilizing physiological characteristics, have been developed, they often struggle in complex environments and require intricate data collection processes.

Recent advancements in convolutional neural networks (CNNs) have shown promise in livestock recognition, particularly through facial recognition technology, which is non-invasive and intuitive. Various studies have demonstrated the effectiveness of CNNs for identifying pigs and cattle, achieving high accuracy rates. In contrast, sheep facial recognition research has also progressed, with notable studies achieving detection accuracies exceeding 99%. However, existing technologies do not integrate individual counting or facial expression analysis, which are crucial for precise feeding and health monitoring. To address these gaps, the paper proposes a three-stage system for sheep recognition that includes multi-part recognition for flock counting, individual identification, and health status assessment through facial expression analysis.

Methods

In this study, the authors utilized video footage of sheep faces captured in a controlled indoor environment at the Luoyang Xiangshun Agricultural and Animal Husbandry Technology Company in Henan Province, China. The experimental subjects comprised 30 adult sheep categorized into four distinct states: nulliparous, fertilized, pregnant, and postpartum. To facilitate the analysis, a Canon camera was employed for semimonthly video tracking of the sheep within their herd, with each recording session lasting a minimum of 30 minutes and operating at a frame rate of 30 frames per second. The specific shooting scenes are illustrated in Figure 1.

Results

The results of this study present a comprehensive evaluation of various recognition algorithms for sheep identification, focusing on the performance of an improved multi-part recognition model based on YOLOv5. The analysis utilized a common dataset to compare several detection algorithms, including YOLOv4-tiny, YOLOv5-x, and YOLOv5-s, across metrics such as frames per second (FPS), mean Average Precision (mAP), parameters, and memory usage. Notably, YOLOv5-s was selected as the baseline model due to its balanced performance, achieving an mAP of 93.29% and an FPS of 64. The enhancements made to YOLOv5s, resulting in models YOLOv5s-MCFB and YOLOv5s-MCFB+, yielded mAP values of 95.77% and 95.84%, respectively, while maintaining reasonable FPS rates of 57 and 53. The YOLOv5s-MCFB+ model demonstrated exceptional counting accuracy, identifying 1,161 out of 1,173 sheep with a detection accuracy of 98.97%.

In addition to the multi-part recognition model, the study also evaluated an improved facial classification network using various architectures, including GhostNet and its variants. The GhostNetVitm model, which incorporated a two-layer LSA-GSA structure and a four-layer Multi-Head Attention mechanism, achieved the highest performance metrics, with precision, recall, and accuracy rates of 99.4%, 99.6%, and 98.7%, respectively. Furthermore, the facial expression classification network based on EfficientNet demonstrated significant improvements in detection accuracy and speed, particularly with EfficientNet0.5, which achieved a precision of 99.5% and a detection speed of 140 FPS. The heatmap visualizations indicated that the model effectively focused on critical facial regions, such as the nose bridge and eye areas, which are vital for distinguishing between healthy and sick sheep. Overall, the enhancements proposed in this research significantly improved the accuracy and efficiency of sheep recognition and classification tasks.

Discussion

In this study, the authors developed a comprehensive system for sheep facial identity recognition and expression analysis, addressing challenges such as image similarity and occlusion in complex breeding environments. The dataset was meticulously curated, employing a perceptual hash algorithm (d-Hash) to filter out similar images, resulting in 3,078 effective images. The system architecture included a multi-part recognition network based on an improved YOLOv5s model, which was optimized for detecting sheep faces, backs, and lambs. Enhancements included a multi-link convolution feature fusion structure and a Repconv convolution layer to improve detection accuracy for varying target sizes, particularly in obstructed conditions.

The facial classification network utilized a modified GhostNet architecture, incorporating a four-layer spatially separable self-attention mechanism (SSSA) to enhance feature extraction from sheep faces, which often exhibit minimal intra-class variability. The expression analysis network was built on the EfficientNet framework, with model compression techniques applied to improve detection speed and accuracy. Evaluation metrics such as average precision (AP), mean average precision (mAP), and accuracy were employed to assess the system’s performance. The findings indicate that while the multi-part recognition network effectively identifies sheep features, further refinement is needed to enhance facial detection accuracy. Additionally, the study highlights the potential for using eye region features to assess sheep health, suggesting a novel approach to health management in sheep herding.

شارك: