نموذج محسّن للتعرف على المشاعر الوجهية يعتمد على هيكل عميق جديد للتلافيف
Improved facial emotion recognition model based on a novel deep convolutional structure

شارك:
المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-79167-8
PMID: https://pubmed.ncbi.nlm.nih.gov/39580589
تاريخ النشر: 2024-11-23
المؤلف: Reham A. Elsheikh وآخرون
الموضوع الرئيسي: التعرف على العواطف والمزاج

نظرة عامة

تقدم ورقة البحث نهجًا مبتكرًا للتعرف على مشاعر الوجه (FER) من خلال تطوير شبكة تلافيفية عميقة مضادة للتشويش (AA-DCN). يتناول هذا النموذج التحديات الكامنة في FER، مثل التغيرات في تعبيرات الوجه، والاحتجاب، وظروف الإضاءة، التي غالبًا ما تؤثر على جودة الصورة. يكتشف نموذج AA-DCN بنجاح ثمانية مشاعر مميزة وتم تقييمه على ثلاثة مجموعات بيانات: مجموعة بيانات Cohn-Kanade الموسعة (CK+)، تعبيرات الوجه الأنثوية اليابانية (JAFFE)، ومجموعة بيانات الوجه العاطفي في العالم الحقيقي (RAF). حقق النموذج دقة ملحوظة بلغت 99.26% على CK+، و98% على JAFFE، و82% على RAF، مما يدل على تحسينات كبيرة في التعرف على المشاعر وتقليل عيوب التشويش.

في الختام، يتفوق نموذج AA-DCN على النماذج التقليدية من خلال دمج تقنيات مضادة للتشويش، لا سيما من خلال استخدام مرشحات تمرير منخفض مضبوطة في طبقات MaxPool. تشير النتائج التجريبية إلى أن هذا النهج يعزز معدلات التعرف، خاصة في السيناريوهات الصعبة مثل الوجوه المحجوبة. تهدف الأعمال المستقبلية إلى توسيع تطبيق عملية مضادة للتشويش لتشمل تعبيرات عاطفية أكثر تعقيدًا، مما يعزز قدرات الشبكات التلافيفية العميقة في مهام FER. تؤكد النتائج على إمكانية دمج مرشحات الضباب لتحسين الدقة في نماذج التعلم العميق للتعرف على المشاعر.

النتائج

في هذا القسم، تقدم البحث نتائج عدة دراسات تجريبية تستخدم نماذج تعلم عميقة متنوعة، بما في ذلك الشبكات التلافيفية العصبية التقليدية (CNNs) مثل Inception وVGG وResNet، بالإضافة إلى نموذج الشبكة التلافيفية العميقة المقترح (DCN) (الخوارزمية-1) ونموذج AA-DCN المتقدم (الخوارزمية-2). تم إجراء التجارب على ثلاثة مجموعات بيانات: CK+ وJAFFE وRAF، باستخدام مقاييس مثل مصفوفة الالتباس، والاسترجاع، والدقة، ودرجة F1 للتقييم. كانت البيئة الحاسوبية تتضمن جهاز كمبيوتر محمول بمعالج Intel Core i7 و8 جيجابايت من ذاكرة الوصول العشوائي، وتم تقييم التجارب أيضًا باستخدام واجهة برمجة تطبيقات Kaggle مع GPU100 لتحسين الأداء.

تم تحسين المعلمات الفائقة لنموذج AA-DCN كما هو موضح في الجدول 12، مع الحفاظ على نفس مقاييس التقييم كما ذُكر سابقًا. تم إجراء مرحلة التدريب بحجم دفعة قدره 32 ومعدل تعلم ثابت قدره 0.0006، مع تحديد عدد الدورات إلى 30، مما أسفر عن أفضل أداء مقارنة بالقيم الأخرى المختبرة. ستركز التحليلات المستقبلية على تأثير دمج طبقات MaxPool المضادة للتشويش في نموذج AA-DCN، مقارنة بأدائه مقابل نموذج DCN القياسي الذي يستخدم طبقات MaxPool العادية.

المناقشة

في مناقشة ورقة البحث، يتم استعراض مختلف التقدمات في أنظمة التعرف على مشاعر الوجه (FER) التي تستخدم تقنيات التعلم العميق. اقترح عمر وآخرون نظام FER من أربع خطوات يتضمن اكتشاف الوجه، وتعلم الميزات عبر شبكة عصبية تلافيفية (CNN)، وزيادة البيانات، محققًا أداءً متفوقًا مقارنة بالطرق الحالية. استكشف تشوداري وآخرون التعلم الانتقالي، ووجدوا أن نموذج MobileNet يتفوق في السرعة والكفاءة، على الرغم من أنه كان محدودًا باستخدام مجموعة بيانات واحدة. تناولت دراسات أخرى، مثل تلك التي أجراها أباتي وآخرون وشايك وآخرون، تحديات مثل الاحتجاب الناتج عن الأقنعة وطوروا نماذج تركز على الميزات المعتمدة على الانتباه، على التوالي، لكنهم واجهوا قيودًا في نطاق التشغيل وتنوع مجموعة البيانات.

تشمل المساهمات الإضافية نموذج Saurav وآخرون المدمج في الوقت الحقيقي Dual Integrated-CNN، الذي، على الرغم من فعالية التعرف على التعبيرات، واجه صعوبة في التصنيف الخاطئ لبعض المشاعر. قدم راجان وآخرون منهجية CNN هجينة أدت بشكل جيد في تمييز التعبيرات السعيدة والمفاجئة لكنها تعثرت في التعرف على الحزن والغضب. تسلط المناقشة الضوء على التحديات المستمرة في FER، مثل الحاجة إلى مجموعات بيانات أكبر وتحسين دقة التصنيف، خاصة في السيناريوهات العاطفية المعقدة. بشكل عام، تؤكد الورقة على إمكانية التعلم العميق في تعزيز FER بينما تحدد أيضًا المجالات الحرجة لمزيد من البحث والتطوير.

القيود

يسلط قسم القيود الضوء على عدة تحديات تواجهها أنظمة التعرف على مشاعر الوجه (FER) كما ورد في الدراسات الحديثة. لاحظ عمر وآخرون (2022) أن التوازن بين زيادة البيانات واستخراج الميزات العميقة يؤثر بشكل كبير على قدرات الكشف، خاصة مع عينات الاختبار غير المألوفة. حقق تشوداري وآخرون (2021) أداءً محسنًا من خلال إزالة الطبقات المتصلة بالكامل من النماذج المدربة مسبقًا؛ ومع ذلك، كانت تقييماتهم محدودة بمجموعة بيانات واحدة، مما أثار مخاوف بشأن القابلية للتعميم.

علاوة على ذلك، أكد أباتي وآخرون (2022) على الصعوبات التي تطرحها الاحتباسات، مثل الوجوه المmasked، التي تؤثر سلبًا على دقة FER. طور شايك وآخرون (2022) شبكة عصبية عميقة تدمج بين آليات الانتباه المحلية والعالمية، متفوقة على النماذج السابقة، ومع ذلك كانت مقاربتهم مقيدة بالصور الأمامية وبيانات الوجه غير المتغيرة في الوقت الحقيقي. أخيرًا، أظهر Saurav وآخرون (2022) كفاءة حاسوبية متفوقة ودقة في التعرف مقارنة بنماذج CNN الحالية، لكن أدائهم في التعرف على الخوف كان ضعيفًا بشكل ملحوظ. تؤكد هذه القيود على الحاجة إلى مجموعات بيانات أوسع ومنهجيات أكثر قوة لتعزيز فعالية أنظمة FER عبر ظروف متنوعة.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-79167-8
PMID: https://pubmed.ncbi.nlm.nih.gov/39580589
Publication Date: 2024-11-23
Author(s): Reham A. Elsheikh et al.
Primary Topic: Emotion and Mood Recognition

Overview

The research paper presents an innovative approach to Facial Emotion Recognition (FER) through the development of an anti-aliased deep convolution network (AA-DCN). This model addresses the inherent challenges of FER, such as variations in facial expressions, occlusions, and lighting conditions, which often degrade image quality. The AA-DCN model successfully detects eight distinct emotions and has been evaluated on three datasets: the Cohn-Kanade Extending (CK+), Japanese female facial expressions (JAFFE), and the Real-world Affective Face (RAF) dataset. The model achieved remarkable accuracies of 99.26% on CK+, 98% on JAFFE, and 82% on RAF, demonstrating significant improvements in emotion recognition and reduction of aliasing artifacts.

In conclusion, the AA-DCN model outperforms traditional models by incorporating anti-aliasing techniques, particularly through the use of tuned low-pass filters in the MaxPool layers. The experimental results indicate that this approach enhances recognition rates, particularly in challenging scenarios such as occluded faces. Future work aims to extend the application of the anti-aliasing process to more complex emotional expressions, thereby further advancing the capabilities of deep convolutional networks in FER tasks. The findings underscore the potential of integrating blur filters to improve accuracy in deep learning models for emotion recognition.

Results

In this section, the research presents the results of several experimental studies utilizing various deep learning models, including classical Convolutional Neural Networks (CNNs) such as Inception, VGG, and ResNet, as well as the proposed Deep Convolutional Network (DCN) model (Algorithm-1) and the advanced AA-DCN model (Algorithm-2). The experiments were conducted on three datasets: CK+, JAFFE, and RAF, employing metrics such as confusion matrix, recall, precision, and F1 score for evaluation. The computational environment included a laptop with an Intel Core i7 processor and 8 GB of RAM, and the experiments were also evaluated using the Kaggle API with a GPU100 for enhanced performance.

The hyper-parameters for the AA-DCN model were optimized as shown in Table 12, maintaining the same evaluation metrics as previously mentioned. The training phase was conducted with a batch size of 32 and a fixed learning rate of 0.0006, with the number of epochs set to 30, which yielded the best performance compared to other tested values. Future analyses will focus on the impact of incorporating antialiased MaxPool layers in the AA-DCN model, comparing its performance against the standard DCN model that utilizes regular MaxPool layers.

Discussion

In the discussion of the research paper, various advancements in facial emotion recognition (FER) systems utilizing deep learning techniques are reviewed. Umer et al. proposed a four-step FER system that includes face detection, feature learning via a convolutional neural network (CNN), and data augmentation, achieving superior performance compared to existing methods. Chowdary et al. explored transfer learning, finding that the MobileNet model excelled in speed and efficiency, although it was limited by the use of a single dataset. Other studies, such as those by Abate et al. and Shaik et al., addressed challenges like occlusion from masks and developed models focusing on attention-based features, respectively, but faced limitations in operational scope and dataset diversity.

Further contributions include Saurav et al.’s real-time Dual Integrated-CNN model, which, despite effective expression recognition, struggled with misclassification in certain emotions. Rajan et al. introduced a hybrid CNN methodology that performed well in distinguishing joyful and surprised expressions but faltered in recognizing sadness and anger. The discussion highlights the ongoing challenges in FER, such as the need for larger datasets and improved classification accuracy, particularly in complex emotional scenarios. Overall, the paper underscores the potential of deep learning in advancing FER while also identifying critical areas for further research and development.

Limitations

The section on limitations highlights several challenges faced by various facial emotion recognition (FER) systems as reported in recent studies. Umer et al. (2022) noted that the balance between data augmentation and deep feature extraction significantly affects detection capabilities, particularly with unfamiliar test samples. Chowdary et al. (2021) achieved improved performance by removing fully connected layers from pre-trained models; however, their evaluation was limited to a single dataset, raising concerns about generalizability.

Further, Abate et al. (2022) emphasized the difficulties posed by occlusions, such as masked faces, which adversely affect FER accuracy. Shaik et al. (2022) developed a deep neural network that integrates both local and global attention mechanisms, outperforming previous models, yet their approach was restricted to frontal images and real-time invariant face data. Lastly, Saurav et al. (2022) demonstrated superior computing efficiency and recognition accuracy compared to existing CNN models, but their performance in recognizing fear was notably poor. These limitations underscore the need for broader datasets and more robust methodologies to enhance FER systems’ effectiveness across diverse conditions.

شارك: