DOI: https://doi.org/10.1038/s41598-024-52989-2
PMID: https://pubmed.ncbi.nlm.nih.gov/38396002
تاريخ النشر: 2024-02-23
المؤلف: Anastasia Pentari وآخرون
الموضوع الرئيسي: التعرف على العواطف والمزاج
نظرة عامة
تتناول الورقة البحثية التقدم في التعرف على عواطف الكلام (SER) في سياق الحوسبة العاطفية، مع تسليط الضوء على تطبيق نظرية الرسم البياني لتصنيف إشارات الكلام الملونة عاطفياً. يقترح المؤلفون مجموعة ميزات جديدة مستمدة من المعلومات الإحصائية والبنائية القائمة على الرسم البياني، والتي تُستخدم لإنشاء هوية عاطفية فريدة لكل متحدث. تستخدم عملية التصنيف مصنف الغابة العشوائية تحت نظام التحقق المتقاطع لترك متحدث واحد (LOSO-CV). تُظهر التجارب المقارنة التي أُجريت على ثلاثة مجموعات بيانات—EMODB وAESDD وDEMoS—أن الطريقة المقترحة تتفوق على الأساليب الحالية الرائدة، محققة زيادات في معدلات الدقة غير الموزونة (UAR) تبلغ حوالي 18% و8% و13%، على التوالي.
في الختام، يدمج الدراسة بنجاح نظرية الرسم البياني في مجال SER، مما يعزز عملية التصنيف من خلال التركيز على الهويات العاطفية المستندة إلى المتحدث بدلاً من الأنماط التقليدية المستندة إلى الكلام. لا تقلل هذه الطريقة من المتطلبات الحاسوبية من خلال استخدام ميزات قائمة على الرسم البياني فحسب، بل تحقق أيضًا أداءً متفوقًا عبر مجموعات بيانات متنوعة. تشمل اتجاهات البحث المستقبلية تطبيق هذه المنهجية على الكلام والاضطرابات النفسية، بالإضافة إلى استكشاف دمج البيانات متعددة الوسائط، بما في ذلك المدخلات البصرية.
الطرق
في هذا القسم، يوضح المؤلفون منهجيتهم للتعرف على عواطف الكلام (SER) باستخدام نهج قائم على الرسم البياني يدمج كل من المعلومات البنائية والإحصائية من بيانات السلاسل الزمنية. تتضمن الخطوة الأولية بناء تمثيل هيكلي للرسم البياني للكلام من خلال نظرية الرسم البياني للرؤية (VG)، تليها حساب الميزات الإحصائية. ثم يعرف المؤلفون نموذجًا قائمًا على المتحدث يستخدم اللحظات الاحتمالية الأربعة الأولى لالتقاط الخصائص العاطفية الفريدة لكل متحدث.
تشمل التقييمات التجريبية مجموعات البيانات المستخدمة، والأساليب المقارنة، وأداء المنهجية المقترحة. يقارن المؤلفون نهجهم ضد طريقتين معروفتين في SER: واحدة تستخدم ميزات صوتية مصنوعة يدويًا (eGeMAPS) مع مصنفات التعلم الآلي (ML)، والأخرى تستخدم طيف ميل المعالج من خلال شبكة متبقية (ResNet). تتضمن مجموعة eGeMAPS 88 معلمة صوتية، بينما تطبق بنية ResNet سلسلة من الطبقات التلافيفية وعمليات التجميع لتصنيف العواطف بشكل فعال. تشير النتائج إلى أن المنهجية القائمة على الرسم البياني المقترحة، خاصة عند دمجها مع النمط القائم على المتحدث، حققت دقة تصنيف متفوقة، حيث وصلت إلى 77.8% (±16%) لمجموعة EMODB و79.1% (±20%) لمجموعة DEMoS، متفوقة على الأساليب المقارنة. ومن الجدير بالذكر أن دمج الأنماط القائمة على المتحدث مع ميزات eGeMAPS أدى إلى زيادة تقريبية بنسبة 20% في الدقة، مما يبرز أهمية هذه الميزات في تعزيز أداء SER.
المناقشة
في هذا القسم، يناقش المؤلفون التقدم في التعرف على عواطف الكلام (SER)، مع تسليط الضوء على التطور من طرق استخراج الميزات التقليدية إلى دمج هياكل التعلم العميق (DL). استخدمت الأساليب المبكرة مجموعات ميزات مثل Interspeech وGeMAPS وopenSMILE، مع الجمع بين مصنفات مثل آلات الدعم الشعاعي ونماذج ماركوف المخفية. ومع ذلك، أدت تعقيدات SER إلى تحول نحو تقنيات DL، بما في ذلك الشبكات العصبية التلافيفية والتعلم الانتقالي مع نماذج مدربة مسبقًا مثل الشبكات المتبقية. تشمل الابتكارات الحديثة أيضًا نماذج المحولات وتمثيلات متعددة الوسائط، على الرغم من أن هذه الطرق تتطلب حسابات مكثفة.
يقترح المؤلفون منهجية جديدة تستفيد من نظرية الرسم البياني لتحليل إشارات الكلام، مع التركيز على كل من المعلومات البنائية والإحصائية. يقدمون “نمط عاطفي قائم على المتحدث”، الذي يحسب اللحظات الإحصائية على الميزات القائمة على الرسم البياني لإنشاء هوية عاطفية فريدة لكل متحدث. تعالج هذه الطريقة التحديات مثل عدم توازن البيانات واختلاف أطوال الكلام، مما يعزز في النهاية دقة التصنيف. يتفوق خط الأنابيب المقترح على الطرق التقليدية، مما يظهر تحسينات كبيرة في التعرف على الحالات العاطفية عبر مجموعات بيانات متنوعة، بما في ذلك EMODB وAESDD وDEMoS. تشير النتائج إلى أن دمج الميزات القائمة على الرسم البياني مع تقنيات التعلم الآلي يمكن أن يؤدي إلى أنظمة SER أكثر فعالية، مما يمهد الطريق لتطبيقات مستقبلية في تقييمات الصحة النفسية.
DOI: https://doi.org/10.1038/s41598-024-52989-2
PMID: https://pubmed.ncbi.nlm.nih.gov/38396002
Publication Date: 2024-02-23
Author(s): Anastasia Pentari et al.
Primary Topic: Emotion and Mood Recognition
Overview
The research paper discusses advancements in Speech Emotion Recognition (SER) within the context of affective computing, highlighting the application of graph theory to classify emotionally-colored speech signals. The authors propose a novel feature set derived from graph-based statistical and structural information, which is utilized to create a unique emotional identity for each speaker. The classification process employs a Random Forest classifier under a Leave-One-Speaker-Out Cross Validation (LOSO-CV) scheme. Comparative experiments conducted on three datasets—EMODB, AESDD, and DEMoS—demonstrate that the proposed method outperforms existing state-of-the-art approaches, achieving average Unweighted Accuracy Rates (UAR) increases of approximately 18%, 8%, and 13%, respectively.
In conclusion, the study successfully integrates graph theory into the SER domain, enhancing the classification process by focusing on speaker-based emotional identities rather than traditional utterance-based motifs. This approach not only reduces computational demands through the use of fewer graph-based features but also achieves superior performance across various datasets. Future research directions include applying this methodology to speech and mental disorders, as well as exploring multimodal data integration, including visual inputs.
Methods
In this section, the authors detail their methodology for speech emotion recognition (SER) using a graph-based approach that integrates both structural and statistical information from time series data. The initial step involves constructing a structural graph representation of speech through Visibility Graph (VG) theory, followed by the computation of statistical features. The authors then define a speaker-based model utilizing the first four probabilistic moments to capture the unique emotional characteristics of each speaker.
The experimental evaluation encompasses the datasets utilized, comparative methods, and the performance of the proposed methodology. The authors benchmark their approach against two established SER methods: one using handcrafted acoustic features (eGeMAPS) with machine learning (ML) classifiers, and the other employing Mel-Spectrograms processed through a Residual Network (ResNet). The eGeMAPS set includes 88 acoustic parameters, while the ResNet architecture applies a series of convolutional layers and pooling operations to classify emotions effectively. The results indicate that the proposed graph-based methodology, particularly when combined with the speaker-based motif, achieved superior classification accuracy, reaching 77.8% (±16%) for EMODB and 79.1% (±20%) for DEMoS, outperforming the comparative methods. Notably, the integration of speaker-based motifs with eGeMAPS features led to an approximate 20% increase in accuracy, underscoring the significance of these features in enhancing SER performance.
Discussion
In this section, the authors discuss advancements in speech emotion recognition (SER), highlighting the evolution from traditional feature extraction methods to the integration of deep learning (DL) architectures. Early approaches utilized feature sets like Interspeech, GeMAPS, and openSMILE, combined with classifiers such as support vector machines and hidden Markov models. However, the complexity of SER prompted a shift towards DL techniques, including convolutional neural networks and transfer learning with pre-trained models like Residual Networks. Recent innovations also involve transformer models and multimodal representations, although these methods are computationally intensive.
The authors propose a novel methodology that leverages graph-based theory to analyze speech signals, focusing on both structural and statistical information. They introduce a “speaker-based emotional motif,” which computes statistical moments over graph-based features to create a unique emotional identity for each speaker. This approach addresses challenges such as data imbalance and varying utterance lengths, ultimately enhancing classification accuracy. The proposed pipeline outperforms traditional methods, demonstrating significant improvements in recognizing emotional states across various datasets, including EMODB, AESDD, and DEMoS. The findings suggest that combining graph-based features with machine learning techniques can lead to more effective SER systems, paving the way for future applications in mental health assessments.
