DOI: https://doi.org/10.1038/s41598-026-35858-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41611801
تاريخ النشر: 2026-01-29
المؤلف: Fereshteh Manafzadeh Heir وآخرون
الموضوع الرئيسي: التعرف على الكلام والتوليف
نظرة عامة
تستكشف هذه الدراسة بنية تعلم عميق هجينة تجمع بين الشبكات العصبية التلافيفية (CNNs) والتعلم المعزز (RL) لتحديد المتحدث، وهو عنصر حاسم في أنظمة المصادقة البيومترية. تم تقييم طريقتين لاستخراج الميزات: الطريقة 1، التي استخدمت تمثيلات طيف ميل مع آليات انتباه ذاتي، والطريقة 2، التي استخدمت تحويل الموجات المستمرة (CWT) مع موجات مورليت. تم اختبار كلا الطريقتين على مجموعة بيانات LibriSpeech dev-clean، التي تضم 2,703 ملف صوتي من 40 متحدثًا، باستخدام التحقق المتقاطع الطبقي 5-fold. أشارت النتائج إلى أن الطريقة 1 تفوقت على الطريقة 2، محققة دقة قدرها 87.60% (95% CI: [83.60%, 91.95%]) وROC-AUC قدره 99.54%، بينما حققت الطريقة 2 دقة قدرها 77.60% (95% CI: [73.12%, 82.08%]) مع ROC-AUC قدره 98.21%. أظهر تحليل ANOVA أن 21 من أصل 22 ميزة صوتية كانت لها قوة تمييزية كبيرة، حيث أظهرت الانتروبيا أقوى تأثير (F = 39.79, η² = 0.37).
سهل دمج التعلم Q في البنية اتخاذ قرارات واعية بالثقة، مما عزز الأداء مقارنة بأسس CNN فقط بفروق ذات دلالة إحصائية (الطريقة 1: +2.80%، p = 0.0142؛ الطريقة 2: +3.40%، p = 0.0089). تؤكد النتائج على أهمية تمثيلات مقياس ميل في التقاط الدقة الزمنية الدقيقة لأنماط الصوت الخاصة بالمتحدث، وهو أمر حاسم للتحديد القوي. على الرغم من النتائج الواعدة، تشير الدراسة إلى التبادلات الحسابية بين الطريقتين، حيث تتطلب الطريقة 1 وقت تدريب أقل بكثير من الطريقة 2. تشمل اتجاهات البحث المستقبلية استكشاف نماذج خفيفة الوزن والتحقق من الأداء عبر بيئات صوتية متنوعة لتعزيز التطبيق العملي للنظام المقترح في سيناريوهات العالم الحقيقي.
الطرق
في هذه الدراسة، تم تنفيذ استراتيجية تحقق متقاطع طبقي 5-fold لضمان تقديرات أداء قوية وتقليل الإفراط في التكيف. كانت مجموعة البيانات تتكون من 2,703 عينة صوتية من 40 متحدثًا، مع الحفاظ على توزيع المتحدثين الأصلي في كل طية. لكل طية \( k \in \{1, \ldots, 5\} \)، تم بناء مجموعة التدريب لضمان تمثيل متوازن للمتحدثين عبر جميع الفئات، المشار إليها بـ \( |D^{(k)}_{\text{train}}| \approx | \{ i : y_i = c \} | \) لكل فئة \( c \in \{1, \ldots, 20\} \). خضعت كل طية للتدريب لمدة 100 عصر، مع مراقبة دقة التحقق للاحتفاظ بحالة النموذج التي حققت أعلى دقة.
شملت المنهجية خط معالجة شامل، كما هو موضح في الشكل 6، الذي يوضح الخطوات من إدخال بيانات LibriSpeech من خلال المعالجة المسبقة، واستخراج الميزات باستخدام طيف ميل وتحويل الموجات المستمرة (CWT)، إلى تطبيق بنى قائمة على CNN ووحدة تعلم معزز هجينة. انتهت هذه الاستراتيجية التدريبية المشتركة إلى نتائج تحديد المتحدث النهائية، مما سمح بتقييم شامل لأداء النموذج عبر التنبؤات المجمعة من جميع مجموعات التحقق الخمس.
النتائج
قيمت الدراسة اثنتين من بنى الشبكات العصبية التلافيفية-التعلم المعزز (CNN-RL) الهجينة لتحديد المتحدث، باستخدام مجموعة بيانات LibriSpeech dev-clean، التي تشمل 2,703 ملف صوتي من 40 متحدثًا. أظهر تحليل ANOVA لـ 22 ميزة صوتية أن 21 ميزة كانت لها قوة تمييزية ذات دلالة إحصائية (p < 0.05)، حيث أظهرت الانتروبيا أقوى حجم تأثير (η² = 0.37). تدعم هذه النتيجة الفكرة القائلة بأن هوية المتحدث تمثل عبر أبعاد صوتية تكاملية متنوعة. من بين الطرق التي تم تقييمها، حققت الطريقة 1، التي تعتمد على طيف ميل، أعلى دقة قدرها 87.60% (95% CI: [83.60%, 91.95%]) وأظهرت قدرة تمييز استثنائية مع ROC-AUC قدره 99.54%. كان الأداء متسقًا عبر خمس طيات تحقق متقاطع، مع انحراف معياري قدره ± 3.53%. بالإضافة إلى ذلك، أشارت استخدام المقاييس الموزونة إلى توزيعات أداء طبيعية إحصائيًا واتفاق كبير يتجاوز الصدفة (Kappa > 0.76)، مما يثبت فعالية أساليب التعلم المعزز العميق الهجينة لتحديد المتحدث بشكل قوي.
المناقشة
في هذه الدراسة، استخدم المؤلفون مجموعة بيانات LibriSpeech “dev-clean”، وهي مجموعة مرجعية لتعرف الكلام التلقائي (ASR)، تضم حوالي 1,000 ساعة من صوت الكلام الإنجليزي عالي الجودة. تم تنسيق مجموعة البيانات بعناية، حيث تضم 2,703 ملف صوتي من 40 متحدثًا فريدًا، بمتوسط مدة قدرها 7.18 ثانية لكل ملف. شملت خط معالجة البيانات تحميل الصوت، وإزالة الصمت، وتطبيع الإشارة، وزيادة البيانات، واستخراج الميزات الخاصة بالطريقة، باستخدام كل من تقنيات طيف ميل وتحويل الموجات المستمرة (CWT). تدمج البنية الهجينة الشبكات العصبية التلافيفية العميقة (CNNs) مع التعلم المعزز (RL) لتعزيز أداء تحديد المتحدث، مما يسمح للنموذج بتحسين قرارات التصنيف بشكل تكيفي بناءً على ثقة التنبؤ.
يتكون النظام المقترح من ثلاثة مكونات رئيسية: مستخرج ميزات، وشبكة Q لاختيار الإجراءات، ورأس تصنيف. تعالج بنية طيف ميل المدخلات ذات الطول المتغير من خلال إطار عمل تلافيفي ثلاثي المراحل، بينما تستخدم بنية CWT تسلسلًا أعمق من أربع مراحل للمدخلات ذات الحجم الثابت. تستفيد كلا البنيتين من آليات الانتباه الذاتي وتنظيم الإسقاط لتحسين التعميم. يجمع الهدف التدريبي بين التصنيف الخاضع للإشراف مع تقدير الثقة القائم على التعلم المعزز، مما يحسن معلمات النموذج باستخدام مُحسِّن AdamW مع جدول زمني لمعدل تعلم تراجع جيبي. كما أن الدراسة أنشأت أيضًا قاعدة بيانات CNN فقط للمقارنة، مما يسمح بتقييم كمي لمساهمة مكون التعلم المعزز في أداء النموذج، الذي تم تقييمه من خلال مجموعة شاملة من المقاييس بما في ذلك الدقة، والدقة، والاسترجاع، والتدابير الإحصائية المتقدمة.
القيود
تسلط قسم القيود الضوء على عدة قيود حاسمة مرتبطة ببنى CNN-RL الهجينة المستخدمة في الدراسة. أولاً، مجموعة البيانات، التي تتكون من 2,703 عينة صوتية من 40 متحدثًا بمدة إجمالية قدرها 5.39 ساعة، صغيرة نسبيًا مقارنةً بمعايير أكبر في تحديد المتحدث. قد تعيق هذه النطاق المحدود قدرة النماذج على التعميم بشكل فعال عبر ظروف صوتية متنوعة، وبيئات التسجيل، وخصائص المتحدثين خارج مجموعات بيانات اللغة الإنجليزية.
ثانيًا، تؤدي استخدام تمثيلات صورة تحويل الموجات المستمرة (CWT) الثابتة 128 × 128 في الطريقة 2 إلى فقدان الدقة الزمنية بسبب الاستيفاء الثنائي، مما يؤدي إلى انخفاض بنسبة 10% في الدقة مقارنةً بالطريقة 1. يشير هذا إلى الحاجة إلى طرق معالجة مسبقة بديلة تحافظ بشكل أفضل على الديناميات الزمنية الدقيقة. بالإضافة إلى ذلك، هناك تفاوت ملحوظ في الكفاءة الحسابية بين الطريقتين، حيث تتطلب الطريقة 1 6 ساعات وتستغرق الطريقة 2 24 ساعة، مما يشير إلى مشكلات محتملة في قابلية التوسع لأساليب CWT في التطبيقات الزمنية الحقيقية. أخيرًا، على الرغم من أن مكون التعلم المعزز قدم إطارًا مفاهيميًا جديدًا، إلا أن مساهمته في الأداء العام كانت متواضعة، مما يشير إلى أن المزيد من الاستكشاف ضروري لتحقيق التوازن بين خسارة التصنيف وأهداف التعلم Q عبر مهام تحديد المتحدث المختلفة.
DOI: https://doi.org/10.1038/s41598-026-35858-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41611801
Publication Date: 2026-01-29
Author(s): Fereshteh Manafzadeh Heir et al.
Primary Topic: Speech Recognition and Synthesis
Overview
This study investigates a hybrid deep learning architecture that combines Convolutional Neural Networks (CNNs) with Reinforcement Learning (RL) for speaker identification, a crucial component of biometric authentication systems. Two feature extraction methods were evaluated: Method 1, which utilized Mel-spectrogram representations with self-attention mechanisms, and Method 2, which employed Continuous Wavelet Transform (CWT) with Morlet wavelets. Both methods were tested on the LibriSpeech dev-clean dataset, comprising 2,703 audio files from 40 speakers, using stratified 5-fold cross-validation. The results indicated that Method 1 outperformed Method 2, achieving an accuracy of 87.60% (95% CI: [83.60%, 91.95%]) and a ROC-AUC of 99.54%, while Method 2 achieved 77.60% accuracy (95% CI: [73.12%, 82.08%]) with a ROC-AUC of 98.21%. ANOVA analysis revealed that 21 out of 22 acoustic features had significant discriminative power, with entropy showing the strongest effect (F = 39.79, η² = 0.37).
The integration of Q-learning in the architecture facilitated confidence-aware decision-making, enhancing performance over CNN-only baselines by statistically significant margins (Method 1: +2.80%, p = 0.0142; Method 2: +3.40%, p = 0.0089). The findings underscore the importance of Mel-scale representations in capturing fine temporal resolution of speaker-specific acoustic patterns, which is critical for robust identification. Despite the promising results, the study notes the computational trade-offs between the methods, with Method 1 requiring significantly less training time than Method 2. Future research directions include exploring lightweight models and validating performance across diverse acoustic environments to enhance the practical applicability of the proposed system in real-world scenarios.
Methods
In this study, a stratified 5-fold cross-validation strategy was implemented to ensure robust performance estimates and mitigate overfitting. The dataset comprised 2,703 audio samples from 40 speakers, with each fold maintaining the original speaker distribution. For each fold \( k \in \{1, \ldots, 5\} \), the training set was constructed to ensure balanced representation of speakers across all classes, denoted as \( |D^{(k)}_{\text{train}}| \approx | \{ i : y_i = c \} | \) for each class \( c \in \{1, \ldots, 20\} \). Each fold underwent training for 100 epochs, with validation accuracy monitored to retain the model state that achieved the highest accuracy.
The methodology included a comprehensive processing pipeline, as illustrated in Figure 6, which details the steps from LibriSpeech data input through preprocessing, feature extraction using Mel-spectrogram and Continuous Wavelet Transform (CWT), to the application of CNN-based architectures and a hybrid reinforcement learning module. This joint training strategy culminated in the final speaker identification results, allowing for a thorough evaluation of model performance across the aggregated predictions from all five validation sets.
Results
The study assessed two hybrid convolutional neural network-reinforcement learning (CNN-RL) architectures for speaker identification, utilizing the LibriSpeech dev-clean dataset, which includes 2,703 audio files from 40 speakers. An ANOVA analysis of 22 acoustic features indicated that 21 features had statistically significant discriminative power (p < 0.05), with entropy showing the strongest effect size (η² = 0.37). This finding supports the notion that speaker identity is represented across various complementary acoustic dimensions. Among the evaluated methods, Method 1, which is based on Mel-spectrograms, achieved the highest accuracy of 87.60% (95% CI: [83.60%, 91.95%]) and demonstrated exceptional discriminative capability with a ROC-AUC of 99.54%. The performance was consistent across five cross-validation folds, exhibiting a standard deviation of ± 3.53%. Additionally, the use of weighted metrics indicated statistically normal performance distributions and substantial agreement beyond chance (Kappa > 0.76), thereby validating the efficacy of hybrid deep reinforcement learning approaches for robust speaker identification.
Discussion
In this study, the authors utilized the LibriSpeech “dev-clean” subset, a benchmark corpus for automatic speech recognition (ASR), comprising approximately 1,000 hours of high-quality English speech audio. The dataset was carefully curated, featuring 2,703 audio files from 40 unique speakers, with an average duration of 7.18 seconds per file. The preprocessing pipeline included audio loading, silence removal, signal normalization, data augmentation, and method-specific feature extraction, employing both Mel-spectrogram and Continuous Wavelet Transform (CWT) techniques. The hybrid architecture integrates deep convolutional neural networks (CNNs) with reinforcement learning (RL) to enhance speaker identification performance, allowing the model to adaptively optimize classification decisions based on prediction confidence.
The proposed system consists of three main components: a feature extractor, a Q-network for action selection, and a classification head. The Mel-spectrogram architecture processes variable-length inputs through a three-stage convolutional framework, while the CWT architecture employs a deeper four-stage hierarchy for fixed-size inputs. Both architectures leverage self-attention mechanisms and dropout regularization to improve generalization. The training objective combines supervised classification with RL-based confidence estimation, optimizing model parameters using the AdamW optimizer with a cosine annealing learning rate schedule. The study also established a CNN-only baseline for comparison, allowing for a quantitative assessment of the RL component’s contribution to model performance, evaluated through a comprehensive set of metrics including accuracy, precision, recall, and advanced statistical measures.
Limitations
The section on limitations highlights several critical constraints associated with the hybrid CNN-RL architectures employed in the study. Firstly, the dataset, comprising 2,703 audio samples from 40 speakers with a total duration of 5.39 hours, is relatively small compared to larger benchmarks in speaker identification. This limited scale may hinder the models’ ability to generalize effectively across varied acoustic conditions, recording environments, and speaker demographics beyond English-language datasets.
Secondly, the use of fixed 128 × 128 Continuous Wavelet Transform (CWT) image representations in Method 2 results in a loss of temporal resolution due to bilinear interpolation, leading to a 10% decrease in accuracy compared to Method 1. This suggests a need for alternative preprocessing methods that better preserve fine-grained temporal dynamics. Additionally, there is a notable disparity in computational efficiency between the two methods, with Method 1 requiring 6 hours and Method 2 taking 24 hours, indicating potential scalability issues for CWT-based approaches in real-time applications. Lastly, while the reinforcement learning component introduced a novel conceptual framework, its contribution to overall performance was modest, indicating that further exploration is necessary to optimize the balance between classification loss and Q-learning objectives across various speaker identification tasks.
