تقسيم وتصنيف مقاطع أغاني الطيور في الوقت الحقيقي لتجارب التعلم
Real-Time Segmentation and Classification of Birdsong Syllables for Learning Experiments

شارك:
المجلة: eNeuro، المجلد: 13، العدد: 7
DOI: https://doi.org/10.1523/eneuro.0023-26.2026
PMID: https://pubmed.ncbi.nlm.nih.gov/42315338
تاريخ النشر: 2026-06-18
المؤلف: Zhenyun Du
الموضوع الرئيسي: التواصل الصوتي الحيواني والسلوك

نظرة عامة

في هذا القسم، يقدم المؤلفون Moove (تحديد عبر الإنترنت باستخدام فقط بدايات العناصر الصوتية)، وهي أداة جديدة تعتمد على الشبكات العصبية مصممة لتقسيم وتصنيف المقاطع الصوتية لأغاني طائر الفينش البنغالي في الوقت الحقيقي. كانت الطرق التقليدية للتدخلات المغلقة في طيور الأغاني محدودة في التعرف على المقاطع الفردية من خلال قوالب تم إنشاؤها يدويًا، والتي تفتقر إلى القدرة على التكيف مع الاحتياجات التجريبية المتنوعة. يستخدم Moove بنية من مرحلتين تكشف عن بدايات ونهايات المقاطع بينما تصنف المقاطع بناءً فقط على المعلومات الصوتية الأولية، مما يسهل التوقيت الدقيق بين سلوك الطائر والتغذية الراجعة المقدمة.

تم التحقق من فعالية Moove من خلال تجارب شملت خمسة طيور فينش بنغالية ذكور بالغة، مما يظهر قدرته على التوصيف السريع والدقيق للمقاطع الصوتية عبر الإنترنت. بالإضافة إلى ذلك، أجرى المؤلفون تجربة تعلم حيث تم تدريب أحد الفينشات على تعديل تسلسل مقاطعها استجابةً للتغذية الراجعة السمعية المموهة بالضوضاء. أظهرت النتائج أن الفينش تعلم بنجاح تجنب تسلسل المقاطع المستهدف، محققًا نتائج مقارنة بالدراسات السابقة في التعلم المعزز. بشكل عام، يمثل Moove تقدمًا كبيرًا في دراسة التعلم الصوتي في طيور الأغاني، حيث يقدم أداة موثوقة للتجارب المغلقة المستقبلية التي تركز على التلاعب بالإشارات الصوتية.

مقدمة

تناقش المقدمة أهمية طيور الأغاني، وخاصة طيور الفينش البنغالية، كنماذج لدراسة التواصل الصوتي المتعلم. تظهر هذه الطيور أصواتًا معقدة تتكون من مقاطع منفصلة يمكن ترتيبها في تسلسلات متنوعة وفقًا لقواعد تركيبية متعلمة. تم تحقيق تقدم حديث في فهم الآليات العصبية الكامنة وراء إنتاج الأغاني والتعلم من خلال تجارب مغلقة تتلاعب بالتغذية الراجعة السمعية والبصرية والحسية في الوقت الحقيقي. تتضمن طريقة ملحوظة استخدام التعلم المعزز لتعديل عناصر الأغاني عن طريق إخفاء المقاطع المستهدفة بضوضاء بيضاء مزعجة، مما يدفع الطيور لتكييف غنائها لتجنب تشغيل الضوضاء.

يسلط النص الضوء على التحديات المرتبطة بالحلول التقنية الحالية للتعرف على المقاطع في الوقت الحقيقي وتقديم التغذية الراجعة، وخاصة قيود الأساليب المعتمدة على القوالب التي تتطلب تعديلات يدوية. لمعالجة هذه القضايا، يقدم المؤلفون Moove (تحديد عبر الإنترنت باستخدام فقط بدايات العناصر الصوتية)، وهي أداة جديدة تستخدم مصنف ثنائي خفيف الوزن للكشف السريع عن بدايات المقاطع. يسمح هذا النهج بالتغذية الراجعة الفورية بينما لا تزال المقاطع تُنطق، مما يعزز كفاءة التجارب المغلقة. تم تصميم Moove ليكون سهل الاستخدام وقابلًا للتكيف، ويعمل دون الحاجة إلى أجهزة متخصصة، مما يسهل التوصيف الفوري لأغاني الطيور.

طرق البحث

تستخدم الدراسة Moove (تحديد عبر الإنترنت باستخدام فقط بدايات العناصر الصوتية)، وهي أداة تعتمد على بايثون مصممة لتقسيم وتصنيف مقاطع أغاني الطيور في الوقت الحقيقي من خلال الشبكات العصبية. تتكون هذه الأداة من مكونين رئيسيين: MooveTAF (التغذية الراجعة السمعية المستهدفة من Moove)، الذي يسهل تسجيل الأغاني واستهدافها في الوقت الحقيقي، وMooveGUI، الذي يدعم معالجة البيانات، ووضع العلامات اليدوية، وتدريب الشبكات العصبية.

يسمح دمج هذه المكونات بالتعامل الفعال مع بيانات أغاني الطيور، مما يمكّن الباحثين من تحليل الأنماط الصوتية بدقة وسرعة محسنتين. يشير استخدام الشبكات العصبية في هذا السياق إلى تقدم كبير في المنهجية لدراسة التواصل الطيور.

النتائج

في هذه الدراسة، قدمنا Moove، وهو نظام يعتمد على الشبكات العصبية مصمم لتقسيم وتصنيف مقاطع الفينش البنغالية في الوقت الحقيقي. يتكون النظام من شبكتين متخصصتين: واحدة تركز على كشف مقاطع المقاطع والأخرى مكلفة بتصنيف أنواع المقاطع. من الجدير بالذكر أن Moove يستخدم فقط أول 30 مللي ثانية من المعلومات الصوتية بعد بداية المقطع، مما يمكّن من التعرف الفعال واستهداف المقاطع أثناء إنتاجها المستمر. يظهر هذا النهج الإمكانية لمعالجة سمعية متقدمة في دراسات التواصل الطيور.

المناقشة

تناقش هذه القسم تطوير ووظائف MooveTAF، وهو برنامج يعتمد على بايثون مصمم لتسجيل الصوت في الوقت الحقيقي واستهداف مقاطع أغاني الطيور. البرنامج متاح عبر الإنترنت ويعمل على أجهزة الكمبيوتر المكتبية القياسية دون الحاجة إلى أجهزة متخصصة. يستخدم MooveTAF بنية ثنائية الشبكة: شبكة تقسيم تكشف عن مقاطع المقاطع في قطع صوتية قصيرة وشبكة تصنيف تصنف هذه المقاطع إلى أنواع المقاطع. يهدف هذا التصميم إلى تحسين كل من الدقة وسرعة الاستدلال، وهو أمر حاسم للتطبيقات في الوقت الحقيقي. تستخدم شبكة التقسيم مشفر تلافيفي وطبقة متعددة (MLP) لتصنيف إطارات الصوت، بينما تساعد خوارزمية النافذة المنزلقة في تحسين كشف بداية المقطع، مما يقلل من الإيجابيات الكاذبة.

يتضمن نظام MooveTAF أيضًا MooveGUI، واجهة رسومية لمعالجة البيانات، ووضع العلامات، وتدريب الشبكات العصبية. يمكن للمستخدمين تقسيم البيانات الصوتية يدويًا، وتدريب الشبكات على مجموعات الأغاني الفردية، وإجراء التصنيف والاستهداف في الوقت الحقيقي أثناء التجارب. يتم تدريب شبكة التصنيف على الطيفيات المشتقة من مقاطع المقاطع، محققة معدلات دقة عالية عبر عدة طيور فينش بنغالية. يبرز البحث قدرة النظام على التكيف مع ظروف تجريبية متنوعة، بما في ذلك تسلسلات الأهداف المحددة من قبل المستخدم ومحفزات التغذية الراجعة، مما يسهل التحكم الدقيق في التجارب المغلقة. بشكل عام، يمثل MooveTAF تقدمًا كبيرًا في التحليل الآلي لأغاني الطيور، مما يوفر للباحثين أداة قوية لدراسة التواصل الصوتي في الأنواع الطيور.

القيود

تسلط قيود تنفيذ Moove الحالي الضوء على عدة مجالات لتحسين محتمل. إن زمن تأخير التغذية الراجعة للنظام البالغ حوالي 40 مللي ثانية كافٍ لتعلم التسلسل ولكن قد يحتاج إلى تقليص لمهام مثل تعلم النغمة. ينشأ هذا التأخير من نافذة التصنيف الافتراضية البالغة 30 مللي ثانية و10 مللي ثانية من الحمل الزائد للنظام، والتي يمكن تقصيرها على حساب الدقة. يمكن للمستخدمين ضبط هذه المقايضة دون تعديل الكود. بالإضافة إلى ذلك، قد تؤدي تحسينات الأجهزة أو الهياكل الشبكية المتقدمة إلى تحسين الأداء دون تمديد النافذة الزمنية.

تتمثل قيود أخرى في التعرف الخاطئ أحيانًا لشبكة التقسيم على ضوضاء القفص كمقاطع، على الرغم من أن ذلك لم يؤثر على الاستهداف في الدراسة الحالية. يمكن أن تستفيد النسخ المستقبلية من دمج فئة ضوضاء مخصصة أثناء التدريب، كما هو الحال في أنظمة أخرى. ستعتمد قابلية تطبيق Moove على الإشارات الصوتية من أنواع مختلفة على هياكلها الصوتية؛ على سبيل المثال، من المحتمل أن تكون طيور الأغاني ذات مجموعات المقاطع المنفصلة، مثل طيور الفينش المخططة، مناسبة تمامًا لمنهجية Moove. في المقابل، قد تتطلب الأنواع ذات الأصوات الأكثر تنوعًا تعديلات كبيرة على عمليات التقسيم والتصنيف. علاوة على ذلك، تفتح قدرة Moove على الاستهداف في الوقت الحقيقي آفاقًا لتطبيقات تجريبية متنوعة، بما في ذلك التلاعبات الصوتية والبصرية المستندة إلى الأغاني، والتحفيز الميكروي العصبي، والاضطرابات الضوئية.

Journal: eNeuro, Volume: 13, Issue: 7
DOI: https://doi.org/10.1523/eneuro.0023-26.2026
PMID: https://pubmed.ncbi.nlm.nih.gov/42315338
Publication Date: 2026-06-18
Author(s): Zhenyun Du
Primary Topic: Animal Vocal Communication and Behavior

Overview

In this section, the authors introduce Moove (Marking Online using only the Onsets of Vocal Elements), a novel neural network-based tool designed for real-time syllable segmentation and classification of Bengalese finch songs. Traditional methods for closed-loop interventions in songbirds have been limited to recognizing single syllables through manually created templates, which lack adaptability for diverse experimental needs. Moove employs a two-stage architecture that detects syllable onsets and offsets while classifying syllables based solely on the initial acoustic information, facilitating precise timing between the bird’s behavior and the feedback provided.

The efficacy of Moove was validated through experiments involving five adult male Bengalese finches, demonstrating its capability for fast and accurate online annotation of syllables. Additionally, the authors conducted a learning experiment where one finch was trained to modify its syllable sequencing in response to auditory feedback masked by noise. The results indicated that the finch successfully learned to avoid the targeted syllable sequence, achieving outcomes comparable to previous reinforcement learning studies. Overall, Moove represents a significant advancement in the study of vocal learning in songbirds, offering a reliable tool for future closed-loop experiments focused on vocal signal manipulation.

Introduction

The introduction discusses the significance of songbirds, particularly Bengalese finches, as models for studying learned vocal communication. These birds exhibit complex vocalizations composed of discrete syllables that can be arranged in various sequences according to learned syntactic rules. Recent advancements in understanding the neuronal mechanisms underlying song production and learning have been achieved through closed-loop experiments that manipulate auditory, visual, and somatosensory feedback in real-time. A notable method involves using reinforcement learning to modify song elements by masking target syllables with aversive white noise, prompting birds to adapt their singing to avoid triggering the noise.

The text highlights the challenges associated with current technical solutions for real-time syllable recognition and feedback delivery, particularly the limitations of template-based approaches that require manual adjustments. To address these issues, the authors introduce Moove (Marking Online using only the Onsets of Vocal Elements), a novel tool that employs a lightweight binary classifier for rapid onset detection of syllables. This approach allows for immediate feedback while the syllable is still being vocalized, enhancing the efficiency of closed-loop experiments. Moove is designed to be user-friendly and adaptable, functioning without the need for specialized hardware, thereby facilitating real-time birdsong annotation.

Methods

The research employs Moove (Marking Online using only the Onsets of Vocal Elements), a Python-based tool designed for the real-time segmentation and classification of birdsong syllables through neural networks. This tool comprises two main components: MooveTAF (Moove Targeted Auditory Feedback), which facilitates song recording and real-time targeting, and MooveGUI, which supports data preprocessing, manual labeling, and the training of neural networks.

The integration of these components allows for efficient handling of birdsong data, enabling researchers to analyze vocal patterns with enhanced accuracy and speed. The use of neural networks in this context signifies a significant advancement in the methodology for studying avian communication.

Results

In this study, we introduced Moove, a neural network-based system designed for the real-time segmentation and classification of Bengalese finch syllables. The system comprises two specialized networks: one focused on detecting syllable segments and another tasked with classifying the types of syllables. Notably, Moove utilizes only the first 30 milliseconds of acoustic information following syllable onset, enabling effective identification and targeting of syllables during their ongoing production. This approach demonstrates the potential for advanced auditory processing in avian communication studies.

Discussion

The section discusses the development and functionality of MooveTAF, a Python-based software designed for real-time audio recording and targeting of birdsong syllables. The software is accessible online and operates on standard desktop PCs without the need for specialized hardware. MooveTAF employs a dual-network architecture: a segmentation network that detects syllable segments in short audio chunks and a classification network that categorizes these segments into syllable types. This design aims to optimize both accuracy and inference speed, crucial for real-time applications. The segmentation network utilizes a convolutional encoder and a multi-layer perceptron (MLP) to classify audio frames, while a sliding window algorithm helps refine syllable onset detection, reducing false positives.

The MooveTAF system also includes MooveGUI, a graphical interface for data preprocessing, labeling, and training of the neural networks. Users can manually segment audio data, train the networks on individual song repertoires, and perform real-time classification and targeting during experiments. The classification network is trained on spectrograms derived from syllable segments, achieving high accuracy rates across multiple Bengalese finches. The paper highlights the system’s ability to adapt to various experimental conditions, including user-defined target sequences and feedback stimuli, thus facilitating precise control in closed-loop experiments. Overall, MooveTAF represents a significant advancement in the automated analysis of birdsong, providing researchers with a robust tool for studying vocal communication in avian species.

Limitations

The limitations of Moove’s current implementation highlight several areas for potential enhancement. The system’s feedback latency of approximately 40 ms is adequate for sequence learning but may need to be reduced for tasks such as pitch learning. This latency arises from a 30 ms default classification window and 10 ms system overhead, which could be shortened at the expense of accuracy. Users can adjust this trade-off without modifying the code. Additionally, hardware optimizations or advanced network architectures might improve performance without extending the temporal window.

Another limitation is the segmentation network’s occasional misidentification of cage noise as syllable segments, although this did not impact targeting in the current study. Future iterations could benefit from incorporating a dedicated noise class during training, as seen in other systems. The applicability of Moove to vocal signals from different species will depend on their vocal structures; for instance, songbirds with discrete syllable repertoires, like zebra finches, are likely to be well-suited for Moove’s methodology. In contrast, species with more variable vocalizations may require significant modifications to the segmentation and classification processes. Furthermore, Moove’s real-time targeting capability opens avenues for diverse experimental applications, including song-triggered manipulations of auditory and visual feedback, neural microstimulation, and optogenetic perturbations.

شارك: