DOI: https://doi.org/10.1016/j.ecoinf.2026.103765
تاريخ النشر: 2026-04-27
المؤلف: Raphael Schwinger وآخرون
الموضوع الرئيسي: التواصل الصوتي الحيواني والسلوك
نظرة عامة
تقدم هذه القسم نظرة شاملة وتحليلًا للتحليل البيولوجي الصوتي الآلي، مع التأكيد على ضرورة نماذج التعلم العميق المتقدمة لمراقبة التنوع البيولوجي والحفاظ عليه. يستعرض المقال نماذج الأساس البيولوجي الصوتي المدربة مسبقًا على نطاق واسع، مع التركيز على قابليتها للنقل عبر مهام التصنيف المختلفة. يوضح بيانات التدريب، والمعالجة المسبقة، والتعزيزات، والهندسة المعمارية، ونماذج التدريب لهذه النماذج، جنبًا إلى جنب مع دراسة تجريبية تقيم أدائها على معايير BEANS وBirdSet. تشير النتائج الرئيسية إلى أن Perch 2.0 يتفوق بأعلى درجة BirdSet (90.78 AUROC) وأفضل نتيجة استكشاف خطي على BEANS (97.78 AUROC)، مما يوضح فعالية التدريب المسبق المراقب المتنوع عبر الأنواع. بالإضافة إلى ذلك، تم تحديد BirdMAE كنموذج رائد لاستراتيجيات الاستكشاف على BirdSet، بينما تتفوق نماذج الصوت العامة المدربة على AudioSet على نماذج أصوات الطيور المتخصصة على BEANS.
تلخص الخاتمة النتائج المستخلصة من التحليل المقارن لثلاثة عشر نموذجًا أساسيًا، مع تسليط الضوء على أهمية الاستكشاف الدقيق في تعظيم أداء النماذج المعتمدة على المحولات. تقترح أن Perch 2.0 هو خيار افتراضي قوي لتصنيف البيولوجيا الصوتية نظرًا لأدائه العالي وعدد المعلمات المنخفض، مما يجعله مناسبًا للتطبيقات ذات الموارد المحدودة. يشير التحليل أيضًا إلى BEATs NLM للتصنيف عبر الأنواع وBirdMAE لتصنيف أصوات الطيور مع الاستكشاف الدقيق. تشمل اتجاهات البحث المستقبلية تطوير نماذج على نطاق واسع مدربة على بيانات المراقبة الصوتية السلبية واستكشاف استراتيجيات التكيف المتقدمة لمهام البيولوجيا الصوتية.
مقدمة
تؤكد المقدمة على الدور الحاسم لمراقبة التنوع البيولوجي في إبلاغ استراتيجيات الحفظ وفهم الديناميات البيئية. تبرز أن مثل هذه المراقبة ضرورية لاكتشاف التغيرات في أعداد الأنواع والعمليات البيئية، مما يساعد بدوره في تدابير الإدارة والحماية الفعالة. على وجه الخصوص، تعتبر مراقبة أعداد الطيور مؤشرًا على التحولات الأوسع في التنوع البيولوجي.
يتم تقديم المراقبة الصوتية السلبية (PAM) كطريقة قابلة للتوسع وأقل تدخلاً لتتبع الأنواع المنتجة للصوت، وخاصة تلك النادرة أو التي يصعب مسحها بسبب عاداتها الليلية أو طبيعتها الخفية. لقد أظهرت PAM أنها قادرة على اكتشاف عدد مكافئ أو أكبر من الأنواع مقارنةً بالاستطلاعات التقليدية التي يقودها الخبراء، مع توفير تغطية زمنية واسعة. ومع ذلك، أدت تحديات توضيح كميات هائلة من بيانات الصوت إلى التركيز على طرق الكشف الآلي عن الأنواع. ظهرت نماذج التعلم العميق (DL) كأدوات قوية لهذا الغرض، حيث أظهرت دقة عالية في تصنيف الأنواع حسب الصوت. ومع ذلك، فإن فعالية أساليب DL تعتمد على توفر مجموعات بيانات كبيرة موضحة، مما يشكل قيدًا على الأنواع النادرة والتي لم يتم دراستها بشكل كافٍ.
طرق
في هذا القسم، يوضح المؤلفون تصميمهم التجريبي، بدءًا من وصف المهام المرجعية المختارة للتقييم. هذه المهام حاسمة لتقييم أداء النماذج قيد التحقيق. بعد ذلك، يوضح المؤلفون عملية اختيار النموذج، والتي تتضمن معايير لاختيار النماذج الأكثر ملاءمة للمهام المطروحة. أخيرًا، يتم تقديم بروتوكول التدريب لتكييف النموذج، مع تسليط الضوء على المنهجيات المستخدمة لضبط النماذج لتحقيق الأداء الأمثل على المعايير المختارة. تضمن هذه الطريقة المنظمة تقييمًا شاملاً لقدرات النماذج.
نتائج
تسلط نتائج الدراسة الضوء على أداء نماذج البيولوجيا الصوتية المختلفة على معايير BEANS وBirdSet، باستخدام استراتيجيات الاستكشاف الخطي والدقيق. تشير درجات AUROC إلى أن النماذج المدربة على مجموعات بيانات الصوت العامة، مثل BEATs وPerch 2.0، تتفوق على العديد من النماذج المتخصصة في البيولوجيا الصوتية. من الجدير بالذكر أن BEATs NLM حقق أعلى درجة على BEANS (98.57 AUROC) مع الاستكشاف الدقيق، بينما تفوق Perch 2.0 على BirdSet (90.78 AUROC) مع تقييم مقيد، مما يوضح فعالية التدريب المتنوع عبر الأنواع. تشير النتائج إلى أن أساليب التعلم الذاتي المتقدمة (SSL) يمكن أن تنتج تمثيلات قوية قابلة للتطبيق على مهام البيولوجيا الصوتية، كما يتضح من الأداء المتفوق لنماذج مثل BirdMAE (86.54 AUROC) مقارنةً بتلك المدربة فقط على بيانات البيولوجيا الصوتية.
كما تكشف التحليلات أن معدلات العينة الأعلى (32 كيلوهرتز) وأطوال النوافذ المحسّنة (5 ثوانٍ) تعزز بشكل كبير من أداء النموذج، خاصةً لمهام تصنيف الطيور. بالمقابل، كانت النماذج التي تستخدم الموجات الصوتية الخام تؤدي بشكل أقل باستمرار مقارنةً بتلك المعتمدة على الطيف. تؤكد الدراسة أنه بينما يعتبر حجم مجموعة البيانات مهمًا، فإن جودة تنسيق البيانات واختيار نماذج التدريب أمران حاسمان لتحقيق نتائج مثالية. تشمل التوصيات لتطوير النماذج المستقبلية الاستفادة من مجموعات البيانات المعروفة، واستكشاف الجمع بين البيانات العامة وبيانات البيولوجيا الصوتية، والنظر في الكفاءة الحاسوبية في اختيار النموذج. تختتم الدراسة بأن Perch 2.0 وBirdMAE هما مرشحان قويان لتطبيقات البيولوجيا الصوتية، مع إمكانية تحسينات إضافية من خلال الضبط الدقيق وتحسين المعلمات.
مناقشة
تسلط قسم المناقشة في الورقة الضوء على التقدم السريع في تطبيق التعلم العميق على التحليل البيولوجي الصوتي، مع التأكيد على أهمية تعلم التمثيل وبروتوكولات التقييم الموحدة. يستعرض المساهمات الهامة في هذا المجال، مثل النظرة الشاملة لـ Stowell حول البيولوجيا الصوتية الحاسوبية، التي تحدد التحديات مثل ندرة البيانات وضرورة وجود ممارسات تقييم قوية. لقد سهل إنشاء معايير مثل BEANS وBIRB وBirdSet المقارنات النظامية للنماذج، خاصةً في معالجة التحديات الخاصة بالمجال في نمذجة البيولوجيا الصوتية للطيور.
يتناول القسم أيضًا تعلم تمثيل البيولوجيا الصوتية، مشيرًا إلى أن الدراسات أظهرت أن التدريب المسبق الخاص بالمجال يعزز بشكل كبير من أداء النموذج في مهام البيولوجيا الصوتية. على سبيل المثال، أظهر Ghani وآخرون أن التضمينات من مصنفات أصوات الطيور المتخصصة تتفوق على نماذج الصوت العامة. كما تؤكد المناقشة على الدور الحاسم لبناء مجموعة البيانات وبروتوكولات التقسيم، كاشفة أن التقسيم غير الصحيح يمكن أن يؤدي إلى تسرب البيانات وتقديرات أداء متفائلة بشكل مفرط. تهدف المعايير الحديثة إلى توحيد هذه البروتوكولات لضمان تقييمات أكثر واقعية لتعميم النموذج. بشكل عام، يؤكد القسم على الحاجة إلى منهجيات تقييم صارمة والاختيار الدقيق لمجموعات بيانات التدريب لدفع مجال تصنيف البيولوجيا الصوتية إلى الأمام بشكل فعال.
DOI: https://doi.org/10.1016/j.ecoinf.2026.103765
Publication Date: 2026-04-27
Author(s): Raphael Schwinger et al.
Primary Topic: Animal Vocal Communication and Behavior
Overview
This section provides a comprehensive overview and analysis of automated bioacoustic analysis, emphasizing the necessity of advanced deep learning models for biodiversity monitoring and conservation. The article reviews large-scale pretrained bioacoustic foundation models, focusing on their transferability across various classification tasks. It details the training data, preprocessing, augmentations, architecture, and training paradigms of these models, alongside an empirical study evaluating their performance on the BEANS and BirdSet benchmarks. Key findings indicate that Perch 2.0 excels with the highest BirdSet score (90.78 AUROC) and the best linear probing result on BEANS (97.78 AUROC), demonstrating the effectiveness of diverse multi-taxa supervised pretraining. Additionally, BirdMAE is identified as the top model for probing strategies on BirdSet, while general-purpose audio models trained on AudioSet outperform specialized bird sound models on BEANS.
The conclusion synthesizes the findings from the comparative analysis of thirteen foundation models, highlighting the significance of attentive probing in maximizing the performance of transformer-based models. It suggests that Perch 2.0 is a strong default choice for bioacoustic classification due to its high performance and low parameter count, making it suitable for resource-constrained applications. The analysis also points to BEATs NLM for taxa-wide classification and BirdMAE for bird sound classification with attentive probing. Future research directions include the development of large-scale models trained on passive acoustic monitoring data and the exploration of advanced adaptation strategies for bioacoustic tasks.
Introduction
The introduction emphasizes the critical role of biodiversity monitoring in informing conservation strategies and understanding ecological dynamics. It highlights that such monitoring is vital for detecting changes in species populations and ecological processes, which in turn aids in effective management and protection measures. Specifically, monitoring bird populations serves as an indicator of broader biodiversity shifts.
Passive acoustic monitoring (PAM) is presented as a scalable and minimally invasive method for tracking sound-producing taxa, particularly those that are rare or difficult to survey due to their nocturnal habits or cryptic nature. PAM has been shown to detect an equivalent or greater number of species compared to traditional expert-led surveys, while also providing extensive temporal coverage. However, the challenge of annotating vast amounts of audio data has led to a focus on automated species detection methods. Deep learning (DL) models have emerged as powerful tools for this purpose, demonstrating high accuracy in classifying species by sound. Nonetheless, the effectiveness of DL approaches is contingent upon the availability of large annotated datasets, which poses a limitation for rare and understudied species.
Methods
In this section, the authors outline their experimental design, beginning with a description of the benchmark tasks chosen for evaluation. These tasks are critical for assessing the performance of the models under investigation. Following this, the authors detail the model selection process, which involves criteria for choosing the most suitable models for the tasks at hand. Finally, the training protocol for model adaptation is presented, highlighting the methodologies employed to fine-tune the models for optimal performance on the selected benchmarks. This structured approach ensures a comprehensive evaluation of the models’ capabilities.
Results
The results of the study highlight the performance of various bioacoustic models on the BEANS and BirdSet benchmarks, utilizing linear and attentive probing strategies. The AUROC scores indicate that models trained on general audio datasets, such as BEATs and Perch 2.0, outperform many bioacoustic-specific models. Notably, BEATs NLM achieved the highest score on BEANS (98.57 AUROC) with attentive probing, while Perch 2.0 excelled on BirdSet (90.78 AUROC) with restricted evaluation, demonstrating the effectiveness of diverse multi-taxa training. The findings suggest that sophisticated self-supervised learning (SSL) methods can yield robust representations applicable to bioacoustic tasks, as evidenced by the superior performance of models like BirdMAE (86.54 AUROC) compared to those trained solely on bioacoustic data.
The analysis also reveals that higher sampling rates (32 kHz) and optimized window lengths (5 seconds) significantly enhance model performance, particularly for bird classification tasks. In contrast, models utilizing raw waveforms consistently underperformed relative to those based on spectrograms. The study emphasizes that while dataset scale is important, the quality of data curation and the choice of training paradigms are critical for achieving optimal results. Recommendations for future model development include leveraging established datasets, exploring the combination of general and bioacoustic data, and considering computational efficiency in model selection. The study concludes that Perch 2.0 and BirdMAE are strong candidates for bioacoustic applications, with the potential for further improvements through fine-tuning and hyperparameter optimization.
Discussion
The discussion section of the paper highlights the rapid advancements in applying deep learning to bioacoustic analysis, emphasizing the importance of representation learning and standardized evaluation protocols. It reviews significant contributions in the field, such as Stowell’s comprehensive overview of computational bioacoustics, which identifies challenges like data scarcity and the necessity for robust evaluation practices. The establishment of benchmarks like BEANS, BIRB, and BirdSet has facilitated systematic model comparisons, particularly in addressing domain-specific challenges in avian bioacoustic modeling.
The section also delves into bioacoustic representation learning, noting that studies have shown domain-specific pretraining significantly enhances model performance in bioacoustic tasks. For instance, Ghani et al. demonstrated that embeddings from specialized bird sound classifiers outperform general audio models. The discussion further underscores the critical role of dataset construction and split protocols, revealing that improper splits can lead to data leakage and overly optimistic performance estimates. Recent benchmarks have aimed to standardize these protocols to ensure more realistic assessments of model generalization. Overall, the section emphasizes the need for rigorous evaluation methodologies and the careful selection of training datasets to advance the field of bioacoustic classification effectively.
