الكشف عن الأحداث النادرة من خلال التجميع التدريجي تحت العينة
Rare event detection by progressive clustering undersampling

شارك:
المجلة: PLoS ONE، المجلد: 21، العدد: 1
DOI: https://doi.org/10.1371/journal.pone.0340758
PMID: https://pubmed.ncbi.nlm.nih.gov/41616015
تاريخ النشر: 2026-01-30
المؤلف: Amr M. Abuzeid وآخرون
الموضوع الرئيسي: تقنيات تصنيف البيانات غير المتوازنة

نظرة عامة

تناقش هذه القسم مشكلة مجموعات البيانات غير المتوازنة في مهام التصنيف، حيث تحتوي فئة واحدة على عدد أقل بكثير من الحالات مقارنة بالفئات الأخرى. تشكل هذه اللاموازنة تحديًا، حيث أن المصنفات القياسية غالبًا ما تعطي الأولوية للفئات الأكبر، مما يؤدي إلى إهمال الفئات الأصغر، التي قد تكون أكثر إفادة. تسلط شيوع مجموعات البيانات غير المتوازنة في تطبيقات مختلفة – مثل كشف الاحتيال، والحقول الطبية الحيوية، والهندسة، والاستشعار عن بعد، والتصنيع – الضوء على الطبيعة الحرجة لهذه المشكلة في تعلم الآلة.

لمعالجة التحديات التي تطرحها مجموعات البيانات غير المتوازنة، تستعرض الورقة عدة حلول مقترحة. تهدف الدراسة إلى تقديم رؤى حول مشكلة مجموعة البيانات غير المتوازنة واستكشاف استراتيجيات فعالة للتخفيف من تأثيرها على أداء التصنيف.

مقدمة

تناقش المقدمة التحدي الكبير الذي تطرحه اللاموازنة في الفئات في تعلم الآلة والتعرف على الأنماط، خاصة في مجموعات البيانات ذات الفئتين حيث تكون فئة واحدة ممثلة بشكل ناقص. هذه اللاموازنة حرجة في التطبيقات الواقعية حيث يمكن أن يؤدي التصنيف الخاطئ للفئة الأقل إلى تكاليف عالية، مثل في كشف الاحتيال، وتشخيص الأمراض النادرة، وتصنيف النصوص. تحدد الورقة استراتيجيتين رئيسيتين لمعالجة هذه المشكلة: الأساليب الداخلية، التي تعدل الخوارزميات لأخذ اللاموازنة في الفئات بعين الاعتبار، والأساليب الخارجية، التي تعالج البيانات مسبقًا للتخفيف من آثار اللاموازنة. بينما تكون الأساليب الداخلية محددة بالخوارزميات، توفر الأساليب الخارجية مرونة أكبر عبر مصنفات مختلفة.

تسلط الورقة الضوء على تطبيق CO2 RBFN (شبكة الوظيفة الأساسية الشعاعية) لمعالجة مشاكل التصنيف غير المتوازن، مشددة على أن تصنيف الفئة الأقل بدقة غالبًا ما يكون أكثر أهمية من الفئة الأكبر. على سبيل المثال، في توقع تفاعلات البروتينات، يبرز انتشار البروتينات غير المتفاعلة هذه الحاجة. بالإضافة إلى ذلك، يشير المؤلفون إلى أنه بينما لا يُعتبر REMED حلاً نهائيًا لتعلم الآلة في التشخيص الطبي، فإنه يقدم نهجًا واعدًا من خلال تعزيز قابلية تفسير المعرفة التشخيصية، وتحسين اتخاذ القرار، وتقليل عدد الاختبارات المطلوبة للتشخيصات الموثوقة.

الطرق

في قسم الطرق، تناقش الأبحاث التعلم الحساس للتكاليف، الذي يبرز أهمية تقليل تكاليف التصنيف الخاطئ المرتبطة بالفئات المختلفة. يتم استخدام مصفوفة تكلفة لتوضيح عواقب تصنيف العينات بشكل خاطئ، حيث تكون تكلفة تصنيف العينات الأقل بشكل خاطئ أكبر من تلك الخاصة بالعينات الأكبر، المشار إليها كـ \( C(\text{Majority}, \text{Minority}) > C(\text{Minority}, \text{Majority}) \). الهدف هو تحسين التصنيف من خلال تعديل نسب العينات، كما هو موضح بواسطة تشارلز [52]، وتنفيذ استراتيجيات حساسة للتكاليف كما أوصى بها بيدرو [53].

تتناول القسم بمزيد من التفصيل طرقًا حساسة للتكاليف، بما في ذلك شبكة عصبية حساسة للتكاليف تستخدم تقنية تحريك العتبة لتقليل التصنيف الخاطئ للعينات ذات التكلفة العالية. بالإضافة إلى ذلك، تم تقديم ثلاث طرق تعزيز – AdaC1 وAdaC2 وAdaC3 – التي تدمج اعتبارات التكلفة في استراتيجيات التحديث الخاصة بها. تتضمن طريقة تجميع حساسة للتكاليف اقترحها يونغ زانغ ودابنغ وانغ دمج آلات الدعم الحساسة للتكاليف (SVM) مع نهج الاستفسار بواسطة اللجنة (QBC) لمعالجة تصنيف البيانات غير المتوازنة. تتضمن هذه الطريقة تقسيم الفئة الأكبر إلى مجموعات فرعية واستخدام AdaBoost للتدريب، تليها توليد عينات مرشحة من خلال QBC وتطبيق SVM حساسة للتكاليف للتعلم.

أخيرًا، يتناول القسم طريقة كسب المعلومات، المعروفة أيضًا باسم المعلومات المتبادلة المتوقعة، التي تقيس المعلومات المكتسبة لتوقع الفئة بناءً على وجود أو غياب المصطلحات في الوثائق. يتم استخدام تقنية تقدير الاحتمالية المتوقعة (ELE) لتخفيف الانفرادات في تقدير الاحتمالات [56].

المناقشة

في قسم المناقشة من ورقة البحث، يتناول المؤلفون التحديات التي تطرحها مجموعات البيانات غير المتوازنة في تصنيف النصوص وكشف تسرب النفط. يبرزون فعالية استراتيجيات أخذ العينات المختلفة، وخاصة أخذ العينات الزائدة وأخذ العينات الفرعية، في التخفيف من مشاكل اللاموازنة في الفئات. تستخدم الدراسة ثلاثة مصنفات – آلة الدعم (SVM)، الجيران الأقرب (KNN)، ونايف بايز – على نصوص طبية حيوية مأخوذة من قاعدة بيانات PubMed، باستخدام ثلاثة قواميس متميزة: NLPBA، BioCreative، ومجموعة فرعية مخصصة من قاعدة بيانات UniProt. من الجدير بالذكر أن مصنف SVM، عند اقترانه بطريقة أخذ العينات الفرعية وقواميس NLPBA والبروتين، حقق أفضل النتائج، متفوقًا على المعايير من مجموعة بيانات TREC Genomics 2005 العامة.

كما يؤكد المؤلفون على الحاجة إلى دراسات مقارنة شاملة حول الأساليب لمعالجة الفئات غير المتوازنة، حيث أن الأساليب الحالية مجزأة. يحددون القضايا الرئيسية مثل مشكلة التباينات الصغيرة، حيث تتجمع حالات الفئة الأقل في مجموعات صغيرة، وتأثير البيانات الضوضائية، التي تؤثر بشكل غير متناسب على الفئات الأقل. بالإضافة إلى ذلك، يناقشون مشكلة تحول مجموعة البيانات، حيث تختلف توزيعات بيانات التدريب والاختبار، مما يعقد مهام التصنيف. تحدد الورقة تقنيات مختلفة لمعالجة مجموعات البيانات غير المتوازنة، مصنفة إلى تقنيات على مستوى البيانات، وتقنيات على مستوى الخوارزمية، وتقنيات حساسة للتكاليف، واختيار الميزات، ونهج التجميع، مما يبرز ضرورة وجود استراتيجيات مخصصة لتحسين أداء التصنيف في ظل وجود اللاموازنة في الفئات.

Journal: PLoS ONE, Volume: 21, Issue: 1
DOI: https://doi.org/10.1371/journal.pone.0340758
PMID: https://pubmed.ncbi.nlm.nih.gov/41616015
Publication Date: 2026-01-30
Author(s): Amr M. Abuzeid et al.
Primary Topic: Imbalanced Data Classification Techniques

Overview

The section discusses the issue of imbalanced datasets in classification tasks, where one class has significantly fewer instances than others. This imbalance poses a challenge, as standard classifiers often prioritize larger classes, leading to the neglect of smaller, yet potentially more informative, classes. The prevalence of imbalanced datasets in various applications—such as fraud detection, biomedical fields, engineering, remote sensing, and manufacturing—highlights the critical nature of this problem in machine learning.

To address the challenges posed by imbalanced datasets, the paper reviews several proposed solutions. The study aims to provide insights into the imbalanced dataset problem and explore effective strategies for mitigating its impact on classification performance.

Introduction

The introduction discusses the significant challenge posed by class imbalance in Machine Learning and Pattern Recognition, particularly in two-class datasets where one class is under-represented. This imbalance is critical in real-world applications where misclassification of the minority class can incur high costs, such as in fraud detection, rare disease diagnosis, and text categorization. The paper outlines two main strategies to address this issue: internal approaches, which modify algorithms to account for class imbalance, and external approaches, which preprocess data to mitigate imbalance effects. While internal methods are algorithm-specific, external methods offer greater flexibility across different classifiers.

The paper highlights the application of CO2 RBFN (Radial Basis Function Network) for tackling imbalanced classification problems, emphasizing that accurately classifying the minority class is often more crucial than the majority class. For instance, in predicting protein-protein interactions, the prevalence of non-interacting proteins underscores this need. Additionally, the authors note that while REMED is not a definitive solution for machine learning in medical diagnostics, it presents a promising approach by enhancing the interpretability of diagnostic knowledge, improving decision-making, and reducing the number of tests required for reliable diagnoses.

Methods

In the section on methods, the research discusses cost-sensitive learning, which emphasizes the importance of minimizing misclassification costs associated with different classes. A cost matrix is employed to illustrate the consequences of misclassifying samples, where the cost of misclassifying minority samples is greater than that of majority samples, denoted as \( C(\text{Majority}, \text{Minority}) > C(\text{Minority}, \text{Majority}) \). The objective is to optimize classification by adjusting the sample proportions, as outlined by Charles [52], and to implement general cost-sensitive strategies as recommended by Pedro [53].

The section further details various cost-sensitive methods, including a cost-sensitive neural network that utilizes a threshold-moving technique to reduce misclassification of high-cost samples. Additionally, three boosting methods—AdaC1, AdaC2, and AdaC3—are introduced, which incorporate cost considerations into their updating strategies. A cost-sensitive ensemble method proposed by Yong Zhang and Dapeng Wang combines cost-sensitive support vector machines (SVM) with a query-by-committee (QBC) approach to address imbalanced data classification. This method involves partitioning the majority class into sub-datasets and employing AdaBoost for training, followed by generating candidate samples through QBC and applying cost-sensitive SVM for learning.

Lastly, the section touches on the Information Gain method, also known as Expected Mutual Information, which quantifies the information gained for category prediction based on the presence or absence of terms in documents. The Expected Likelihood Estimation (ELE) smoothing technique is utilized to mitigate singularities in probability estimation [56].

Discussion

In the discussion section of the research paper, the authors address the challenges posed by imbalanced datasets in text classification and oil spill detection. They highlight the effectiveness of various sampling strategies, specifically Oversampling and Subsampling, in mitigating class imbalance issues. The study employs three classifiers—Support Vector Machine (SVM), k-Nearest Neighbors (KNN), and Naive Bayes—on biomedical texts sourced from the PubMed database, utilizing three distinct dictionaries: NLPBA, BioCreative, and a custom subset from the UniProt database. Notably, the SVM classifier, when paired with the Subsampling method and the NLPBA and Protein dictionaries, yielded the best results, outperforming benchmarks from the TREC Genomics 2005 public corpus.

The authors also emphasize the need for comprehensive comparative studies on methods for addressing imbalanced classes, as existing approaches are fragmented. They identify key issues such as the small disjuncts problem, where minority class instances cluster in small groups, and the impact of noisy data, which disproportionately affects minority classes. Additionally, they discuss the dataset shift problem, where training and test data distributions differ, complicating classification tasks. The paper outlines various techniques to tackle imbalanced datasets, categorized into data-level, algorithmic-level, cost-sensitive, feature selection, and ensemble approaches, underscoring the necessity for tailored strategies to improve classification performance in the presence of class imbalance.

شارك: