تحسين اختيار الميزات في بيانات الميكروأري السرطانية باستخدام إطار عمل تطوري مدفوع بالهيب لمساحات عالية الأبعاد
Optimizing feature selection in cancer microarray data using a heap-driven evolutionary framework for high-dimensional spaces

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-026-37803-5
PMID: https://pubmed.ncbi.nlm.nih.gov/41617939
تاريخ النشر: 2026-01-30
المؤلف: Mohammed Alweshah وآخرون
الموضوع الرئيسي: تعبير الجينات وتصنيف السرطان

نظرة عامة

اختيار الميزات (FS) هو خطوة معالجة مسبقة أساسية في مجالات مختلفة، وخاصة في تعدين البيانات الطبية، حيث يعزز كفاءة خوارزميات التعلم من خلال تقليل التكرار والخصائص غير ذات الصلة في مجموعات البيانات. تتطلب التعقيدات المتزايدة للبيانات تطوير تقنيات يمكن أن تقلل بفعالية من الأبعاد مع تحسين دقة التنبؤ. لقد فضلت التطورات الأخيرة طرق التحسين الميتاheuristic على الأساليب التقليدية بسبب قدرتها على تقديم حلول فعالة وفي الوقت المناسب لمشاكل FS. من بين هذه الطرق، يتم تسليط الضوء على المحسن القائم على الكومة (HBO) كنهج ميتاheuristic جديد يدمج هيستوجرام التدرجات الموجهة مع مصنف الجار الأقرب باستخدام طريقة التغليف. تم تقييم فعالية HBO مقابل سبع تقنيات FS معروفة عبر تسع مجموعات بيانات عالية الأبعاد، مما أظهر دقة متفوقة ومجموعات ميزات مخفضة في مجموعتين من البيانات، بالإضافة إلى معدلات تقارب أسرع.

قدمت الدراسة أيضًا نسخة محدثة من HBO، تُسمى المحسن القائم على الكومة الثنائية (BHBO)، والتي تم اختبارها على نفس مجموعات البيانات المرجعية. أشارت النتائج إلى أن HBO قلل بفعالية من عدد الميزات مع الحفاظ على أو تحسين دقة التصنيف مقارنة بأساليب أخرى، بما في ذلك BPSO وBGSA وBBA وBALO وGA وBHHO وBSSA. اختلف أداء الطريقة المقترحة اعتمادًا على سياق المشكلة المحدد، مما يبرز أهمية تحقيق التوازن بين الدقة وتقليل الأبعاد. بشكل عام، تشير النتائج إلى أن نهج HBO هو حل واعد لتحسين عملية FS في مجموعات البيانات عالية الأبعاد، مع إمكانية مزيد من التحسين والتطبيق في مجالات متنوعة.

مقدمة

تناقش المقدمة التحديات التي تطرحها البيانات عالية الأبعاد في تعلم الآلة، وخاصة في التطبيقات الطبية مثل مجموعات بيانات الميكروأري السرطانية. مع زيادة الأبعاد، يتوسع فضاء البحث بشكل أسي، مما يعقد المهام مثل التصنيف والانحدار. يتفاقم هذا التعقيد بسبب وجود ميزات غير ذات صلة أو مكررة، والتي يمكن أن تؤثر سلبًا على أداء المصنف وتزيد من التكاليف الحاسوبية. لمعالجة هذه القضايا، تقترح تقنيات اختيار الميزات (FS) تقليل عدد الميزات عن طريق القضاء على تلك التي لا تساهم بشكل ذي معنى في عملية التعلم. ومع ذلك، فإن البحث عن مجموعات الميزات المثلى هو NP-hard، مما يتطلب تطوير طرق FS أكثر كفاءة.

تقدم الورقة المحسن القائم على الكومة (HBO) كإطار ميتاheuristic مصمم لتعزيز FS في مجموعات البيانات عالية الأبعاد، مستهدفة بشكل خاص بيانات الميكروأري السرطانية. على الرغم من التقدم في تقنيات FS، غالبًا ما تفشل الأساليب الحالية في التعامل مع التعقيد الحاسوبي للتطبيقات الواقعية. يهدف HBO إلى تقديم حل من خلال تحقيق التوازن بين الاستكشاف والاستغلال في البحث عن الميزات ذات الصلة، مما يحسن دقة التنبؤ ويقلل من وقت الحساب. تسلط المقدمة الضوء على أهمية FS الفعالة في معالجة البيانات الطبية، وخاصة في تحديد العلامات الحيوية وتحسين الاستراتيجيات العلاجية في أبحاث السرطان.

النتائج

في هذا القسم، يقدم المؤلفون نتائج التجارب التي أجريت على تسع مجموعات بيانات ميكروأري سرطانية عالية الأبعاد تتميز بأحجام عينات صغيرة وعدد كبير من الفئات. تم تنفيذ التجارب باستخدام MATLAB R2016a على إعدادات الأجهزة القياسية، مع تكرار كل تجربة 30 مرة لتقليل التأثيرات العشوائية. أظهر نهج HBO المقترح (تحسين ثنائي هجين) أداءً متفوقًا مقارنة بسبع طرق معروفة، محققًا دقة متوسطة تبلغ 95.2%، ومعدل تقليل الميزات بنسبة 85.3%، وسرعة تقارب تبلغ 45 تكرارًا في المتوسط.

لتقديم تقييم شامل، وسع المؤلفون مقاييسهم لتشمل الدقة، والاسترجاع، ودرجة F1، والحساسية، والنوعية، وAUC-ROC، وجذر متوسط مربع الخطأ (RMSE)، ووقت الحساب. تم تقييم الأهمية الإحصائية باستخدام اختبار فريدمان، الذي أشار إلى تصنيف عام كبير لـ HBO (chi-square=48.2، p<0.001)، مع اختبارات Nemenyi اللاحقة التي كشفت عن اختلافات كبيرة بين HBO والأساليب ذات التصنيف الأدنى. أكدت اختبارات ويلكوكسون ذات الرتبة الموقعة أيضًا تفوق HBO على خوارزميات معينة، مثل PSO وACO، مع درجات Z تبلغ -4.62 و-4.31، على التوالي (كلاهما p<0.001). أشار اختبار الإشارة إلى أن HBO تفوق على الأساليب الأساسية في 28 من 30 تجربة (p<0.001)، مما يبرز قوته وفعاليته في سياق تصنيف السرطان.

المناقشة

في قسم المناقشة من ورقة البحث، يؤكد المؤلفون على أهمية تحقيق التوازن بين الاستكشاف والاستغلال في خوارزميات التحسين لتحقيق سرعات تقارب عالية، وخاصة في أطر اختيار الميزات (FS). يستعرضون خوارزميات ميتاheuristic المختلفة وتطبيقاتها في FS، مع تسليط الضوء على المساهمات الكبيرة من دراسات مثل دراسة Li وآخرون، الذين اقترحوا خوارزمية تطورية متعددة الأهداف (DMEA-FS) لاختيار الميزات على نطاق واسع. يشير المؤلفون إلى أن المحسن القائم على الكومة (HBO) يتفوق على الأساليب الحالية، مثل خوارزميات البحث الميميتي وخوارزميات البحث عن البيض، من حيث سرعة التقارب وقابلية التوسع لبيانات البيولوجيا عالية الأبعاد، وخاصة في تحليل الميكروأري.

يناقش القسم أيضًا دمج HBO مع تقنيات استخراج الميزات مثل هيستوجرام التدرجات الموجهة (HOG) ومصنفات مثل الجار الأقرب (KNN). تم تصميم خوارزمية HBO لتقليل دالة هدف مركبة توازن بين دقة التصنيف وملاءمة مجموعة الميزات، مما يضمن التعامل الفعال مع التحديات التي تطرحها مجموعات البيانات عالية الأبعاد. يجادل المؤلفون بأن الهيكل المستوحى من المجتمع لـ HBO ونمذجة البحث الهرمية تعزز أدائها في تحديد الميزات ذات الصلة المرتبطة بمسارات الأمراض، مما يوفر إطارًا قويًا لتحليل البيانات الجينومية في السياقات الطبية.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-026-37803-5
PMID: https://pubmed.ncbi.nlm.nih.gov/41617939
Publication Date: 2026-01-30
Author(s): Mohammed Alweshah et al.
Primary Topic: Gene expression and cancer classification

Overview

Feature selection (FS) is an essential preprocessing step in various fields, particularly in medical data mining, where it enhances the efficiency of learning algorithms by reducing redundancy and irrelevant characteristics in datasets. The increasing complexity of data necessitates the development of techniques that can effectively minimize dimensionality while optimizing prediction accuracy. Recent advancements have favored metaheuristic optimization methods over traditional approaches due to their ability to provide timely and effective solutions to FS problems. Among these, the Heap Based Optimizer (HBO) is highlighted as a novel metaheuristic approach that integrates the Histogram of Oriented Gradients with the K-Nearest Neighbor classifier using a wrapper method. The effectiveness of HBO was evaluated against seven established FS techniques across nine high-dimensional datasets, demonstrating superior accuracy and reduced feature sets in two datasets, along with faster convergence rates.

The study further introduced an upgraded version of HBO, termed Binary Heap Based Optimizer (BHBO), which was tested on the same benchmark datasets. Results indicated that HBO effectively reduced the number of features while maintaining or improving classification accuracy compared to other methods, including BPSO, BGSA, BBA, BALO, GA, BHHO, and BSSA. The performance of the proposed method varied depending on the specific problem context, emphasizing the importance of balancing accuracy and dimensionality reduction. Overall, the findings suggest that the HBO approach is a promising solution for optimizing the FS process in high-dimensional datasets, with potential for further refinement and application in diverse domains.

Introduction

The introduction discusses the challenges posed by high-dimensional data in machine learning, particularly in medical applications such as cancer microarray datasets. As dimensionality increases, the search space expands exponentially, complicating tasks like classification and regression. This complexity is exacerbated by the presence of irrelevant or redundant features, which can degrade classifier performance and increase computational costs. To address these issues, feature selection (FS) techniques are proposed to reduce the number of features by eliminating those that do not contribute meaningfully to the learning process. However, the search for optimal feature subsets is NP-hard, necessitating the development of more efficient FS methods.

The paper introduces the Heap-Based Optimizer (HBO) as a metaheuristic framework designed to enhance FS in high-dimensional datasets, specifically targeting cancer microarray data. Despite advancements in FS techniques, existing methods often fall short in handling the computational complexity of real-world applications. The HBO aims to provide a solution by balancing exploration and exploitation in the search for relevant features, thereby improving prediction accuracy and reducing computation time. The introduction highlights the importance of effective FS in medical data processing, particularly in identifying biomarkers and optimizing therapeutic strategies in cancer research.

Results

In this section, the authors present the results of experiments conducted on nine high-dimensional cancer microarray datasets characterized by small sample sizes and a large number of classes. The experiments were executed using MATLAB R2016a on a standard hardware setup, with each experiment repeated 30 times to mitigate stochastic effects. The proposed HBO (Hybrid Binary Optimization) approach demonstrated superior performance compared to seven established methods, achieving an average accuracy of 95.2%, a feature reduction rate of 85.3%, and a convergence speed of 45 iterations on average.

To provide a comprehensive evaluation, the authors expanded their metrics to include precision, recall, F1-score, sensitivity, specificity, AUC-ROC, root mean square error (RMSE), and computational time. Statistical significance was assessed using the Friedman test, which indicated a significant overall ranking for HBO (chi-square=48.2, p<0.001), with post-hoc Nemenyi tests revealing significant differences between HBO and lower-ranked methods. Wilcoxon signed-rank tests further confirmed HBO's superiority over specific algorithms, such as PSO and ACO, with Z-scores of -4.62 and -4.31, respectively (both p<0.001). The sign test indicated that HBO outperformed baseline methods in 28 out of 30 runs (p<0.001), underscoring its robustness and effectiveness in the context of cancer classification.

Discussion

In the discussion section of the research paper, the authors emphasize the importance of balancing exploration and exploitation in optimization algorithms to achieve high convergence speeds, particularly in feature selection (FS) frameworks. They review various metaheuristic algorithms and their applications in FS, highlighting significant contributions from studies such as Li et al., who proposed a multi-objective evolutionary algorithm (DMEA-FS) for large-scale feature selection. The authors note that their proposed Heap-Based Optimizer (HBO) outperforms existing methods, such as memetic and cuckoo search algorithms, in terms of convergence speed and scalability for high-dimensional biological data, specifically in microarray analysis.

The section also discusses the integration of HBO with feature extraction techniques like Histogram of Oriented Gradients (HOG) and classifiers such as k-Nearest Neighbors (KNN). The HBO algorithm is designed to minimize a composite objective function that balances classification accuracy and feature subset compactness, ensuring effective handling of the challenges posed by high-dimensional datasets. The authors argue that HBO’s socio-inspired structure and hierarchical modeling of search agents enhance its performance in identifying relevant features linked to disease pathways, thereby providing a robust framework for genomic data analysis in medical contexts.

شارك: