تحسين وزغ الغطاس والشبكة العصبية متعددة الطبقات لتعزيز تصنيف سرطان الرئة
Greylag goose optimization and multilayer perceptron for enhancing lung cancer classification

شارك:
المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-72013-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39390014
تاريخ النشر: 2024-10-10
المؤلف: El‐Sayed M. El‐kenawy وآخرون
الموضوع الرئيسي: تعلم الآلة في المعلوماتية الحيوية

نظرة عامة

تتناول ورقة البحث القضية الحرجة لسرطان الرئة، الذي يتميز بنمو غير طبيعي للخلايا في أنسجة الرئة، مما يؤدي إلى مرضية ووفاة كبيرة. إن التعرف في الوقت المناسب والتصنيف الدقيق أمران أساسيان للعلاج الفعال والتشخيص. تستخدم الدراسة خوارزميات تحسين، وتحديداً تحسين الإوز الرمادي (GGO)، لتعزيز أداء نماذج التعلم الآلي في تصنيف سرطان الرئة من خلال تحسين اختيار الميزات من مجموعات البيانات المعقدة. أظهر GGO دقة تصنيف متفوقة، حيث حقق 98.4% باستخدام نموذج الشبكة العصبية متعددة الطبقات (MLP)، متفوقاً على خوارزميات تحسين ثنائية أخرى مثل bSC وbMVO وbPSO وbWOA وbGWO وbFOA. تم استخدام تحليلات إحصائية، بما في ذلك اختبار ويلكوكسون للرتب الموقعة وANOVA، للتحقق من النتائج، التي تم دعمها أيضاً من خلال تمثيلات رسومية.

في الختام، تحسن خوارزمية GGO بشكل كبير دقة تصنيف سرطان الرئة من خلال إعداد البيانات الفعال واختيار الميزات. حقق مصنف MLP، الذي تم تحسينه بواسطة GGO، دقة قدرها 98.4%، مما يشير إلى إمكانيته للتطبيق العملي في تشخيص سرطان الرئة. ستركز الأعمال المستقبلية على استخدام مجموعات بيانات أكبر واستكشاف نماذج تحسين وتعلم عميق متنوعة لتعزيز معدلات التنبؤ المبكر وتعميم النتائج عبر سياقات مختلفة. تسلط الدراسة الضوء على أهمية اختيار الميزات وتقليل الأبعاد في تحسين دقة التنبؤ ومعالجة الإفراط في التخصيص في تحليل السرطان.

الطرق

في هذا القسم، يوضح المؤلفون الطرق التجريبية المستخدمة لتقييم خوارزمية تحسين الإوز الرمادي (GGO) المقترحة مقابل ست تقنيات تحسين معروفة: SC وMVO وPSO وWOA وGWO وFOA. تم اختيار هذه الخوارزميات بسبب فعاليتها المعترف بها في أدبيات التحسين. شمل التقييم تطبيق وظائف رياضية تقليدية كمعايير لتحديد قيمها الدنيا ضمن منطقة بحث محددة، وهي ممارسة شائعة في أبحاث التحسين.

تم إجراء التجارب على منصة مزودة بذاكرة رئيسية سعة 16 جيجابايت، ومعالج Intel Core i7، وبطاقة رسومات GeForce RTX 2070 Super مع 8 جيجابايت من ذاكرة الوصول العشوائي. شملت بيئة البرمجيات Ubuntu 20.04 وTensorFlow 1.15 وCUDA 9.0 وcuDNN 7.1 وPython 3.7، باستخدام بيئة تطوير متكاملة (IDE) Spider. سهلت هذه الإعدادات مقارنة شاملة لإظهار الأداء المتفوق والفعالية لخوارزمية GGO مقارنة بالطرق المختارة للتحسين.

النتائج

في هذا القسم، تقدم الدراسة نتائج تجارب اختيار الميزات والتصنيف باستخدام خوارزميات تحسين متنوعة، بما في ذلك تحسين الإوز الرمادي الثنائي (bGGO) المقترح. تشير النتائج إلى أن bGGO تفوق على تقنيات اختيار الميزات الثنائية الأخرى، محققاً أدنى القيم لمتوسط الخطأ (0.7743) ومتوسط حجم الاختيار (0.7271) ومتوسط اللياقة (0.8375) وأفضل لياقة (0.7393) وأسوأ لياقة (0.8378) والانحراف المعياري للياقة (0.6598). تم توضيح فعالية bGGO بشكل أكبر من خلال تصورات متنوعة، بما في ذلك مخططات QQ وخرائط الحرارة، التي تؤكد موثوقيته في تحديد ميزات سرطان الرئة.

بالإضافة إلى ذلك، تم تقييم أداء التصنيف لعدة نماذج تعلم آلي، تم تعزيزها بتقنية bGGO. أظهر نموذج الشبكة العصبية متعددة الطبقات (MLP) أعلى مقاييس الأداء، بما في ذلك الدقة (0.9838) والحساسية (0.9773) والمساحة تحت منحنى ROC (AUC) (0.993). أكدت المقارنات مع تقنيات تحسين أخرى، مثل محسن الذئب الرمادي (GWO) ومحسن سرب الجسيمات (PSO)، تفوق نهج GGO + MLP. قدمت التحليلات الإحصائية، بما في ذلك KDE ومخططات الصندوق، رؤى حول توزيع دقة النماذج وأبرزت أهمية اختيار الميزات في تحسين نتائج التصنيف. بشكل عام، تؤكد النتائج فعالية تقنية bGGO المقترحة في تعزيز كل من اختيار الميزات ودقة التصنيف في سياق اكتشاف سرطان الرئة.

المناقشة

في قسم المناقشة من ورقة البحث، يتم تسليط الضوء على مختلف التقدمات في تشخيص سرطان الرئة الآلي باستخدام تقنيات التعلم الآلي (ML) والتعلم العميق (DL). بشكل ملحوظ، طور محمد وآخرون شبكة عصبية تلافيفية هجينة (CNN) مدمجة مع خوارزمية بحث تحسين الإيبولا (EOSA)، محققين دقة تصنيف قدرها 0.9321 على مجموعة بيانات سرطان الرئة. تفوق نموذجهم على approaches الأخرى من حيث الخصوصية والحساسية عبر الحالات الطبيعية والخبيثة والحميدة. بالمثل، قدم رين وآخرون إطار عمل جديد، LCGANT، الذي استخدم شبكة تلافيفية مولدة تنافسية عميقة (GAN) ونموذج VGG-DF، محققاً دقة مثيرة للإعجاب قدرها 99.84% ± 0.156%. استخدمت دراسات أخرى، مثل تلك التي أجراها أنانيا بهاتاشارجي وآخرون وVijh وآخرون، مصنفات وتقنيات تحسين متنوعة، مسجلة معدلات دقة عالية (تصل إلى 95% و98.08%، على التوالي) في التمييز بين العقيدات الرئوية الخبيثة وغير الخبيثة.

يناقش القسم أيضاً خوارزمية تحسين الإوز الرمادي (GGO)، التي تعزز اختيار الميزات لنماذج الشبكة العصبية متعددة الطبقات (MLP). تصنف خوارزمية GGO الأفراد ديناميكياً إلى مجموعات استكشاف واستغلال لتحسين الحلول بشكل فعال. تستخدم مبادئ رياضية مثل عدم المساواة المثلثية وقانون الأعداد الكبيرة لتحسين قدرات التقارب والاستكشاف. أظهر نهج GGO + MLP أداءً متفوقاً في تصنيف حالات سرطان الرئة مقارنة بخوارزميات التحسين الأخرى، كما يتضح من مخططات الدقة والاختبارات الإحصائية، مما يؤكد فعاليتها في تعزيز دقة التشخيص في نماذج التنبؤ بسرطان الرئة.

القيود

تظهر الدراسة حول خوارزمية GGO لتحسين تصنيف سرطان الرئة نتائج واعدة، خاصة في اختيار الميزات ودقة التصنيف. ومع ذلك، هناك عدة قيود تستدعي الاعتبار للبحوث المستقبلية. أولاً، تم تقييم أداء الخوارزمية باستخدام مجموعة بيانات واحدة، والتي قد لا تعكس تنوع السيناريوهات الواقعية بشكل كافٍ. وبالتالي، قد تتعرض عمومية النتائج للخطر، حيث قد تختلف النتائج عبر مجموعات بيانات مختلفة.

ثانياً، تتأثر دقة التصنيف بشكل كبير بجودة واكتمال بيانات الإدخال. يمكن أن تؤثر التناقضات أو القيم المفقودة داخل مجموعة البيانات بشكل كبير على النتائج، مما يشير إلى أن الأعمال المستقبلية يجب أن تركز على استخدام مجموعات بيانات أكثر شمولاً ومعالجة قضايا جودة البيانات لتعزيز قوة خوارزمية GGO في تصنيف سرطان الرئة.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-72013-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39390014
Publication Date: 2024-10-10
Author(s): El‐Sayed M. El‐kenawy et al.
Primary Topic: Machine Learning in Bioinformatics

Overview

The research paper addresses the critical issue of lung cancer, characterized by abnormal cell growth in lung tissues, which leads to significant morbidity and mortality. Timely identification and accurate staging are essential for effective treatment and prognosis. The study employs optimization algorithms, specifically the Greylag Goose Optimization (GGO), to enhance the performance of machine learning models in lung cancer classification by refining feature selection from complex datasets. The GGO demonstrated superior classification accuracy, achieving 98.4% with a Multilayer Perceptron (MLP) model, outperforming other binary optimization algorithms such as bSC, bMVO, bPSO, bWOA, bGWO, and bFOA. Statistical analyses, including the Wilcoxon signed-rank test and ANOVA, were utilized to validate the results, which were further supported by graphical representations.

In conclusion, the GGO algorithm significantly improves lung cancer classification accuracy through effective data preparation and feature selection. The MLP classifier, optimized by GGO, achieved an accuracy of 98.4%, indicating its potential for practical application in lung cancer diagnosis. Future work will focus on utilizing larger datasets and exploring various optimization and deep learning models to enhance early prediction rates and generalize findings across different contexts. The study highlights the importance of feature selection and dimensionality reduction in improving prediction accuracy and addressing overfitting in cancer analysis.

Methods

In this section, the authors detail the experimental methods used to evaluate the proposed Greylag Goose Optimization (GGO) algorithm against six established optimization techniques: SC, MVO, PSO, WOA, GWO, and FOA. These algorithms were selected due to their recognized effectiveness in the optimization literature. The assessment involved applying traditional mathematical functions as benchmarks to identify their minimum values within a specified search area, a common practice in optimization research.

The experiments were conducted on a platform equipped with 16 GB of primary memory, an Intel Core i7 CPU, and a GeForce RTX 2070 Super GPU with 8 GB of RAM. The software environment included Ubuntu 20.04, TensorFlow 1.15, CUDA 9.0, cuDNN 7.1, and Python 3.7, utilizing the Spider Integrated Development Environment (IDE). This setup facilitated a comprehensive comparison to demonstrate the superior performance and effectiveness of the GGO algorithm relative to the selected optimization methods.

Results

In this section, the study presents the results of feature selection and classification experiments using various optimization algorithms, including the proposed binary Greylag Goose Optimization (bGGO). The results indicate that bGGO outperformed other binary feature selection techniques, achieving the lowest values for average error (0.7743), average select size (0.7271), average fitness (0.8375), best fitness (0.7393), worst fitness (0.8378), and standard deviation fitness (0.6598). The effectiveness of bGGO is further illustrated through various visualizations, including QQ plots and heatmaps, which confirm its reliability in identifying lung cancer features.

Additionally, the classification performance of several machine learning models, enhanced by the bGGO technique, was evaluated. The Multilayer Perceptron (MLP) model demonstrated the highest performance metrics, including accuracy (0.9838), sensitivity (0.9773), and Area Under the ROC Curve (AUC) (0.993). Comparisons with other optimization techniques, such as the Gray Wolf Optimizer (GWO) and Particle Swarm Optimizer (PSO), reaffirmed the superiority of the GGO + MLP approach. Statistical analyses, including KDE and Box Plots, provided insights into the distribution of model accuracies and highlighted the importance of feature selection in improving classification outcomes. Overall, the findings underscore the effectiveness of the proposed bGGO technique in enhancing both feature selection and classification accuracy in the context of lung cancer detection.

Discussion

In the discussion section of the research paper, various advancements in automated lung cancer diagnosis using machine learning (ML) and deep learning (DL) techniques are highlighted. Notably, Mohamed et al. developed a hybrid convolutional neural network (CNN) combined with the Ebola optimization search algorithm (EOSA), achieving a classification accuracy of 0.9321 on a lung cancer dataset. Their model outperformed other approaches in terms of specificity and sensitivity across normal, malignant, and benign cases. Similarly, Ren et al. introduced a novel framework, LCGANT, which utilized a deep convolutional generative adversarial network (GAN) and a VGG-DF model, achieving an impressive accuracy of 99.84% ± 0.156%. Other studies, such as those by Ananya Bhattacharjee et al. and Vijh et al., employed various classifiers and optimization techniques, reporting high accuracy rates (up to 95% and 98.08%, respectively) in distinguishing between malignant and non-malignant lung nodules.

The section also discusses the Greylag Goose Optimization (GGO) algorithm, which enhances feature selection for multilayer perceptron (MLP) models. The GGO algorithm dynamically categorizes individuals into exploration and exploitation groups to optimize solutions effectively. It employs mathematical principles such as the Triangle Inequality and the Law of Large Numbers to improve convergence and exploration capabilities. The proposed GGO + MLP approach demonstrated superior performance in classifying lung cancer cases compared to other optimization algorithms, as evidenced by accuracy plots and statistical tests, confirming its efficacy in enhancing diagnostic accuracy in lung cancer prediction models.

Limitations

The study on the GGO algorithm for enhancing lung cancer classification demonstrates promising outcomes, particularly in feature selection and classification accuracy. However, several limitations warrant consideration for future research. Firstly, the algorithm’s performance was evaluated using a single dataset, which may not adequately reflect the diversity of real-world scenarios. Consequently, the generalizability of the findings could be compromised, as results may differ across various datasets.

Secondly, the accuracy of classification is heavily influenced by the quality and completeness of the input data. Inconsistencies or missing values within the dataset can significantly impact the results, suggesting that future work should focus on utilizing more comprehensive datasets and addressing data quality issues to enhance the robustness of the GGO algorithm in lung cancer classification.

شارك: