DOI: https://doi.org/10.3389/frai.2026.1790249
PMID: https://pubmed.ncbi.nlm.nih.gov/41940101
تاريخ النشر: 2026-03-19
المؤلف: Waqas Haider Bangyal وآخرون
الموضوع الرئيسي: تعبير الجينات وتصنيف السرطان
نظرة عامة
تتناول الدراسة الحاجة الملحة لتصنيف أنواع السرطان بشكل فعال لتعزيز الكشف المبكر والعلاج، مستفيدة من بيانات RNA-Seq لتحليل الجينات المعبر عنها بشكل مختلف. غالبًا ما تركز التحليلات التقليدية على أنواع السرطان الفردية كمشاكل تصنيف ثنائية، متجاهلة الطيف الأوسع من أنواع السرطان. تستخدم هذه البحث مجموعة بيانات من مستودع ماندلي، محولة قيم RNA-Seq إلى صور ثنائية الأبعاد للتصنيف عبر خمسة خوارزميات تعلم نقل (TL): VGG16، VGG19، ResNet50، ResNet101، وResNet152. تم تقييم المصنفات باستخدام استراتيجيات تقسيم بيانات متنوعة، مما كشف أن تقسيم 70-30 حقق أفضل أداء، حيث حققت VGG16 أعلى دقة بنسبة 95%.
تشير النتائج إلى أن التعلم بالنقل يمثل تقدمًا كبيرًا في الذكاء الاصطناعي، متجاوزًا الطرق التقليدية للتصنيف وتقليل الأبعاد. تظهر النتائج إمكانيات بيانات RNA-Seq في تصنيف أنواع السرطان المتعددة، مع ظهور VGG16 كنموذج الأكثر فعالية. يمكن أن تركز الأعمال المستقبلية على دمج عمليات استخراج الميزات واختيارها لتعزيز تحديد علامات المرض من خلال مزيد من التحقيق في الجينات الواعدة.
النتائج
في قسم النتائج، يقدم المؤلفون توصيفًا شاملاً لمختلف أنواع السرطان، مع تسليط الضوء على النتائج الرئيسية من تحليلاتهم. تشير النتائج إلى وجود ملفات تعريف جزيئية وجينية متميزة عبر أنواع السرطان المختلفة، والتي قد يكون لها آثار على العلاجات المستهدفة والطب الشخصي. تؤكد المناقشة على أهمية هذه التوصيفات في فهم سلوك الورم واستجابات العلاج، مما يشير إلى أن الأساليب المخصصة يمكن أن تعزز النتائج السريرية للمرضى الذين يعانون من أنواع سرطان محددة. بشكل عام، تؤكد النتائج على ضرورة استمرار البحث في تباين السرطان لتحسين استراتيجيات التشخيص والعلاج.
المناقشة
في هذه الدراسة، تم استخدام مجموعة بيانات تعبير الجينات التي تشمل خمسة أنواع من السرطان – سرطان الثدي الغازي (BRCA)، سرطان الكلى الخلوي الواضح (KIRC)، سرطان الرحم (UCEC)، سرطان الرئة الغدي (LUAD)، وسرطان الرئة الخلوي الحرشفي (LUSC) – لتقييم فعالية نماذج التعلم بالنقل (TL) المختلفة لتصنيف السرطان. تتكون مجموعة البيانات من 2,086 عينة مع 971 ميزة، والتي تم تحويلها إلى صور ثنائية الأبعاد من خلال مرحلة معالجة مسبقة تضمنت التطبيع وإضافة الصفر. استخدمت الدراسة خمسة نماذج TL: VGG16، VGG19، ResNet50، ResNet101، وResNet152، مع تقييم أدائها تحت تقسيمات تدريب-اختبار مختلفة (50-50، 60-40، 70-30، و80-20) مع وبدون زيادة البيانات.
أشارت النتائج إلى أن VGG16 حققت أعلى دقة بنسبة 95% لتقسيم 70-30 بدون زيادة، بينما حققت ResNet101 وResNet152 أفضل أداء مع زيادة البيانات، حيث وصلت دقتهما إلى 92% و93%، على التوالي. من الجدير بالذكر أن النماذج واجهت صعوبات في التصنيفات الخاطئة بين أنواع السرطان السريرية المتشابهة، مما يبرز الحاجة إلى تحسين طرق استخراج الميزات واختيارها لتحسين دقة التصنيف. كشفت التحليلات الإحصائية عن اختلافات كبيرة في الأداء بين النماذج، حيث تفوقت VGG16 وVGG19 على نماذج ResNet. تشير النتائج إلى أن النهج المقترح للتعلم بالنقل فعال في تصنيف السرطان، مع إمكانية العمل المستقبلي الموجه نحو دمج عمليات استخراج الميزات واختيارها لتحديد علامات المرض.
DOI: https://doi.org/10.3389/frai.2026.1790249
PMID: https://pubmed.ncbi.nlm.nih.gov/41940101
Publication Date: 2026-03-19
Author(s): Waqas Haider Bangyal et al.
Primary Topic: Gene expression and cancer classification
Overview
The study addresses the critical need for effective cancer type classification to enhance early detection and treatment, leveraging RNA-Seq data to analyze differentially expressed genes. Traditional analyses often focused on single cancer types as binary classification problems, neglecting the broader spectrum of cancer types. This research utilizes a dataset from the Mendeley repository, transforming RNA-Seq values into 2D images for classification via five Transfer Learning (TL) algorithms: VGG16, VGG19, ResNet50, ResNet101, and ResNet152. The classifiers were evaluated using various data splitting strategies, revealing that a 70-30 split yielded the best performance, with VGG16 achieving the highest accuracy of 95%.
The findings indicate that transfer learning represents a significant advancement in artificial intelligence, surpassing traditional methods for classification and dimensionality reduction. The results demonstrate the potential of RNA-Seq data in classifying multiple cancer types, with VGG16 emerging as the most effective model. Future work could focus on integrating feature extraction and selection processes to enhance the identification of disease biomarkers through further investigation of promising genes.
Results
In the Results section, the authors present a comprehensive characterization of various cancer types, highlighting key findings from their analyses. The results indicate distinct molecular and genetic profiles across different cancers, which may have implications for targeted therapies and personalized medicine. The discussion emphasizes the importance of these characterizations in understanding tumor behavior and treatment responses, suggesting that tailored approaches could enhance clinical outcomes for patients with specific cancer types. Overall, the findings underscore the necessity for continued research into the heterogeneity of cancer to improve diagnostic and therapeutic strategies.
Discussion
In this study, a gene expression dataset encompassing five cancer types—Breast Invasive Carcinoma (BRCA), Kidney Renal Clear Cell Carcinoma (KIRC), Uterine Corpus Endometrial Carcinoma (UCEC), Lung Adenocarcinoma (LUAD), and Lung Squamous Cell Carcinoma (LUSC)—was utilized to assess the efficacy of various transfer learning (TL) models for cancer classification. The dataset comprised 2,086 samples with 971 features, which were transformed into 2D images through a preprocessing phase that included normalization and zero-padding. The study employed five TL models: VGG16, VGG19, ResNet50, ResNet101, and ResNet152, evaluating their performance under different training-testing splits (50-50, 60-40, 70-30, and 80-20) with and without data augmentation.
The results indicated that VGG16 achieved the highest accuracy of 95% for the 70-30 split without augmentation, while ResNet101 and ResNet152 performed best with augmentation, reaching accuracies of 92% and 93%, respectively. Notably, the models struggled with misclassifications among clinically similar cancer types, highlighting the need for enhanced feature extraction and selection methods to improve classification accuracy. Statistical analysis revealed significant performance differences among the models, with VGG16 and VGG19 outperforming ResNet models. The findings suggest that the proposed TL approach is effective for cancer classification, with potential future work aimed at integrating feature extraction and selection processes to identify disease biomarkers.
