DOI: https://doi.org/10.1109/access.2026.3678764
تاريخ النشر: 2026-01-01
المؤلف: Hussain Ahmad Madni وآخرون
الموضوع الرئيسي: تعلم الآلة في الرعاية الصحية
نظرة عامة
تتناول هذه الورقة البحثية التحديات المتعلقة بمجموعات البيانات السريرية المحدودة واللوائح الصارمة للخصوصية في تطوير نماذج تصنيف دقيقة لتشخيص السرطان، مع التركيز بشكل خاص على سرطان القولون والرئة. يقترح المؤلفون إطار عمل لتوليد البيانات الاصطناعية يستخدم نموذج انتشار مستقر معزز بأوزان التكيف منخفض الرتبة (LoRA) لإنشاء صور اصطناعية عالية الدقة من عدد قليل من العينات الحقيقية. يحسن هذا النهج بشكل كبير أداء نماذج التعلم العميق، مثل محول الصور الفعال من حيث البيانات – كبير (DeiT-L) والتدريب المسبق اللغوي-الصوري التبايني (CLIP)، خاصة في سيناريوهات البيانات المنخفضة. تشير النتائج إلى أن البيانات الاصطناعية يمكن أن تعزز بشكل فعال مجموعات البيانات الحقيقية المحدودة، مما يعزز تعميم النموذج ودقة التصنيف.
تظهر الورقة أيضًا أن تنوع وكمية عينات القليل من اللقطات المستخدمة لتوجيه توليد الصور الاصطناعية أمر حاسم لإنتاج مجموعات بيانات واقعية تشبه بشكل وثيق توزيعات البيانات السريرية الحقيقية. تؤكد التقييمات النوعية والكمية، بما في ذلك تحليل مسافة فريشيت (FID)، أن زيادة التنوع في عينات القليل من اللقطات تؤدي إلى تحسين جودة الصور الاصطناعية. ومن الجدير بالذكر أن النماذج المدربة حصريًا على البيانات الاصطناعية حققت أداءً تنافسيًا مقارنة بتلك التي تم ضبطها على مجموعات البيانات الحقيقية. يعزز دمج التمثيلات متعددة الوسائط من خلال ترميز النص الموجه في إطار عمل قائم على CLIP نتائج التصنيف. بشكل عام، يظهر النهج الهجين المقترح، الذي يجمع بين عدد محدود من العينات الحقيقية والبيانات الاصطناعية، كاستراتيجية عملية وفعالة لنشر نماذج التعلم العميق في البيئات السريرية التي تعاني من نقص البيانات، مما يؤدي إلى تحسينات متسقة في مقاييس التصنيف مثل AUROC وF1-score.
مقدمة
في مقدمة هذه الورقة البحثية، يبرز المؤلفون الدور الحاسم للأطباء الشرعيين في التقييم النسيجي للسرطان، خاصة في التمييز بين الأنسجة الخبيثة وغير الخبيثة وتصنيف أنواع السرطان بدقة مثل سرطان القولون والرئة. تتطلب تعقيدات التقييمات النسيجية، التي تفاقمت بسبب تباين الأورام والعيوب التقنية، تطوير أدوات تشخيصية مكملة لتعزيز القابلية للتكرار وتقليل تباين المراقب. يشير المؤلفون إلى أنه على الرغم من أن تقنيات التعلم العميق قد أظهرت وعدًا في أتمتة تصنيف الأنسجة واكتشاف السرطان، فإن فعالية هذه النماذج تعتمد بشكل كبير على توفر مجموعات بيانات تدريب كبيرة ومتنوعة وعالية الجودة، والتي غالبًا ما تكون محدودة بسبب التحديات في توضيح البيانات، واللوائح المتعلقة بالخصوصية، والحواجز المؤسسية.
لمعالجة هذه القيود، تقترح الورقة استخدام توليد البيانات الاصطناعية، خاصة من خلال نماذج الانتشار، التي يمكن أن تخلق صورًا عالية الدقة تحاكي شكل الأنسجة الحقيقية مع الحفاظ على خصوصية المرضى. يؤكد المؤلفون على إمكانيات التعلم من القليل من اللقطات والنماذج التوليدية لإنتاج مجموعات بيانات اصطناعية واسعة من مدخلات قليلة، مما يعزز تدريب النموذج. يقدمون إطار عمل موحد يدمج توليد بيانات القليل من اللقطات المعتمد على الانتشار مع طرق التكيف الفعالة من حيث المعلمات، مثل التكيف منخفض الرتبة (LoRA)، ويقيمون بشكل منهجي تأثيرها المشترك عبر هياكل المصنفات المختلفة. يهدف هذا النهج إلى تحسين أداء التصنيف لأنواع سرطان القولون والرئة مع معالجة التحديات التي تطرحها ندرة البيانات في علم الأمراض النسيجي. تشمل المساهمات الرئيسية توليد صور اصطناعية عالية الدقة، وتطوير نهج جديد لتوليد البيانات، وتقييم شامل للمنهجية المقترحة على مجموعات بيانات مرجعية.
الطرق
تنقسم المنهجية الموضحة في هذه الورقة البحثية إلى قسمين رئيسيين: توليد البيانات والتصنيف. في مرحلة توليد البيانات، يتم إنشاء صور اصطناعية عالية الدقة لأنواع سرطان القولون والرئة باستخدام نهج قائم على الانتشار، مع استخدام نموذج الانتشار المستقر 2.1. يقوم هذا النموذج، المدرب مسبقًا على مجموعة ضخمة من الصور، بإنشاء مرئيات من ضوضاء عشوائية موجهة بواسطة مطالبات نصية. تقلل عملية التدريب من خسارة إزالة الضوضاء، \(L_D\)، لتحسين جودة الصورة، على الرغم من أنها تواجه تحديات في تمثيل حالات السرطان المحددة بدقة. لمعالجة هذه القيود، يتم استخدام طريقة DataDream، التي تتضمن التكيف منخفض الرتبة (LoRA) لضبط النموذج بكفاءة مع عدد محدود من العينات النسيجية، مما يجعلها مناسبة للسيناريوهات التي تعاني من نقص البيانات.
تُعزز عملية توليد الصور الاصطناعية من خلال التصميم التكراري لمطالبات محددة للفئة تصف الخصائص المرئية الحرجة اللازمة للتمثيل الدقيق. يتم صياغة هذه المطالبات بعناية لتعظيم الدقة الشكلية للصور المولدة، مع التركيز على عينات عالية الجودة من مجموعة بيانات LC25000. تُطبق تقنية LoRA على طبقات الإسقاط الخطي لآليات انتباه النموذج، مما يحسن فقط معلمات LoRA مع الحفاظ على أوزان النموذج الأصلية مجمدة. يتم إجراء الضبط الدقيق باستخدام مُحسِّن AdamW، مع استخدام الإيقاف المبكر لمنع الإفراط في التكيف. ستقوم النتائج التجريبية اللاحقة بتقييم كل من جودة البيانات الاصطناعية وأداء نماذج التصنيف المدربة على هذه المجموعات، وبالتالي تقييم فعالية المنهجية المقترحة.
المناقشة
في هذا القسم، يناقش المؤلفون التجارب التصنيفية التي أجريت باستخدام نماذج مختلفة قائمة على الشبكات العصبية الالتفافية والتحويلية المدربة على كل من مجموعات البيانات الحقيقية والاصطناعية. كان الهدف الرئيسي هو تقييم ما إذا كانت البيانات الاصطناعية، التي تم إنشاؤها من عدد محدود من الصور الحقيقية، يمكن أن تعزز أداء النموذج في اكتشاف الخصائص الشكلية في الصور النسيجية. تم ضبط الشبكات العصبية الالتفافية (CNNs) لتصنيف ثنائي ومتعدد الفئات، مستفيدة من قدرتها على استخراج الميزات الهرمية، بينما استخدمت هياكل التحويل آليات الانتباه الذاتي لالتقاط السياق العالمي، وهو ما يكون مفيدًا بشكل خاص في تمييز الميزات الدقيقة في الأمراض النادرة.
استخدمت الدراسة نماذج CLIP (التدريب المسبق اللغوي-الصوري التبايني) لتسهيل التعلم عبر الأنماط، مما يسمح بمقارنة المدخلات المرئية والمطالبات النصية في مساحة تضمين مشتركة. تم استخدام نوعين من CLIP، ViT-B/16 وViT-B/32، لتحليل التوازن بين دقة التمثيل وكفاءة الحساب. تم تدريب النماذج باستخدام دالة خسارة الانتروبيا المتقاطعة ومُحسِّن AdamW، مع بروتوكول تدريب منظم لضمان مقارنات عادلة بين النماذج المدربة على البيانات الحقيقية والاصطناعية. تشير النتائج إلى أن استراتيجية توليد البيانات الاصطناعية المقترحة تعزز بشكل فعال دقة التصنيف وقدرات التعميم في ظل ظروف نقص البيانات، مما يظهر إمكانية تطبيقها في البيئات السريرية حيث تكون البيانات الطبية المعلّمة محدودة.
القيود
تقدم الدراسة عدة قيود يجب أخذها في الاعتبار على الرغم من نتائجها الواعدة. أولاً، على الرغم من أن البيانات الاصطناعية عززت أداء التصنيف في السيناريوهات التي تعاني من نقص البيانات، قد لا تمثل الصور المولدة بشكل كافٍ التنوع البيولوجي والشكلي الموجود في البيئات السريرية الحقيقية، والتي يمكن أن تختلف بشكل كبير بسبب عوامل مثل الممارسات المؤسسية، وبروتوكولات الصبغ، وتقنيات المسح. ثانيًا، يعتمد تقييم جودة الصورة الاصطناعية، الذي يستند بشكل أساسي إلى مسافة فريشيت (FID)، على ميزات من نماذج مدربة على صور طبيعية. على الرغم من أن FID يقدم مقارنة نسبية عبر الظروف التجريبية، إلا أنه قد لا يلتقط تمامًا الميزات النسيجية ذات الصلة سريريًا.
بالإضافة إلى ذلك، تقتصر الأبحاث على التركيز على نوعين فقط من السرطان ومجموعة بيانات عامة واحدة، مما يستلزم مزيدًا من التحقق عبر مجموعات بيانات متعددة المراكز ونطاق أوسع من الأمراض لتعزيز القابلية السريرية. قد يحد الاعتماد على المطالبات المصممة من قبل الخبراء والاختيار اليدوي لعينات القليل من اللقطات عالية الجودة أيضًا من قابلية توسيع النهج. ستسعى الأبحاث المستقبلية إلى أتمتة تحسين المطالبات، وتوسيع تغطية الأمراض، وإشراك الخبراء السريريين في تقييم واقعية الصور الاصطناعية. علاوة على ذلك، بينما يهدف الأسلوب إلى تقليل مشاركة الصور الخام للمرضى، فإنه لا يوفر ضمانات خصوصية رسمية، حيث يتم ضبط نموذج الانتشار باستخدام عينات حقيقية دون تدابير الخصوصية التفاضلية. ستتناول التحقيقات المستقبلية هذه المخاوف المتعلقة بالخصوصية من خلال دمج آليات الحفاظ على الخصوصية الرسمية وإجراء اختبارات القوة المعاكسة لتقييم المخاطر المحتملة لتسرب البيانات.
DOI: https://doi.org/10.1109/access.2026.3678764
Publication Date: 2026-01-01
Author(s): Hussain Ahmad Madni et al.
Primary Topic: Machine Learning in Healthcare
Overview
This research paper addresses the challenges of limited clinical datasets and strict privacy regulations in developing accurate classification models for cancer diagnostics, specifically focusing on colorectal and lung carcinoma. The authors propose a synthetic data generation framework that utilizes a stable diffusion model enhanced with Low-Rank Adaptation (LoRA) weights to create high-fidelity synthetic images from a small number of real samples. This approach significantly improves the performance of deep learning models, such as the Data-efficient Image Transformer-Large (DeiT-L) and Contrastive Language-Image Pretraining (CLIP), particularly in low-data scenarios. The findings indicate that synthetic data can effectively augment limited real datasets, enhancing model generalization and classification accuracy.
The paper further demonstrates that the diversity and quantity of few-shot samples used for guiding synthetic image generation are crucial for producing realistic datasets that closely resemble real clinical data distributions. Qualitative and quantitative evaluations, including Fréchet Inception Distance (FID) analysis, confirm that increased diversity in few-shot samples leads to better synthetic image quality. Notably, models trained exclusively on synthetic data achieved competitive performance compared to those fine-tuned on real datasets. The integration of multimodal representations through prompt-guided text encoding in a CLIP-based framework further enhances classification outcomes. Overall, the proposed hybrid approach, which combines a limited number of real samples with synthetic data, emerges as a practical and effective strategy for deploying deep learning models in data-scarce clinical environments, yielding consistent improvements in classification metrics such as AUROC and F1-score.
Introduction
In the introduction of this research paper, the authors highlight the critical role of pathologists in the histopathological evaluation of cancer, particularly in distinguishing between malignant and non-malignant tissues and accurately subtyping cancers such as colorectal and lung carcinoma. The complexity of histological assessments, exacerbated by tumor heterogeneity and technical artifacts, necessitates the development of complementary diagnostic tools to enhance reproducibility and reduce observer variability. The authors note that while deep learning techniques have shown promise in automating tissue classification and cancer detection, the effectiveness of these models is heavily dependent on the availability of large, diverse, and high-quality training datasets, which are often limited due to challenges in data annotation, privacy regulations, and institutional barriers.
To address these limitations, the paper proposes the use of synthetic data generation, particularly through diffusion models, which can create high-resolution images that mimic real tissue morphology while preserving patient privacy. The authors emphasize the potential of few-shot learning and generative models to produce extensive synthetic datasets from minimal input, thereby enhancing model training. They introduce a unified framework that integrates diffusion-based few-shot data synthesis with parameter-efficient adaptation methods, such as Low-Rank Adaptation (LoRA), and systematically evaluates their combined impact across various classifier architectures. This approach aims to improve the classification performance of colorectal and lung cancer subtypes while addressing the challenges posed by data scarcity in histopathology. The main contributions include the generation of high-fidelity synthetic images, the development of a novel data generation approach, and a comprehensive evaluation of the proposed methodology on benchmark datasets.
Methods
The methodology outlined in this research paper is divided into two main subsections: Data Generation and Classification. In the Data Generation phase, high-fidelity synthetic images of colon and lung cancer subtypes are created using a diffusion-based approach, specifically employing the Stable Diffusion 2.1 model. This model, pre-trained on a vast image corpus, generates visuals from random noise guided by textual prompts. The training process minimizes a denoising loss, \(L_D\), to refine image quality, although it faces challenges in accurately representing specific cancer conditions. To address these limitations, the DataDream method is utilized, which incorporates Low-Rank Adaptation (LoRA) to efficiently fine-tune the model with a limited number of histopathological samples, making it suitable for scenarios with scarce data.
The synthetic image generation process is enhanced through the iterative design of class-specific prompts that describe critical visual characteristics necessary for accurate representation. These prompts are meticulously crafted to maximize the morphological fidelity of the generated images, with an emphasis on high-quality samples from the LC25000 dataset. The LoRA technique is applied to the linear projection layers of the model’s attention mechanisms, optimizing only the LoRA parameters while keeping the original model weights frozen. Fine-tuning is conducted using the AdamW optimizer, with early stopping employed to prevent overfitting. The subsequent experimental results will assess both the quality of the synthetic data and the performance of classification models trained on these datasets, thereby evaluating the effectiveness of the proposed methodology.
Discussion
In this section, the authors discuss the classification experiments conducted using various convolutional and transformer-based models trained on both real and synthetic datasets. The primary objective was to evaluate whether synthetic data, generated from a limited number of real images, could enhance model performance in detecting morphological characteristics in histopathological images. Convolutional Neural Networks (CNNs) were fine-tuned for binary and multiclass classification, leveraging their ability to extract hierarchical features, while transformer architectures utilized self-attention mechanisms to capture global context, which is particularly beneficial for distinguishing subtle features in rare diseases.
The study employed CLIP (Contrastive Language-Image Pretraining) models to facilitate cross-modal learning, allowing for the comparison of visual inputs and textual prompts in a shared embedding space. Two variants of CLIP, ViT-B/16 and ViT-B/32, were utilized to analyze the trade-off between representation granularity and computational efficiency. The models were trained using a cross-entropy loss function and the AdamW optimizer, with a structured training protocol to ensure fair comparisons between models trained on real and synthetic data. The findings indicate that the proposed synthetic data generation strategy effectively enhances classification accuracy and generalization capabilities under data-scarce conditions, demonstrating its potential applicability in clinical settings where annotated medical data is limited.
Limitations
The study presents several limitations that must be considered despite its promising findings. Firstly, while synthetic data enhanced classification performance in data-scarce scenarios, the generated images may not adequately represent the biological and morphological diversity found in real clinical environments, which can vary significantly due to factors such as institutional practices, staining protocols, and scanning technologies. Secondly, the evaluation of synthetic image quality, primarily based on the Fréchet Inception Distance (FID), relies on features from models trained on natural images. Although FID offers a relative comparison across experimental conditions, it may not fully capture clinically relevant histopathological features.
Additionally, the research is constrained by its focus on only two cancer types and a single public dataset, necessitating further validation across multi-center datasets and a wider range of diseases to enhance clinical applicability. The reliance on expert-designed prompts and the manual selection of high-quality few-shot samples may also limit the scalability of the approach. Future research will aim to automate prompt optimization, expand disease coverage, and involve clinical experts in evaluating the realism of synthetic images. Furthermore, while the method aims to minimize the sharing of raw patient images, it does not provide formal privacy guarantees, as the diffusion model is fine-tuned using real samples without differential privacy measures. Future investigations will address these privacy concerns by incorporating formal privacy-preserving mechanisms and conducting adversarial robustness testing to evaluate potential risks of data leakage.
