تحسين تصنيف الجسيمات تحت المرئية في مجهر التصوير التدفق عبر توليد الصور المعتمد على الذكاء الاصطناعي
Improved sub-visible particle classification in flow imaging microscopy via generative AI-based image synthesis

شارك:
المجلة: Journal of Pharmaceutical Sciences، المجلد: 115، العدد: 5
DOI: https://doi.org/10.1016/j.xphs.2026.104225
PMID: https://pubmed.ncbi.nlm.nih.gov/41763492
تاريخ النشر: 2026-02-26
المؤلف: Utku Özbulak وآخرون
الموضوع الرئيسي: التصوير الرقمي لأمراض الدم

نظرة عامة

تستكشف هذه الدراسة تطبيق مجهر تصوير التدفق مع التعلم العميق لتحليل الجسيمات تحت المرئية (SvP)، مع التركيز بشكل خاص على التحديات التي تطرحها ندرة البيانات وعدم التوازن بين أنواع الجسيمات. تسلط الدراسة الضوء على الصعوبة في تمييز الجسيمات الأقل شيوعًا، مثل زيت السيليكون وفقاعات الهواء، عن الجسيمات البروتينية الأكثر شيوعًا بسبب توافر الصور المحدود للأولى. لمواجهة هذه المشكلة، يقترح المؤلفون نموذج انتشار جديد يولد صورًا اصطناعية عالية الدقة لتعزيز مجموعات بيانات التدريب، مما يسهل التدريب الفعال لشبكات الأعصاب العميقة متعددة الفئات.

تظهر النتائج أن دمج هذه الصور الاصطناعية يعزز بشكل كبير توازن مجموعات البيانات ويحسن الأداء التنبؤي للمصنفات. لا يقلل هذا النهج من الاعتماد على التوصيف اليدوي فحسب، بل يظهر أيضًا إمكانيات الذكاء الاصطناعي التوليدي في مراقبة الجودة في صناعة الأدوية. يقترح المؤلفون أن نماذج الانتشار يمكن أن توفر حلولًا قابلة للتوسع لندرة البيانات، مما يؤدي في النهاية إلى اتخاذ قرارات أكثر موثوقية تعتمد على البيانات في هذا المجال.

مقدمة

تناقش المقدمة أهمية الجسيمات تحت المرئية (SvPs) في العلاجات القائمة على البروتين، مع تسليط الضوء على قدرتها على التسبب في آثار سلبية مثل الاستجابات المناعية وتقليل فعالية الأدوية. وضعت الهيئات التنظيمية مثل دستور الأدوية الأمريكي (USP) ودستور الأدوية الأوروبي إرشادات لمراقبة المواد الجسيمية، مع التركيز بشكل أساسي على عدد الجسيمات مع تجاهل الخصائص الكيميائية والشكلية التي تعتبر حاسمة لتقييم المخاطر. يمكن أن تنشأ SvPs من مصادر داخلية، جوهرية، وخارجية، مع ارتباط بعض الأنواع بانخفاض الفعالية العلاجية والاستقرار. على سبيل المثال، يمكن أن تؤدي تجمعات البروتين إلى استجابات مناعية، بينما قد تحفز جسيمات زيت السيليكون، على الرغم من كونها عمومًا منخفضة السمية، مزيدًا من تجمع البروتين.

لمواجهة تحدي تحديد وتصنيف SvPs، يقترح المؤلفون استخدام مجهر تصوير التدفق (FIM) مع تقنيات التعلم العميق. ومع ذلك، فإن فعالية التعلم العميق تعوقها الحاجة إلى مجموعات بيانات كبيرة ومتوازنة، خاصة لأنواع الجسيمات الممثلة تمثيلًا ناقصًا مثل قطرات زيت السيليكون وفقاعات الهواء. يتناول المؤلفون هذا عدم التوازن في البيانات من خلال تطوير نموذج توليدي قائم على الانتشار لإنشاء صور اصطناعية عالية الدقة لهذه الفئات الأقل تمثيلًا. يهدف هذا النهج إلى تعزيز تدريب المصنفات متعددة الفئات من خلال تكملة مجموعات البيانات الحالية، مما يحسن دقة التصنيف ويدعم مراقبة الجودة الصيدلانية القوية. كما يلتزم المؤلفون بتعزيز العلوم المفتوحة من خلال إصدار نموذجهم والموارد المرتبطة به علنًا.

الطرق

في هذه الدراسة، تم استخدام ثمانية أجسام مضادة وحيدة النسيلة (mAbs) للتحقيق في تكوين تجمعات البروتين تحت ظروف ضغط مختلفة. تم إعداد الأجسام المضادة في نظامي عازلة متميزين: عازلة أسيتات 10 مليمول عند pH 5 بتركيز 0.5 ملغ/مل لتجارب ضغط الحرارة، وعازلة هيستيدين 10 مليمول عند pH 6 تحتوي على 1 ملغ/مل من الأجسام المضادة، 0.9% NaCl، و0.05% بوليسوربات 20 لظروف الضغط الميكانيكي. بالإضافة إلى ذلك، تم استخدام 0.9% NaCl لتسهيل استعادة قطرات زيت السيليكون من الحقن المجهزة بالسيليكون (5 مل بلاستيباك لوير لوك، BD، فرانكلين ليكس، الولايات المتحدة). قبل التجربة، تم تصفية جميع المحاليل من خلال مرشحات 0.22 ميكرون لضمان التعقيم والنقاء، مع كون جميع المواد الكيميائية المستخدمة من الدرجة الصيدلانية أو أعلى.

توفر المنهجية الموضحة إطارًا قويًا لتقييم تأثير الضغط الحراري والميكانيكي على استقرار وتجمع الأجسام المضادة، وهو أمر حاسم لفهم سلوك البروتينات العلاجية في ظروف الصياغة والتخزين المختلفة.

المناقشة

تؤكد قسم المناقشة في ورقة البحث على فعالية النماذج التوليدية القائمة على الانتشار في معالجة مشكلة عدم التوازن في البيانات السائدة في تصنيف الجسيمات تحت المرئية (SvP). من خلال توليد صور اصطناعية عالية الدقة للفئات الممثلة تمثيلًا ناقصًا، مثل قطرات زيت السيليكون وفقاعات الهواء، تظهر الدراسة أنه يمكن بناء مجموعات بيانات تدريب متوازنة دون الحاجة إلى توصيف يدوي مكثف. لا يعزز هذا النهج فقط أداء المصنفات متعددة الفئات، بل يظهر أيضًا قابلية تكيف الطريقة عبر أنواع الجسيمات المختلفة وطرائق التصوير، مما يجعلها أداة متعددة الاستخدامات لمراقبة الجودة في صناعة الأدوية.

تشير النتائج إلى أن دمج الصور الاصطناعية يحسن بشكل كبير الأداء التنبؤي، كما يتضح من النتائج من كل من نماذج ResNet-18 وResNet-50. تسلط الدراسة الضوء على إمكانيات الذكاء الاصطناعي التوليدي لتبسيط عمليات مراقبة الجودة، وتقليل الاعتماد على التوصيف اليدوي، ودعم الامتثال التنظيمي. علاوة على ذلك، يدعو المؤلفون إلى التطبيق الأوسع للنماذج التوليدية في التصوير الطبي الحيوي، مقترحين أن هذا الإطار يمكن أن يسهل حلول التعلم الآلي القوية في المجالات الحساسة للسلامة. لتعزيز القابلية للتكرار ومزيد من البحث، تم إتاحة مجموعات البيانات التي تم إنشاؤها، والنماذج المدربة، والمصنفات للجمهور.

Journal: Journal of Pharmaceutical Sciences, Volume: 115, Issue: 5
DOI: https://doi.org/10.1016/j.xphs.2026.104225
PMID: https://pubmed.ncbi.nlm.nih.gov/41763492
Publication Date: 2026-02-26
Author(s): Utku Özbulak et al.
Primary Topic: Digital Imaging for Blood Diseases

Overview

This research explores the application of flow imaging microscopy combined with deep learning for sub-visible particle (SvP) analysis, specifically focusing on the challenges posed by data scarcity and class imbalance among particle types. The study highlights the difficulty in distinguishing less prevalent particles, such as silicone oil and air bubbles, from more common protein particles due to the limited availability of images for the former. To tackle this issue, the authors propose a novel diffusion model that generates high-fidelity synthetic images to augment training datasets, thereby facilitating the effective training of multi-class deep neural networks.

The findings demonstrate that the integration of these synthesized images significantly enhances the balance of the datasets and improves the predictive performance of the classifiers. This approach not only reduces the reliance on manual annotation but also showcases the potential of generative AI in pharmaceutical quality control. The authors suggest that diffusion models could provide scalable solutions to data scarcity, ultimately leading to more reliable, data-driven decision-making in the field.

Introduction

The introduction discusses the significance of sub-visible particles (SvPs) in protein-based therapeutics, highlighting their potential to induce adverse effects such as immunogenic responses and reduced drug efficacy. Regulatory bodies like the United States Pharmacopeia (USP) and the European Pharmacopoeia have established guidelines for monitoring particulate matter, focusing primarily on particle count while neglecting the chemical and morphological characteristics that are crucial for risk assessment. SvPs can originate from inherent, intrinsic, and extrinsic sources, with certain types linked to decreased therapeutic potency and stability. For instance, protein aggregates can elicit immunogenic reactions, while silicone oil particles, although generally low in toxicity, may catalyze further protein aggregation.

To tackle the challenge of identifying and classifying SvPs, the authors propose using flow-imaging microscopy (FIM) combined with deep learning techniques. However, the effectiveness of deep learning is hindered by the need for large, balanced datasets, particularly for underrepresented particle types like silicone oil droplets and air bubbles. The authors address this data imbalance by developing a diffusion-based generative model to create high-fidelity synthetic images of these minority classes. This approach aims to enhance the training of multi-class classifiers by supplementing existing datasets, thereby improving classification accuracy and supporting robust pharmaceutical quality control. The authors also commit to promoting open science by publicly releasing their model and associated resources.

Methods

In this study, eight commercial monoclonal antibodies (mAbs) were utilized to investigate the formation of protein aggregates under different stress conditions. The mAbs were prepared in two distinct buffer systems: a 10 mM acetate buffer at pH 5 with a concentration of 0.5 mg/ml for heat stress experiments, and a 10 mM histidine buffer at pH 6 containing 1 mg/ml of mAbs, 0.9% NaCl, and 0.05% polysorbate 20 for mechanical stress conditions. Additionally, 0.9% NaCl was employed to facilitate the recovery of silicone oil droplets from siliconized syringes (5 ml Plastipak Luer Lock, BD, Franklin Lakes, US). Prior to experimentation, all solutions were filtered through 0.22 µm filters to ensure sterility and purity, with all chemicals used being of pharmaceutical grade or higher.

The methodology outlined provides a robust framework for assessing the impact of thermal and mechanical stress on mAb stability and aggregation, which is critical for understanding the behavior of therapeutic proteins in various formulation and storage conditions.

Discussion

The discussion section of the research paper emphasizes the effectiveness of diffusion-based generative models in addressing the data imbalance problem prevalent in sub-visible particle (SvP) classification. By generating high-fidelity synthetic images of underrepresented classes, such as silicone oil droplets and air bubbles, the study demonstrates that balanced training datasets can be constructed without extensive manual annotation. This approach not only enhances the performance of multi-class classifiers but also showcases the adaptability of the method across various particle types and imaging modalities, making it a versatile tool for quality control in pharmaceutical manufacturing.

The findings indicate that the integration of synthetic images significantly improves predictive performance, as evidenced by the results from both ResNet-18 and ResNet-50 models. The study highlights the potential of generative AI to streamline quality control processes, reduce reliance on manual labeling, and support regulatory compliance. Furthermore, the authors advocate for the broader application of generative models in biomedical imaging, suggesting that this framework could facilitate robust machine learning solutions in safety-critical domains. To foster reproducibility and further research, the generated datasets, trained models, and classifiers have been made publicly available.

شارك: