MedPix 2.0: مجموعة بيانات طبية متعددة الأبعاد شاملة لتطبيقات الذكاء الاصطناعي المتقدمة مع توليد معزز بالاسترجاع ورسوم بيانية للمعرفة
MedPix 2.0: A Comprehensive Multimodal Biomedical Data Set for Advanced AI Applications with Retrieval Augmented Generation and Knowledge Graphs

شارك:
المجلة: Data Science and Engineering، المجلد: 11، العدد: 2
DOI: https://doi.org/10.1007/s41019-025-00297-8
تاريخ النشر: 2025-07-01
المؤلف: Irene Siragusa وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تقدم الورقة MedPix 2.0، مجموعة بيانات متعددة الوسائط تتضمن تقارير سريرية ومسحات CT/MR، وتتناول تحدي نقص البيانات الطبية عالية الجودة بسبب مخاوف الخصوصية. قام المؤلفون بتطوير خط أنابيب شبه آلي لاستخراج البيانات وتنظيمها، مما أسفر عن قاعدة بيانات منظمة في MongoDB. تم إنشاء واجهة مستخدم رسومية سهلة الاستخدام لتسهيل التنقل الفعال واستعلام مجموعة البيانات، مما يمكّن المستخدمين من تنزيل مخرجات منظمة مناسبة لتدريب نماذج الذكاء الاصطناعي.

تسلط الدراسة الضوء على التطبيقات المحتملة لـ MedPix 2.0 في تطوير نماذج الذكاء الاصطناعي متعددة الوسائط، مثل أنظمة استخراج المعلومات وتحليل الصور الطبية الآلي. يوضح المؤلفون فائدة مجموعة البيانات من خلال نموذج DR-Minerva، وهو نموذج لغة رؤية قائم على استرجاع المعلومات المعزز (VLM) يتنبأ بأجزاء الجسم والوسائط من الصور المدخلة. تهدف الأعمال المستقبلية إلى تعزيز قدرات النموذج من خلال دمج رسم بياني للمعرفة (KG) وتحسين واجهة المستخدم الرسومية بأدوات تصور متقدمة. يسمح الطابع القابل للتوسع لقاعدة بيانات MongoDB بإضافة حالات سريرية جديدة مع الالتزام باللوائح المتعلقة بالخصوصية، مما يجعل MedPix 2.0 موردًا قيمًا لأنظمة دعم القرار الطبي ومزيد من البحث في هذا المجال.

مقدمة

تناقش مقدمة هذه الورقة البحثية الاندماج المتزايد للذكاء الاصطناعي (AI) في إدارة وتحليل البيانات البيولوجية الطبية، مع تسليط الضوء على إمكانيته في تعزيز دقة التشخيص ودعم العلاج الشخصي. يؤكد المؤلفون على الحاجة الملحة لنماذج الذكاء الاصطناعي الموثوقة، والتي تتطلب مجموعات بيانات كبيرة لأداء فعال. ومع ذلك، فإن الطبيعة الحساسة للبيانات البيولوجية الطبية تطرح تحديات كبيرة في توفر البيانات، مما يؤدي إلى مخاوف الخصوصية التي تعيق إنشاء مجموعات بيانات عامة. لمعالجة هذه القضايا، تم إنشاء الفضاء الأوروبي لبيانات الصحة (EHDS)، مما يعزز إطارًا للوصول الآمن إلى البيانات الصحية الإلكترونية، وبالتالي يسهل تطوير تطبيقات الذكاء الاصطناعي في الرعاية الصحية.

تقدم الورقة MedPix®، وهي قاعدة بيانات طبية مفتوحة الوصول بارزة، وتعرض إعادة تنظيم جديدة لبياناتها في هيكل MongoDB غير العلاقي، يسمى MedPix 2.0. تعزز مجموعة البيانات الجديدة إمكانية الوصول وسهولة الاستخدام لتدريب نماذج لغة الرؤية (VLMs) من خلال السماح للمستخدمين بإنشاء مجموعات بيانات تدريب مخصصة وتبسيط استعلامات البيانات من خلال واجهة مستخدم رسومية سهلة الاستخدام. علاوة على ذلك، يوضح المؤلفون بنية وتدريب نموذجهم VLM، DR-Minerva، الذي يستخدم مجموعة بيانات MedPix 2.0 ويشمل رسم بياني للمعرفة (KG) لاستعلام المعلومات الطبية. يهدف هذا النظام المتكامل إلى العمل كأداة لدعم القرار الطبي، مما يوفر اقتراحات تشخيصية بناءً على الصور الطبية. تحدد الورقة هيكله، مع تخصيص الأقسام اللاحقة لحالة الفن في مجموعات البيانات متعددة الوسائط، وعملية بناء مجموعة البيانات، والنتائج التجريبية، والاتجاهات المستقبلية.

النتائج

في هذا القسم، يتم تقديم نتائج خط أنابيب متعدد الوسائط يستخدم نموذج لغة الرؤية (VLM) من البداية إلى النهاية لدعم التشخيص الطبي. تدمج البنية نموذج DR-Minerva مع رسم بياني للمعرفة (KG)، كما هو موضح في القسم 4.3. تقيم الدراسة قابلية استخدام مجموعة البيانات للمهام المعقدة من خلال استعلام KGs مع مطالبات ذات تعقيد متنوع: بسيطة، ومتوسطة التعقيد، ومعقدة. تتضمن المطالبات تدريجيًا معلومات إضافية عن المريض، culminating in a complex prompt that simulates a doctor’s query, including clinical history alongside imaging data.

تتم مناقشة مقاييس الأداء لتقييم مخرجات النموذج مقابل مجموعة من “الإجابات الذهبية” المستمدة من MedPix 2.0، مع التركيز على درجة BERT كمقياس رئيسي. يتم تبرير هذا الاختيار من خلال الحاجة إلى التشابه الدلالي بدلاً من التداخل النصي الصارم، حيث يمكن أن تختلف الاستجابات السريرية في التعبير بينما تنقل نفس المعنى. تشير النتائج إلى أن المطالبة البسيطة تحقق أفضل أداء لـ KGs مع عدد أكبر من العلاقات، بينما تتفوق المطالبة المعقدة في KGs مع اتصالات أقل، مما يبرز تأثير هيكل KG على توليد الاستجابة. تؤكد النتائج على أهمية تصميم المطالبات وخصائص KG في تحسين قدرات دعم التشخيص للنموذج.

المناقشة

في هذا القسم، يناقش المؤلفون تطوير وأهمية MedPix 2.0، مجموعة بيانات متعددة الوسائط مصممة لتطبيقات الذكاء الاصطناعي في المجال الطبي. تتناول مجموعة البيانات التحديات الشائعة المرتبطة بالبيانات الطبية، مثل مخاوف الخصوصية ونقص البيانات، من خلال استخدام خط أنابيب شبه آلي لتنظيم الحالات السريرية، ومسحات CT، وMR من قاعدة بيانات MedPix الأصلية. يتم تنظيم البيانات الهيكلية في مستندات JSON مخزنة في مثيل MongoDB، مما يسمح باستعلامات واسترجاع فعالة. لا يسهل هذا الإعداد تدريب نماذج الذكاء الاصطناعي فحسب، بل يعزز أيضًا سهولة الاستخدام للباحثين والمهنيين في الرعاية الصحية.

يبرز المؤلفون تعددية استخدام مجموعة البيانات، مشيرين إلى تطبيقاتها المحتملة في مهام متعددة مدفوعة بالذكاء الاصطناعي، بما في ذلك استخراج المعلومات من التقارير السريرية وتحليل الصور الطبية الآلي. كما يقدمون DR-Minerva، وهي بنية قائمة على RAG تدمج رسم بياني للمعرفة (KG) لتحسين الاقتراحات التشخيصية بناءً على مجموعة البيانات. يؤكد المؤلفون على قابلية توسيع هيكل MongoDB، مما يسمح بالتوسعات المستقبلية والامتثال للوائح الخصوصية، مما يجعل MedPix 2.0 موردًا قيمًا لتقدم الذكاء الاصطناعي في الرعاية الصحية. ستركز الأعمال المستقبلية على تعزيز قدرات النموذج التوليدية وتحسين واجهة المستخدم لتوفير تصور أفضل للبيانات وتفاعلها.

القيود

تناقش قسم القيود مجموعة بيانات MedPix 2.0، المصممة لتدريب الهياكل العصبية العميقة، وتبرز النموذج التجريبي الأولي الذي يجمع بين Flamingo القائم على RAG، Dr-Minerva، وLlama 3.1 لتوليد استجابات مطولة. تشير النتائج، التي تم تقييمها باستخدام مقاييس مثل الدقة، والاسترجاع، ودرجة F1، إلى أن النموذج يظهر أداءً قويًا في توليد استجابات متماسكة. ومع ذلك، يعترف المؤلفون بأن التركيز الرئيسي لعملهم ليس على بناء نموذج موثوق تمامًا، حيث لم يتم إجراء التحقق من قبل المهنيين الطبيين – وهي خطوة حاسمة لدعم التشخيص.

علاوة على ذلك، بينما تقدم MedPix 2.0 سهولة الاستخدام عبر إصداراتها فقط الصور، والنص فقط، ومتعددة الوسائط، يقترح المؤلفون أن دمج هذه المجموعة مع مجموعات بيانات أخرى يمكن أن يعزز هيكلها وفعاليتها، مما يحسن مهام تدريب الذكاء الاصطناعي الموجهة نحو دعم التشخيص. يمكن أن يؤدي هذا الدمج إلى مورد أكثر شمولاً لتطوير نماذج تشخيصية موثوقة.

Journal: Data Science and Engineering, Volume: 11, Issue: 2
DOI: https://doi.org/10.1007/s41019-025-00297-8
Publication Date: 2025-07-01
Author(s): Irene Siragusa et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

The paper presents MedPix 2.0, a multimodal dataset comprising clinical reports and CT/MR scans, addressing the challenge of limited high-quality medical data due to privacy concerns. The authors developed a semi-automated pipeline for data extraction and curation, resulting in a structured MongoDB database. A user-friendly Graphical User Interface (GUI) was created to facilitate efficient navigation and querying of the dataset, enabling users to download structured outputs suitable for training AI models.

The study highlights the potential applications of MedPix 2.0 in developing AI multimodal models, such as Information Extraction systems and automated medical image analysis. The authors demonstrate the utility of the dataset through the DR-Minerva model, a Retrieve Augmented Generation-based Vision Language Model (VLM) that predicts body parts and modalities from input images. Future work aims to enhance the model’s capabilities by integrating a Knowledge Graph (KG) and improving the GUI with advanced visualization tools. The scalable nature of the MongoDB database allows for the addition of new clinical cases while adhering to privacy regulations, making MedPix 2.0 a valuable resource for medical decision support systems and further research in the field.

Introduction

The introduction of this research paper discusses the increasing integration of Artificial Intelligence (AI) in the management and analysis of biomedical data, highlighting its potential to enhance diagnostic accuracy and support personalized treatment. The authors emphasize the critical need for trustworthy AI models, which require substantial datasets for effective performance. However, the sensitive nature of biomedical data poses significant challenges in data availability, leading to privacy concerns that hinder the creation of public datasets. To address these issues, the European Health Data Space (EHDS) has been established, promoting a framework for secure access to electronic health data, thereby facilitating the development of AI applications in healthcare.

The paper introduces MedPix®, a prominent open-access medical image database, and presents a novel reorganization of its data into a non-relational MongoDB structure, termed MedPix 2.0. This new dataset enhances accessibility and usability for training Vision Language Models (VLMs) by allowing users to create tailored training datasets and simplifying data queries through a user-friendly Graphical User Interface (GUI). Furthermore, the authors detail the architecture and training of their VLM, DR-Minerva, which utilizes the MedPix 2.0 dataset and incorporates a Knowledge Graph (KG) for querying medical information. This integrated system aims to function as a medical decision support tool, providing diagnostic suggestions based on medical images. The paper outlines its structure, with subsequent sections dedicated to the state of the art in multimodal datasets, the dataset building process, experimental results, and future directions.

Results

In this section, the results of a multimodal pipeline utilizing an end-to-end Vision-Language Model (VLM) for medical diagnosis support are presented. The architecture integrates the DR-Minerva model with a Knowledge Graph (KG), as detailed in Section 4.3. The study evaluates the usability of the dataset for complex tasks by querying the KGs with prompts of varying complexity: simple, mid-complex, and complex. The prompts progressively incorporate additional patient information, culminating in a complex prompt that simulates a doctor’s query, including clinical history alongside imaging data.

Performance metrics for evaluating the model’s output against a set of “golden” answers derived from MedPix 2.0 are discussed, with a focus on BERT score as the primary metric. This choice is justified by the need for semantic similarity rather than strict textual overlap, as clinical responses can vary in expression while conveying the same meaning. Results indicate that the simple prompt yields the best performance for KGs with a higher number of relations, while the complex prompt excels in KGs with fewer connections, highlighting the influence of KG structure on response generation. The findings underscore the importance of prompt design and KG characteristics in optimizing the model’s diagnostic support capabilities.

Discussion

In this section, the authors discuss the development and significance of MedPix 2.0, a multimodal dataset designed for AI applications in the medical domain. The dataset addresses common challenges associated with medical data, such as privacy concerns and data scarcity, by utilizing a semi-automated pipeline to curate clinical cases, CT, and MR scans from the original MedPix database. The structured data is organized into JSON documents stored in a MongoDB instance, allowing for efficient querying and retrieval. This setup not only facilitates the training of AI models but also enhances usability for researchers and healthcare professionals.

The authors highlight the dataset’s versatility, noting its potential applications in various AI-driven tasks, including information extraction from clinical reports and automated medical image analysis. They also introduce DR-Minerva, a RAG-based architecture that integrates a knowledge graph (KG) to improve diagnostic suggestions based on the dataset. The authors emphasize the scalability of the MongoDB structure, which allows for future expansions and compliance with privacy regulations, making MedPix 2.0 a valuable resource for advancing AI in healthcare. Future work will focus on enhancing the model’s generative capabilities and improving the user interface for better data visualization and interaction.

Limitations

The section on limitations discusses the MedPix 2.0 dataset, designed for training deep neural architectures, and highlights the initial experimental model combining a RAG-based Flamingo, Dr-Minerva, and Llama 3.1 for generating verbose responses. The results, evaluated using metrics such as Precision, Recall, and F1 score, indicate that the model demonstrates robust performance in generating coherent responses. However, the authors acknowledge that the primary focus of their work is not on constructing a fully reliable model, as validation by medical professionals—a critical step for diagnostic support—has not been conducted.

Furthermore, while MedPix 2.0 offers ease of use across its Image-only, Text-only, and Multimodal versions, the authors suggest that integrating this dataset with other datasets could enhance its structure and effectiveness, thereby improving AI training tasks aimed at diagnostic support. This integration could lead to a more comprehensive resource for developing reliable diagnostic models.

شارك: