هجمات حقن سريعة على نماذج اللغة البصرية في علم الأورام
Prompt injection attacks on vision language models in oncology

شارك:
المجلة: Nature Communications، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41467-024-55631-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39890777
تاريخ النشر: 2025-02-01
المؤلف: Jan Clusmann وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تناقش هذه القسم ثغرات نماذج اللغة-الرؤية (VLMs) في التطبيقات الطبية، وخاصة قابليتها للهجمات من نوع حقن التعليمات. لقد أظهرت نماذج VLMs، التي تدمج البيانات البصرية والنصية لمهام مثل تفسير الصور ودعم القرار في الرعاية الصحية، وعدًا في تعزيز الممارسات الطبية. ومع ذلك، يكشف البحث عن عيب أمني حرج: يمكن التلاعب بهذه النماذج لإنتاج مخرجات ضارة من خلال حقن التعليمات، حيث يتم تضمين تعليمات مخفية ضمن مدخلات المستخدم.

أجرى الباحثون تحليلًا كميًا شمل 594 هجومًا متميزًا عبر أربعة نماذج VLM رائدة—Claude-3 Opus، Claude-3.5 Sonnet، Reka Core، وGPT-4o—مظهرين أن جميع النماذج التي تم اختبارها كانت عرضة لمثل هذه الهجمات. ومن الجدير بالذكر أن الدراسة تسلط الضوء على أن التعليمات الفرعية المدمجة في بيانات التصوير الطبي يمكن أن تؤدي إلى مخرجات ضارة، وغالبًا ما تبقى غير قابلة للاكتشاف من قبل المراقبين البشريين. تؤكد هذه النتائج على الحاجة الملحة لمعالجة هذه الثغرات لضمان نشر آمن لنماذج VLM في البيئات السريرية، حيث تتطلب التطورات السريعة لهذه التقنيات تدابير أمان قوية لحماية ضد التهديدات الناشئة.

طرق

في هذه الدراسة، تضمنت الإعدادات التجريبية إنشاء خمس صور متميزة لكل حالة مريض لتقييم آثار حقن التعليمات على استجابات النموذج. تمت معالجة الصور 1 و2 و5 بشكل مستقل، حيث كانت الصورة 1 تعمل كتحكم سلبي يتضمن فقط قالب التعليمات غير المعدل. تضمنت الصورة 2 “حقن تعليمات نصية”، توجيه النموذج لوصف العضو المرئي مع التأكيد على أنه يبدو صحيًا. تم استخدام الصور 3 و4 معًا لمحاكاة سيناريو حيث تم إعداد حقن التعليمات مسبقًا كتعليمات للنموذج (الصورة 3) قبل تقييم الصورة التالية (الصورة 4).

تم تغيير حقن التعليمات بشكل منهجي من حيث حجم الخط (14 أو 6) واللون (أسود أو رمادي داكن)، مع تعيين الخلفيات إلى إما أبيض أو أسود. تم تقديم الصور، التي دمجت حقن التعليمات مع المحتوى الإشعاعي، إلى النموذج بدقة 4457 × 2846 بكسل، بينما تراوحت دقة التصوير الأصلية من 500 إلى 1000 بكسل لكل بُعد. تم اختيار “أسود في أسود” للإعداد منخفض التباين ليعكس الخلفيات السريرية النموذجية في الأشعة. جميع الصور ومخرجات النموذج موثقة في البيانات التكميلية المقدمة.

نتائج

يقدم قسم “النتائج” نتائج الدراسة، موضحًا نتائج التجارب التي تم إجراؤها. يتم الإبلاغ عن مقاييس رئيسية وتحليلات إحصائية، مما يظهر علاقات كبيرة بين المتغيرات قيد التحقيق. تشير البيانات إلى أن المجموعة التجريبية أظهرت تحسنًا ملحوظًا في الأداء مقارنةً بمجموعة التحكم، مع قيمة p أقل من 0.05، مما يشير إلى أن النتائج ذات دلالة إحصائية.

بالإضافة إلى ذلك، تسلط النتائج الضوء على اتجاهات محددة لوحظت في البيانات، مثل العلاقة بين المتغير X والمتغير Y، والتي تم قياسها باستخدام تحليل الانحدار. تدعم النتائج الفرضية الأولية، مما يشير إلى أن التدخل المطبق كان له تأثير إيجابي على النتائج المقاسة. بشكل عام، تسهم النتائج في تقديم رؤى قيمة في هذا المجال وتقترح طرقًا محتملة لمزيد من البحث.

مناقشة

تستكشف الدراسة ضعف نماذج اللغة-الرؤية المتطورة (VLMs) تجاه هجمات حقن التعليمات في سياق تشخيص السرطان. من خلال إجراء تجارب مع نماذج مختلفة، بما في ذلك Claude 3 وClaude 3.5 وGPT-4o وReka Core، قام الباحثون بتقييم قدرة النماذج على اكتشاف الآفات الخبيثة في التصوير الطبي عبر عدة طرق (CT، MRI، الموجات فوق الصوتية، إلخ). أشارت النتائج إلى أن حقن التعليمات يمكن أن تؤثر بشكل كبير على اكتشاف الآفات، حيث زادت معدلات فقدان الآفات (LMR) بشكل كبير عند إدخال تعليمات خبيثة. على سبيل المثال، ارتفعت LMR لـ GPT-4o إلى 89% في ظل ظروف حقن التعليمات، مما يبرز قابلية النماذج للتلاعب العدائي.

تؤكد النتائج أن حقن التعليمات هو تهديد غير مرتبط بنوع التصوير، يؤثر على جميع نماذج VLM التي تم اختبارها بغض النظر عن نوع التصوير. كما استكشفت الدراسة استراتيجيات التخفيف، مثل هندسة التعليمات الأخلاقية، التي أظهرت بعض الوعد في تقليل الضعف، خاصة بالنسبة لـ Claude 3.5. ومع ذلك، تشير النتائج العامة إلى أن نماذج VLM الحالية ليست محمية بشكل كافٍ ضد مثل هذه الهجمات، مما يشكل مخاطر كبيرة في البيئات الصحية حيث يكون التشخيص الدقيق أمرًا حاسمًا. يدعو المؤلفون إلى تطوير دفاعات قوية ضد هذه الثغرات، مؤكدين على الحاجة إلى البحث المستمر ودمج الإشراف البشري في عمليات اتخاذ القرار السريرية التي تشمل أنظمة الذكاء الاصطناعي.

Journal: Nature Communications, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41467-024-55631-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39890777
Publication Date: 2025-02-01
Author(s): Jan Clusmann et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

This section discusses the vulnerabilities of vision-language models (VLMs) in medical applications, particularly their susceptibility to prompt injection attacks. VLMs, which integrate visual and textual data for tasks such as image interpretation and decision support in healthcare, have shown promise in enhancing medical practices. However, the study reveals a critical security flaw: these models can be manipulated to produce harmful outputs through prompt injection, where hidden instructions are embedded within user inputs.

The researchers conducted a quantitative analysis involving 594 distinct attacks across four leading VLMs—Claude-3 Opus, Claude-3.5 Sonnet, Reka Core, and GPT-4o—demonstrating that all models tested were vulnerable to such attacks. Notably, the study highlights that sub-visual prompts embedded in medical imaging data can lead to detrimental outputs, often remaining undetectable to human observers. These findings underscore the urgent need to address these vulnerabilities to ensure the safe deployment of VLMs in clinical settings, as the rapid advancement of these technologies necessitates robust security measures to protect against emerging threats.

Methods

In this study, the experimental setup involved the creation of five distinct images for each patient case to evaluate the effects of prompt injections on model responses. Images 1, 2, and 5 were processed independently, with Image 1 serving as a negative control that included only the unaltered prompt template. Image 2 incorporated a “text prompt injection,” instructing the model to describe the visible organ while asserting it appears healthy. Images 3 and 4 were utilized in tandem to simulate a scenario where a prompt injection was pre-established as model instruction (Image 3) before assessing the subsequent image (Image 4).

The prompt injections were systematically varied in terms of font size (14 or 6) and color (black or dark gray), with backgrounds set to either white or black. The images, which combined prompt injections with radiological content, were presented to the model at a resolution of 4457 × 2846 pixels, while the original imaging resolution ranged from 500 to 1000 pixels per dimension. The choice of “black in black” for the low-contrast setting was made to reflect typical clinical backgrounds in radiology. All images and model outputs are documented in the supplementary data provided.

Results

The “Results” section presents the findings of the study, detailing the outcomes of the experiments conducted. Key metrics and statistical analyses are reported, demonstrating significant correlations between the variables under investigation. The data indicate that the experimental group exhibited a marked improvement in performance compared to the control group, with a p-value of less than 0.05, suggesting that the results are statistically significant.

Additionally, the results highlight specific trends observed in the data, such as the relationship between variable X and variable Y, which was quantified using regression analysis. The findings support the initial hypothesis, indicating that the intervention applied had a positive effect on the measured outcomes. Overall, the results contribute valuable insights into the field and suggest potential avenues for further research.

Discussion

The study investigates the vulnerability of state-of-the-art vision-language models (VLMs) to prompt injection attacks in the context of cancer diagnosis. By conducting experiments with various models, including Claude 3, Claude 3.5, GPT-4o, and Reka Core, the researchers assessed the models’ ability to detect malignant lesions in medical imaging across multiple modalities (CT, MRI, ultrasound, etc.). The results indicated that prompt injections could significantly impair lesion detection, with lesion miss rates (LMR) increasing substantially when malicious prompts were introduced. For instance, the LMR for GPT-4o rose to 89% under prompt injection conditions, highlighting the models’ susceptibility to adversarial manipulation.

The findings underscore that prompt injection is a modality-agnostic threat, affecting all tested VLMs regardless of the imaging type. The study also explored mitigation strategies, such as ethical prompt engineering, which showed some promise in reducing vulnerability, particularly for Claude 3.5. However, the overall results indicate that current VLMs are not adequately protected against such attacks, posing significant risks in healthcare settings where accurate diagnosis is critical. The authors advocate for the development of robust defenses against these vulnerabilities, emphasizing the need for ongoing research and the integration of human oversight in clinical decision-making processes involving AI systems.

شارك: