DOI: https://doi.org/10.1038/s41591-023-02728-3
PMID: https://pubmed.ncbi.nlm.nih.gov/38317019
تاريخ النشر: 2024-02-01
المؤلف: Matthew Groh وآخرون
الموضوع الرئيسي: كشف وإدارة الميلانوما الجلدية
نظرة عامة
يقدم هذا القسم مصفوفة ارتباك تقارن بين تسميات الإجماع، التي تم التحقق منها من قبل ثلاثة أطباء جلدية معتمدين من المجلس، مع التشخيصات الأولية المقدمة من أطباء جلدية معتمدين من المجلس. تصنف المصفوفة مختلف الأمراض الجلدية إلى تسميات محددة: ‘O’ للأخرى، والتي تشمل 38 حالة جلدية متميزة؛ ‘AD’ لالتهاب الجلد التأتبي؛ ‘CTCL’ لليمفوما الخلوية التائية الجلدية؛ ‘D’ لالتهاب الجلد العضلي؛ ‘LP’ لليكن الخلوي؛ ‘L’ لمرض لايم؛ ‘PR’ لوردية قشرية؛ ‘PRP’ لوردية حمراء شعرية؛ و ‘SS’ للزهري الثانوي.
تهدف هذه المقارنة إلى تقييم دقة التشخيص والاتفاق بين أطباء الجلدية، مما يوفر رؤى حول موثوقية التشخيصات السريرية في علم الأمراض الجلدية. تسلط التصنيف التفصيلي للأمراض داخل مصفوفة الارتباك الضوء على تعقيد تشخيص الأمراض الجلدية وإمكانية حدوث تصنيفات خاطئة.
الطرق
في هذه الدراسة، تم تطوير موقع ويب مخصص لتسهيل تجربة دقة التشخيص، يمكن الوصول إليه على https://diagnosing-diagnosis.media.mit.edu. قدم المشاركون أولاً موافقة مستنيرة وأجابوا على أسئلة تمهيدية قبل تلقي التعليمات التي أكدت على عدم دقة نموذج الذكاء الاصطناعي دون الكشف عن معدلات الدقة المحددة. كانت هذه الطريقة تهدف إلى منع التحيز في التزام المشاركين باقتراحات الذكاء الاصطناعي ولتمثيل سيناريوهات العالم الحقيقي حيث قد يكون الأطباء غير مدركين لدقة الذكاء الاصطناعي في سياقهم المحدد. بدأت التجربة بصورة بسيطة لامرأة تعاني من حب الشباب، مما سمح للباحثين بتقييم تفاعل المشاركين وجديتهم.
تم تكليف المشاركين بتشخيص الحالة الجلدية من خلال تقديم ما يصل إلى ثلاثة تشخيصات تفريقية، مع الإشارة إلى مستويات ثقتهم، واتخاذ قرار بشأن الإحالات لمزيد من التقييم. كانت التصميم يفضل تنسيقات الاستجابة الحرة على خيارات متعددة الاختيارات لتعزيز الصلاحية البيئية، حيث أن الاستجابات الحرة تقلل من التخمين ولا تهيئ المشاركين بتشخيصات محتملة. لمساعدة المشاركين، تم تنفيذ نص تنبؤي استنادًا إلى قائمة شاملة من 445 تشخيصًا ممكنًا، مما يعزز الاستجابات الموحدة. بعد تقديماتهم، تلقى المشاركون تعليقات حول توقع الذكاء الاصطناعي وتم منحهم خيارات لتعديل تشخيصاتهم وفقًا لذلك. استمرت التجربة مع صور متعددة، culminating in a comparative analysis of diagnostic accuracy among different participant groups after ten images.
النتائج
يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات المدروسة، حيث تكشف التحليلات الإحصائية عن قيمة p أقل من 0.05، مما يشير إلى أن النتائج ذات دلالة إحصائية. بالإضافة إلى ذلك، تظهر النتائج اتجاهًا واضحًا في سلوك النظام قيد التحقيق، كما هو موضح في الأشكال والجداول المرفقة.
علاوة على ذلك، تشير نتائج تحليل التباين (ANOVA) إلى أن الفروق بين المجموعات كبيرة، مما يدعم الفرضية المطروحة في المقدمة. تشمل النتائج أيضًا نتائج عددية محددة، مثل القيم المتوسطة والانحرافات المعيارية، والتي توفر نظرة شاملة على توزيع البيانات. بشكل عام، تسهم هذه النتائج في تقديم رؤى قيمة حول سؤال البحث وتضع الأساس لمزيد من المناقشة والتداعيات في الأقسام اللاحقة.
المناقشة
في هذه الدراسة، قمنا بتقييم دقة التشخيص لمجموعات الأطباء المختلفة—أطباء الجلدية المعتمدين من المجلس (BCDs)، المقيمين في الأمراض الجلدية، أطباء الرعاية الأولية (PCPs)، وأطباء آخرين—عند تشخيص الأمراض الجلدية الالتهابية في بيئة محاكاة للتطبيب عن بُعد. تم جمع ما مجموعه 14,261 تشخيصًا تفريقيًا من 1,118 مشاركًا عبر 364 صورة، مع التركيز على ثمانية أمراض جلدية رئيسية. أشارت النتائج إلى أن BCDs حققوا دقة تشخيصية من بين الثلاثة الأوائل بنسبة 38%، بينما كان لدى PCPs دقة بنسبة 19%. ومن الجدير بالذكر أن أداء التشخيص اختلف بشكل كبير بناءً على لون البشرة؛ حيث تم تشخيص صور البشرة الداكنة بدقة أقل مقارنة بتلك ذات البشرة الفاتحة، مع إظهار BCDs و PCPs انخفاضًا في دقة الثلاثة الأوائل بنسبة خمس نقاط مئوية للبشرة الداكنة.
استكشفت الدراسة أيضًا تأثير دعم القرار من أنظمة التعلم العميق (DLS) على دقة التشخيص. حسّن دعم DLS دقة BCDs من 27% إلى 36% و PCPs من 13% إلى 22%. ومع ذلك، كشفت النتائج أنه بينما كان دعم DLS عمومًا يعزز الأداء التشخيصي، فإنه أيضًا زاد من الفجوات الحالية في الدقة بالنسبة لـ PCPs عند تشخيص صور البشرة الداكنة. يبرز هذا الحاجة إلى النظر بعناية في التحيزات الخوارزمية وأهمية التدريب والتنوع في تعرض المرضى لتحسين نتائج التشخيص عبر ألوان البشرة المختلفة. بشكل عام، تؤكد النتائج على تعقيدات تشخيص الأمراض الجلدية في سياق التطبيب عن بُعد والدور المحتمل للذكاء الاصطناعي في تعزيز اتخاذ القرارات السريرية.
DOI: https://doi.org/10.1038/s41591-023-02728-3
PMID: https://pubmed.ncbi.nlm.nih.gov/38317019
Publication Date: 2024-02-01
Author(s): Matthew Groh et al.
Primary Topic: Cutaneous Melanoma Detection and Management
Overview
The section presents a confusion matrix that compares consensus labels, verified by three board-certified dermatologists, with primary diagnoses submitted by board-certified dermatologists. The matrix categorizes various skin diseases into specific labels: ‘O’ for other, encompassing 38 distinct skin conditions; ‘AD’ for atopic dermatitis; ‘CTCL’ for cutaneous T-cell lymphoma; ‘D’ for dermatomyositis; ‘LP’ for lichen planus; ‘L’ for Lyme disease; ‘PR’ for pityriasis rosea; ‘PRP’ for pityriasis rubra pilaris; and ‘SS’ for secondary syphilis.
This comparison aims to assess the diagnostic accuracy and agreement among dermatologists, providing insights into the reliability of clinical diagnoses in dermatology. The detailed categorization of diseases within the confusion matrix highlights the complexity of skin disease diagnosis and the potential for misclassification.
Methods
In this study, a custom website was developed to facilitate a diagnostic accuracy experiment, accessible at https://diagnosing-diagnosis.media.mit.edu. Participants first provided informed consent and answered preliminary questions before receiving instructions that emphasized the AI model’s imperfect accuracy without disclosing specific accuracy rates. This approach aimed to prevent bias in participants’ adherence to the AI’s suggestions and to simulate real-world scenarios where physicians may be unaware of an AI’s accuracy in their specific context. The experiment commenced with a straightforward image of a woman with acne, allowing researchers to gauge participant engagement and seriousness.
Participants were tasked with diagnosing the skin condition by providing up to three differential diagnoses, indicating their confidence levels, and deciding on referrals for further evaluation. The design favored free-response formats over multiple-choice options to enhance ecological validity, as free responses mitigate guessing and do not prime participants with potential diagnoses. To assist participants, predictive text was implemented based on a comprehensive list of 445 possible diagnoses, promoting standardized responses. Following their submissions, participants received feedback on the AI’s prediction and were given options to modify their diagnoses accordingly. The experiment continued with multiple images, culminating in a comparative analysis of diagnostic accuracy among different participant groups after ten images.
Results
The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. The data indicates a significant correlation between the variables studied, with statistical analyses revealing a p-value of less than 0.05, suggesting that the results are statistically significant. Additionally, the results demonstrate a clear trend in the behavior of the system under investigation, as illustrated by the accompanying figures and tables.
Furthermore, the analysis of variance (ANOVA) results indicate that the differences among the groups are substantial, supporting the hypothesis posited in the introduction. The findings also include specific numerical outcomes, such as mean values and standard deviations, which provide a comprehensive overview of the data distribution. Overall, these results contribute valuable insights into the research question and lay the groundwork for further discussion and implications in subsequent sections.
Discussion
In this study, we evaluated the diagnostic accuracy of various physician groups—board-certified dermatologists (BCDs), dermatology residents, primary-care physicians (PCPs), and other physicians—when diagnosing inflammatory skin diseases in a simulated store-and-forward teledermatology setting. A total of 14,261 differential diagnoses were collected from 1,118 participants across 364 images, with a focus on eight main skin diseases. The results indicated that BCDs achieved a top-3 diagnostic accuracy of 38%, while PCPs had an accuracy of 19%. Notably, diagnostic performance varied significantly based on skin tone; images of dark skin were diagnosed with lower accuracy compared to those of light skin, with BCDs and PCPs showing reductions in top-1 and top-3 accuracies by five percentage points for dark skin.
The study also explored the impact of decision support from deep learning systems (DLS) on diagnostic accuracy. DLS assistance improved BCDs’ top-1 accuracy from 27% to 36% and PCPs’ from 13% to 22%. However, the results revealed that while DLS support generally enhanced diagnostic performance, it also exacerbated existing accuracy disparities for PCPs when diagnosing dark skin images. This highlights the need for careful consideration of algorithmic biases and the importance of training and diversity in patient exposure for improving diagnostic outcomes across different skin tones. Overall, the findings underscore the complexities of diagnosing skin diseases in a teledermatology context and the potential role of AI in enhancing clinical decision-making.
