DOI: https://doi.org/10.1016/s2589-7500(24)00224-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39701919
تاريخ النشر: 2024-12-18
المؤلف: Joseph Alderman وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم
نظرة عامة
تؤكد جامعة برمنغهام التزامها بالاهتمام الدقيق بالعناصر المتاحة؛ ومع ذلك، تعترف بأنه قد تحدث أخطاء أحيانًا في عملية التحميل. بالإضافة إلى ذلك، قد يتم تصنيف بعض الوثائق على أنها حساسة تجاريًا أو مقيدة بطرق أخرى. في مثل هذه الحالات، تشجع الجامعة الأفراد على الإبلاغ عن المخاوف من خلال الاتصال بعنوان البريد الإلكتروني المخصص، UBIRA@lists.bham.ac.uk، لتسهيل إزالة الوثيقة المشكلة.
مقدمة
تناقش مقدمة هذه الورقة البحثية النمو السريع للذكاء الاصطناعي (AI) في الرعاية الصحية على مدار العقد الماضي، مع التأكيد على إمكانيته في تعزيز الوصول إلى التشخيصات والعلاجات المتخصصة بينما يخفف الضغط على الموارد الصحية المحدودة. ومع ذلك، تؤكد على الحاجة الملحة لمعالجة المخاطر المرتبطة بالتحيزات الخوارزمية التي قد تفاقم عدم المساواة الصحية. مستشهدة بدراسات مثل تلك التي أجراها أوبيرماير وآخرون، توضح الورقة كيف يمكن أن تؤدي التحيزات في خوارزميات الذكاء الاصطناعي إلى تقدير منهجي غير كافٍ للاحتياجات الصحية للفئات المهمشة، مما يبرز أهمية تمثيل البيانات والدقة في تدريب هذه الخوارزميات.
يحدد المؤلفون عدة آليات يمكن أن تنشأ من خلالها التحيزات، بما في ذلك التمثيل غير الكافي للفئات الأقلية في مجموعات بيانات الصحة، والبيانات غير المكتملة أو غير الدقيقة للسكان المحرومين، وترميز عدم المساواة الاجتماعية داخل مجموعات البيانات. يدعون إلى الشفافية والتقييم الذاتي النقدي لمجموعات البيانات لفهم قيودها وتحاملاتها بشكل أفضل. كما توضح المقدمة برنامج STANDING Together، وهو مبادرة دولية تهدف إلى تطوير توصيات لتحسين تنوع البيانات، والشمولية، وقابلية التعميم في تقنيات الصحة المعتمدة على الذكاء الاصطناعي. تختتم الورقة بالتأكيد على أن اعتماد هذه التوصيات أمر ضروري لضمان أن تكون تقنيات الصحة المعتمدة على الذكاء الاصطناعي عادلة وفعالة لجميع السكان.
الطرق
تم تطوير الطرق المستخدمة في توصيات STANDING Together من خلال برنامج بحثي مختلط الأساليب يمتد من ديسمبر 2021 إلى نوفمبر 2023. تم إنشاء مجموعة عمل دولية تضم 23 خبيرًا من مجالات متنوعة—بما في ذلك الرعاية الصحية، وعلوم الكمبيوتر، والسياسة، والتنظيم، والأكاديميا—لإدارة المشروع. تضمنت هذه المجموعة اثنين من المشاركين المرضى لضمان أن تكون التوصيات مستندة إلى تجارب حقيقية من عدم المساواة الصحية. بالإضافة إلى ذلك، اجتمعت لجنة مشاركة المرضى والجمهور (PPIE)، المكونة من 12 عضوًا، ربع سنويًا لتقديم رؤى تضمن أن تكون لغة ومحتوى التوصيات شاملة ومحترمة. تم الحصول على الموافقة الأخلاقية للدراسة من لجنة المراجعة الأخلاقية بجامعة برمنغهام، وتم نشر التفاصيل الأولية علنًا على موقع المعهد الوطني للبحوث الصحية والرعاية قبل بدء المشروع.
في سياق الذكاء الاصطناعي (AI)، تُعرف طرق العدالة بأنها أساليب إحصائية تقيم مقاييس الأداء عبر مجموعات ديموغرافية مختلفة تتأثر بخوارزمية. تتوفر تقنيات خوارزمية متنوعة لتطوير نماذج تلتزم بقيود العدالة، بهدف تقليل الفجوات في الأداء بين المجموعات الفرعية. تؤكد توصية معينة (2.2f) على ضرورة أن يقوم مستخدمو البيانات بتوثيق أي طرق تم استخدامها لضبط الأداء عبر المجموعات بشكل متعمد. يشمل ذلك تفصيل الأسباب والأهداف والأساليب والمقاييس وأي تغييرات في العتبات المحددة لمجموعات مختلفة طوال دورة حياة تقنية الصحة المعتمدة على الذكاء الاصطناعي.
النتائج
في قسم النتائج، تؤكد التوصية 2.3c على أهمية الإبلاغ عن النتائج من التقييمات السابقة لتقنيات الصحة المعتمدة على الذكاء الاصطناعي ومجموعات البيانات المستخدمة في تطويرها. تقترح أن يقوم مستخدمو البيانات بمراجعة التقييمات المتاحة بدقة لفهم آثارها على مجموعات مختلفة ضمن السكان المستهدفين. يشمل ذلك تقييم المخاطر المحتملة للأذى المرتبطة بالتكنولوجيا، كما هو مفصل في الملحق (ص. 39). تؤكد التوصية على ضرورة الشفافية والعناية الواجبة في تقييم تقنيات الصحة المعتمدة على الذكاء الاصطناعي لضمان السلامة والفعالية لجميع الفئات السكانية.
المناقشة
طور برنامج STANDING Together مجموعة شاملة من التوصيات تهدف إلى معالجة التحيزات وتعزيز العدالة في تقنيات الصحة المعتمدة على الذكاء الاصطناعي. نشأت التوصيات من مراجعة منهجية لمعايير بيانات الصحة الحالية واستطلاع رأي خبراء بيانات الصحة، مما أبلغ قائمة طويلة أولية من العناصر المرشحة. تم تنقيح هذه القائمة من خلال دراسة دلفي معدلة تضمنت عدة جولات من التصويت والتعليقات من مجموعة متنوعة من أصحاب المصلحة، بما في ذلك المرضى وممثلي الجمهور. انتهى عملية التوافق في اجتماع نهائي حضوري حيث صوت 26 مشاركًا على التوصيات، مما أسفر عن 29 عنصرًا نهائيًا تعكس مجموعة واسعة من آراء الخبراء ورؤى أصحاب المصلحة.
تم تصميم عملية دلفي التكرارية لضمان أن تكون التوصيات ذات صلة ثقافية وقابلة للتطبيق عبر سياقات متنوعة. تم طلب التعليقات في كل مرحلة، بما في ذلك الاستشارات العامة ومدخلات من مجموعة استشارية دولية، لتعزيز فائدة التوصيات. تهدف التوصيات النهائية، التي سيتم إصدارها في أواخر عام 2023، إلى التخفيف من التحيزات في تقنيات الصحة المعتمدة على الذكاء الاصطناعي من خلال مراعاة عوامل مثل أصل البيانات، والافتقار، والتقاطع بين محددات الصحة. يبرز هذا النهج أهمية الحوار الشامل وبناء التوافق في تطوير الإرشادات التي يمكن أن تعالج بشكل فعال الفجوات في نتائج الصحة المرتبطة بتطبيقات الذكاء الاصطناعي.
القيود
يؤكد قسم القيود على ضرورة توثيق مجموعة البيانات بشكل شامل لتحديد التحيزات المحتملة، والأخطاء، والعوامل الأخرى التي قد تؤثر على صلاحية مجموعة البيانات العامة وقابلية تعميمها. يبرز أهمية التعرف على هذه القيود لضمان أن يتمكن مستخدمو تقنيات الصحة المعتمدة على الذكاء الاصطناعي من تقييم مدى ملاءمة البيانات بدقة في سياقات مختلفة (التوصية 1.3a).
علاوة على ذلك، يؤكد على أنه يجب على مستخدمي البيانات الإبلاغ بوضوح عن قيود مجموعات البيانات المستخدمة وآثارها على الاستخدام المقصود لتقنيات الصحة المعتمدة على الذكاء الاصطناعي (التوصية 2.3a). يشمل ذلك التحقيق الشامل فيما إذا كانت القيود تختلف بشكل منهجي عبر مجموعات ديموغرافية مختلفة، وخاصة تلك المصنفة على أنها “غير معروفة”، “يفضل عدم القول”، أو “أخرى”. يعد هذا الإبلاغ أمرًا حيويًا لفهم الفجوات المحتملة في أداء تقنيات الصحة المعتمدة على الذكاء الاصطناعي عبر مجموعات سكانية مختلفة.
DOI: https://doi.org/10.1016/s2589-7500(24)00224-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39701919
Publication Date: 2024-12-18
Author(s): Joseph Alderman et al.
Primary Topic: Artificial Intelligence in Healthcare and Education
Overview
The University of Birmingham emphasizes its commitment to careful curation of available items; however, it acknowledges that errors may occasionally occur in the uploading process. Additionally, certain documents may be classified as commercially sensitive or otherwise restricted. In such instances, the university encourages individuals to report concerns by contacting the designated email address, UBIRA@lists.bham.ac.uk, to facilitate the removal of the problematic document.
Introduction
The introduction of this research paper discusses the rapid growth of artificial intelligence (AI) in healthcare over the past decade, emphasizing its potential to enhance access to specialized diagnoses and treatments while alleviating pressure on limited health resources. However, it underscores the critical need to address the risks associated with algorithmic biases that may exacerbate health inequities. Citing studies such as that by Obermeyer et al., the paper illustrates how biases in AI algorithms can lead to systematic underestimation of healthcare needs for marginalized groups, highlighting the importance of data representation and accuracy in training these algorithms.
The authors identify several mechanisms through which biases can arise, including inadequate representation of minority groups in health datasets, incomplete or inaccurate data for underserved populations, and the encoding of societal inequalities within datasets. They advocate for transparency and self-critical assessment of datasets to better understand their limitations and biases. The introduction also outlines the STANDING Together program, an international initiative aimed at developing recommendations for improving data diversity, inclusivity, and generalizability in AI health technologies. The paper concludes by asserting that the adoption of these recommendations is essential for ensuring that AI health technologies are equitable and effective for all populations.
Methods
The methods employed in the STANDING Together recommendations were developed through a mixed-methods research program spanning from December 2021 to November 2023. An international working group comprising 23 experts from diverse fields—including healthcare, computer science, policy, regulation, and academia—was established to oversee the project. This group included two patient co-investigators to ensure that the recommendations were informed by lived experiences of health inequalities. Additionally, a Patient and Public Involvement and Engagement (PPIE) committee, consisting of 12 members, convened quarterly to provide insights that ensured the language and content of the recommendations were inclusive and respectful. Ethical approval for the study was obtained from the University of Birmingham’s Ethical Review Committee, and preliminary details were made publicly available on the National Institute for Health and Care Research website prior to the project’s initiation.
In the context of artificial intelligence (AI), fairness methods are defined as statistical approaches that assess performance metrics across different demographic groups affected by an algorithm. Various algorithmic techniques are available to develop models that adhere to fairness constraints, aiming to minimize performance disparities among subgroups. One specific recommendation (2.2f) emphasizes the necessity for data users to document any methods employed to intentionally adjust performance across groups. This includes detailing the rationale, goals, methods, metrics, and any variations in thresholds set for different groups throughout the lifecycle of the AI health technology.
Results
In the results section, Recommendation 2.3c emphasizes the importance of reporting findings from preexisting assessments of AI health technologies and the datasets utilized in their development. It suggests that data users should thoroughly review available assessments to understand their implications for various groups within the intended use population. This includes evaluating potential risks of harm associated with the technology, as detailed in the appendix (p. 39). The recommendation underscores the necessity of transparency and due diligence in the evaluation of AI health technologies to ensure safety and efficacy for all user demographics.
Discussion
The STANDING Together program developed a comprehensive set of recommendations aimed at addressing biases and promoting equity in AI health technologies. The recommendations emerged from a systematic review of existing health data standards and a scoping survey of health data experts, which informed an initial longlist of candidate items. This list was refined through a modified Delphi study involving multiple rounds of voting and feedback from a diverse group of stakeholders, including patients and public representatives. The consensus process culminated in a final in-person meeting where 26 participants voted on the recommendations, resulting in 29 finalized items that reflect a broad range of expert opinions and stakeholder insights.
The iterative Delphi process was designed to ensure that the recommendations are culturally relevant and applicable across various contexts. Feedback was solicited at each stage, including public consultations and input from an international advisory group, to enhance the recommendations’ utility. The final recommendations, which will be released in late 2023, aim to mitigate biases in AI health technologies by considering factors such as data origin, missingness, and the intersectionality of health determinants. This approach underscores the importance of inclusive dialogue and consensus-building in developing guidelines that can effectively address disparities in health outcomes associated with AI applications.
Limitations
The section on limitations emphasizes the necessity for comprehensive dataset documentation that identifies potential biases, errors, and other factors that could influence the dataset’s overall validity and generalizability. It highlights the importance of recognizing these limitations to ensure that users of AI health technologies can accurately assess the applicability of the data in various contexts (Recommendation 1.3a).
Furthermore, it stresses that data users must explicitly report the limitations of the datasets employed and their implications for the intended use of AI health technologies (Recommendation 2.3a). This includes a thorough investigation of whether limitations differ systematically across various demographic groups, particularly those categorized as “unknown,” “prefer not to say,” or “other.” Such reporting is crucial to understanding potential disparities in the performance of AI health technologies across different populations.
