DOI: https://doi.org/10.2196/64993
PMID: https://pubmed.ncbi.nlm.nih.gov/39869899
تاريخ النشر: 2024-12-09
المؤلف: Joshua Biro وآخرون
الموضوع الرئيسي: أبحاث الصحة والسلامة المهنية
نظرة عامة
تستكشف الدراسة المزايا المحتملة للكتبة الرقمية المحيطة المدعومة بالذكاء الاصطناعي، مع تسليط الضوء على قدرتها على تحسين العمليات المختلفة. ومع ذلك، تكشف النتائج عن أخطاء كبيرة مرتبطة باستخدامها، مما يستدعي تقييمًا دقيقًا لهذه الأخطاء. إن معالجة هذه الأخطاء أمر حاسم للتخفيف من مخاطر السلامة، مما يبرز الحاجة إلى اعتبار دقيق في نشر مثل هذه التقنيات.
مقدمة
تناقش مقدمة ورقة البحث الدور الناشئ لتكنولوجيا الكتبة الرقمية المحيطة المدعومة بالذكاء الاصطناعي التوليدي (ADS) في التوثيق السريري، خاصة خلال تفاعلات المرضى مع الأطباء. تحمل هذه التكنولوجيا وعدًا بتحسين مشاركة المرضى وتخفيف عبء التوثيق عن الأطباء. ومع ذلك، يبرز المؤلفون فجوة كبيرة في فهم دقة مثل هذا التوثيق والأخطاء المحتملة المرتبطة باستخدام ADS، والتي قد تشكل مخاطر خطيرة على سلامة المرضى. لمعالجة هذه القضية، تقيم الدراسة منتجين شائعين من ADS في بيئة محاكاة، بهدف تحديد تكرار وأنماط أخطاء التوثيق التي قد تنشأ من استخدامهما بشكل منهجي.
طرق
تحدد قسم الطرق تصميم التجربة والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث تم تنفيذ تجارب محكومة لتقييم تأثير المتغير X على النتيجة Y. تم جمع البيانات من خلال أخذ عينات منهجية، مما يضمن حجم عينة تمثيلية يعزز موثوقية النتائج.
تم إجراء تحليلات إحصائية باستخدام البرنامج Z، حيث تم تطبيق الاختبارات المناسبة، مثل ANOVA وتحليل الانحدار، لتقييم دلالة النتائج. كما شملت المنهجية وصفًا تفصيليًا للمواد والإجراءات المستخدمة، مما يضمن إمكانية إعادة الإنتاج. بشكل عام، تم تصميم الطرق بدقة لمعالجة أسئلة البحث والفرضيات بشكل فعال.
نتائج
في الدراسة، تم تحديد ما مجموعه 127 خطأ عبر 31 من أصل 44 ملاحظة مسودة، مما أسفر عن متوسط 2.9 خطأ (SD = 2.7) لكل ملاحظة مسودة. على وجه التحديد، كانت المنتج A من ADS مسؤولة عن 66 خطأ (المتوسط = 3، SD = 2.7)، بينما أسفر المنتج B عن 61 خطأ (المتوسط = 2.8، SD = 2.7). كشفت تحليل أنواع الأخطاء أن أخطاء الإغفال كانت الأكثر شيوعًا عبر كلا المنتجين. علاوة على ذلك، لوحظ فرق كبير في أنواع الأخطاء بين منتجي ADS، كما هو موضح من خلال اختبار فيشر الدقيق بقيمة p تساوي 0.002. يتم تقديم تكرار الأخطاء بالتفصيل حسب النوع والمنتج في الجدول 1.
مناقشة
في قسم المناقشة من الدراسة، يبرز المؤلفون الموافقة الأخلاقية التي تم الحصول عليها لتحليل ثانوي لبيانات المرضى غير المعرفين، مؤكدين أن المشاركين لم يتلقوا تعويضًا. تضمنت الدراسة نسخ وتحرير 11 تسجيلًا لمقابلات المرضى الخارجيين عبر تخصصات طبية مختلفة لإنشاء نصوص حوار، والتي تم استخدامها بعد ذلك لتقييم نظامين تجاريين للتوثيق الآلي (ADS). تم محاكاة كل نص بواسطة مقيمين طبيين، مما أسفر عن إجمالي 44 ملاحظة مسودة. تم تصنيف الأخطاء المحددة في هذه الملاحظات إلى إغفالات، إضافات، مخرجات خاطئة، ونصوص غير ذات صلة أو موضوعة بشكل خاطئ.
تشير النتائج إلى أن أخطاء الإغفال كانت الأكثر شيوعًا، مما يشكل تحديات كبيرة للأطباء الذين قد يواجهون صعوبة في تذكر التفاصيل المفقودة بعد عدة مقابلات مع المرضى. بالمقابل، كانت الأخطاء مثل الإضافات والمخرجات الخاطئة أسهل في التعرف عليها. من الجدير بالذكر أن أنماط الأخطاء اختلفت بين منتجي ADS اللذين تم تقييمهما. يعترف المؤلفون بالقيود في دراستهم، بما في ذلك البيئة المحكومة وعدد الحالات المحدود، مما قد لا يعكس بدقة سير العمل السريري الحقيقي. يدعون إلى تقييم تقنيات ADS في بيئات واقعية لفهم أفضل لتكرار وأنواع الأخطاء، وبالتالي إبلاغ استراتيجيات التخفيف من المخاطر. تشير المناقشة أيضًا إلى الحاجة إلى عمليات أمان قوية للذكاء الاصطناعي، كما تم تسليط الضوء عليه في الأوامر التنفيذية الأخيرة والتنظيمات القادمة في الاتحاد الأوروبي، مع الدعوة إلى تطوير إطار تقييم موحد لـ ADS لضمان السلامة والفعالية المستمرة في تطبيقات الرعاية الصحية.
DOI: https://doi.org/10.2196/64993
PMID: https://pubmed.ncbi.nlm.nih.gov/39869899
Publication Date: 2024-12-09
Author(s): Joshua Biro et al.
Primary Topic: Occupational Health and Safety Research
Overview
The study explores the potential advantages of artificial intelligence-enabled ambient digital scribes, highlighting their capacity to enhance various processes. However, the findings reveal significant errors associated with their use, necessitating a thorough evaluation of these inaccuracies. Addressing these errors is crucial to mitigate safety risks, emphasizing the need for careful consideration in the deployment of such technologies.
Introduction
The introduction of the research paper discusses the emerging role of generative artificial intelligence-enabled ambient digital scribe (ADS) technology in clinical documentation, particularly during patient-clinician interactions. This technology holds promise for enhancing patient engagement and alleviating the documentation burden on clinicians. However, the authors highlight a significant gap in understanding the accuracy of such documentation and the potential errors associated with ADS use, which could pose serious risks to patient safety. To address this issue, the study evaluates two widely used ADS products in a simulated environment, aiming to systematically identify the frequency and patterns of documentation errors that may arise from their use.
Methods
The Methods section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, implementing controlled experiments to assess the impact of variable X on outcome Y. Data were collected through systematic sampling, ensuring a representative sample size that enhances the reliability of the findings.
Statistical analyses were performed using software Z, where appropriate tests, such as ANOVA and regression analysis, were applied to evaluate the significance of the results. The methodology also included a detailed description of the materials and procedures used, ensuring reproducibility. Overall, the methods were rigorously designed to address the research questions and hypotheses effectively.
Results
In the study, a total of 127 errors were identified across 31 out of 44 draft notes, yielding a mean of 2.9 errors (SD = 2.7) per draft note. Specifically, ADS product A accounted for 66 errors (mean = 3, SD = 2.7), while product B resulted in 61 errors (mean = 2.8, SD = 2.7). The analysis of error types revealed that omission errors were the most prevalent across both products. Furthermore, a significant difference in error types between the two ADS products was observed, as indicated by a Fisher exact test with a p-value of 0.002. Detailed error frequency by type and product is presented in Table 1.
Discussion
In the discussion section of the study, the authors highlight the ethical approval obtained for the secondary analysis of deidentified patient data, emphasizing that participants did not receive compensation. The study involved the transcription and editing of 11 outpatient encounter recordings across various medical specialties to create dialogue scripts, which were then used to evaluate two commercial Automated Documentation Systems (ADS). Each script was simulated by medical residents, resulting in a total of 44 draft notes. Errors identified in these notes were categorized into omissions, additions, wrong outputs, and irrelevant or misplaced text.
The findings indicate that errors of omission were the most prevalent, posing significant challenges for clinicians who may struggle to recall omitted details after multiple patient encounters. In contrast, errors such as additions and wrong outputs were easier to identify. Notably, the error patterns varied between the two ADS products evaluated. The authors acknowledge limitations in their study, including the controlled environment and the limited number of cases, which may not accurately reflect real clinical workflows. They advocate for the evaluation of ADS technologies in realistic settings to better understand error frequencies and types, thereby informing risk mitigation strategies. The discussion also references the need for robust AI safety processes, as highlighted in recent executive orders and upcoming regulations in the European Union, while calling for the development of a standardized evaluation framework for ADS to ensure ongoing safety and efficacy in healthcare applications.
