DOI: https://doi.org/10.2196/65848
PMID: https://pubmed.ncbi.nlm.nih.gov/40493514
تاريخ النشر: 2025-03-21
المؤلف: Emma Coen وآخرون
الموضوع الرئيسي: التدخلات الرقمية في الصحة النفسية
نظرة عامة
تسلط الأبحاث الضوء على الحاجة الملحة لنماذج تقديم خدمات مبتكرة في الفحص الجيني بسبب الطلب المتزايد على الاختبارات الجينية على مستوى السكان وموارد الاستشارة الجينية المحدودة. تركز الدراسة على تطوير وتقييم روبوت محادثة مدعوم بنموذج اللغة الكبير GPT-4 لمساعدة في إرجاع نتائج الفحص الجيني الإيجابية، تحديدًا ضمن برنامج “في حمضنا النووي SC” في ولاية كارولينا الجنوبية، الذي يقدم فحصًا جينيًا مجانيًا لعدد كبير من السكان. تم تصميم روبوت المحادثة باستخدام عملية هندسة موجهة من ثلاث خطوات تضمنت توليد معزز بالاسترجاع وتقنيات قليلة اللقطات، مستخدمًا مواد تدريبية مثل الأسئلة الشائعة للمرضى ونصوص الاستشارة الجينية.
شمل تقييم أداء روبوت المحادثة تقييمات الخبراء بناءً على سيناريوهين افتراضيين للمرضى، مع تقييم معايير مثل النغمة والوضوح والدقة. حصل روبوت المحادثة على متوسط درجة 3.86 (SD 0.89)، مع تقييمات مرتفعة بشكل خاص للنغمة وسهولة الاستخدام (متوسط 4.25، SD 0.71). بينما تمكن روبوت المحادثة من إدارة الاستفسارات المفتوحة والحفاظ على حدود المحادثة، أظهر حاجة للتحسين في دقة البرنامج المحددة (متوسط 3.25، SD 1.39). تشير النتائج إلى أن روبوتات المحادثة المدعومة بنموذج اللغة الكبير يمكن أن تدعم بشكل معقول إرجاع النتائج الجينية الإيجابية، مع فوائد محتملة لكفاءة الرعاية الصحية وتعليم المرضى. ستركز الأعمال المستقبلية على تطوير نموذج روبوت محادثة هجين يجمع بين مكونات نموذج اللغة الكبير والقواعد لتحسين تقديم المعلومات الجينية ومعالجة قيود موارد الاستشارة الجينية الحالية.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الطلب المتزايد على الاختبارات الجينية والتحديات التي تفرضها محدودية الوصول إلى الخبرة الجينية. لمعالجة هذه القضايا، ظهرت نماذج تقديم خدمات جينية مبتكرة، وخاصة دمج روبوتات المحادثة، كأدوات فعالة لتعزيز التواصل طوال عملية الخدمة الجينية. توفر روبوتات المحادثة منصة قابلة للتوسع وسهلة الوصول لتقديم التعليم الصحي الشخصي، داعمة مهام مثل الاستشارة قبل الاختبار، والموافقة المستنيرة، والتواصل بشأن نتائج الاختبارات الجينية. أظهرت الدراسات السابقة، بما في ذلك تجربة BRIDGE، أن الأساليب المعتمدة على روبوتات المحادثة يمكن أن تكون فعالة مثل الرعاية القياسية في إكمال التعليم الجيني قبل الاختبار والاختبار.
على الرغم من مزايا روبوتات المحادثة، لا يزال تطبيقها في تقديم نتائج الاختبارات الجينية الإيجابية غير مستكشف بشكل كافٍ. تشمل الممارسات الحالية بشكل أساسي التواصل المباشر بسبب الطبيعة الحساسة للمعلومات. ومع ذلك، تشير البيانات النوعية إلى أن المرضى منفتحون على تلقي النتائج عبر روبوتات المحادثة، التي تقدم الراحة وفرصة للتفكير والاستفسار. تقترح الورقة تطوير روبوت محادثة هجين يجمع بين مكونات قائمة على القواعد ومكونات نموذج اللغة الكبير (LLM) لتسهيل التفاعلات الدقيقة بشأن نتائج الفحص الجيني الإيجابية على مستوى السكان. أهداف المشروع هي هندسة روبوت محادثة قائم على LLM لهذا الغرض وتقييم فعاليته من خلال دراسات حالة افتراضية وتقييمات الخبراء. تهدف هذه الأعمال إلى تعزيز مشاركة المرضى وفهمهم في عملية تشخيص السرطان الوراثي.
طرق
توضح قسم “الطرق” تصميم التجربة والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث تم استخدام التحليلات الإحصائية لتقييم البيانات التي تم جمعها من تجارب مختلفة. تضمنت المنهجيات المحددة تجارب مختبرية محكومة، حيث تم التلاعب بالمتغيرات بشكل منهجي لمراقبة تأثيراتها على النتائج المعنية.
شمل جمع البيانات استخدام أدوات معيارية لضمان الموثوقية والصلاحية. تم إجراء التحليل باستخدام أدوات برمجية تسهل النمذجة الإحصائية المعقدة، مما يسمح بتقييم العلاقات بين المتغيرات. تم اشتقاق النتائج الرئيسية من التحليلات الانحدارية، التي قدمت رؤى حول القوة التنبؤية للمتغيرات المستقلة على النتائج التابعة. بشكل عام، كانت الطرق المستخدمة صارمة وتهدف إلى ضمان قوة النتائج.
نتائج
في هذا القسم، يقدم المؤلفون نتائج تقييمهم لروبوت محادثة مصمم للمساعدة في تقديم نتائج إيجابية من الفحص الجيني. باستخدام تقنية توليد معزز بالاسترجاع (RAG)، قاموا بتصميم موجه محدد لروبوت المحادثة، الذي تم تقييمه بعد ذلك من قبل ثمانية خبراء. شمل التقييم الداخلي استجابات روبوت المحادثة لـ 14 سؤالًا عبر ثمانية مجالات ضمن سيناريوهين افتراضيين. حصل روبوت المحادثة على متوسط درجة إجمالية قدرها 3.88، مع أعلى التقييمات في مجال النغمة وأدنىها في دقة البرنامج، مما يشير إلى مجالات للتحسين.
تسلط النتائج الضوء على إمكانية تحسين الموجه بشكل أكبر ودمج نموذج اللغة الكبير (LLM) لروبوت المحادثة مع نظام قائم على القواعد موجود. يهدف هذا الدمج إلى إنشاء روبوت محادثة هجين قادر على دعم إرجاع نتائج الفحص الجيني بشكل فعال. من الجدير بالذكر أن التباين في الدرجات بين المقيمين يشير إلى آراء مختلفة حول أداء روبوت المحادثة، مما قد يوجه التحسينات المستقبلية.
مناقشة
في هذا القسم، يناقش المؤلفون تطوير وتقييم روبوت محادثة مصمم لمساعدة الأفراد في تلقي نتائج من برنامج الفحص الجيني السكاني (PGS) في جامعة الطب في ولاية كارولينا الجنوبية (MUSC). يهدف برنامج PGS، الذي بدأ في نوفمبر 2021، إلى تقديم الفحص الجيني لـ 100,000 فرد، مع تحقيق تقدم كبير من خلال تجنيد أكثر من 59,000 مشارك وتحديد حالات جينية متنوعة. تم تدريب روبوت المحادثة باستخدام نهج هندسة موجهة يتضمن تقنيات توليد معزز بالاسترجاع (RAG)، مما يسمح له بتوليد استجابات مستندة إلى السياق من خلال الوصول إلى مواد إضافية في الوقت الحقيقي. شمل تقييم أداء روبوت المحادثة تقييم خبراء المجال لاستجاباته لسيناريوهات مرضى افتراضية، مع التركيز على معايير مثل الدقة والوضوح وسهولة الاستخدام.
تشير النتائج إلى أن روبوت المحادثة قدم أداءً جيدًا من حيث الحدود، ودقة المجال، والصلابة، مع متوسط تقييم للخبراء قدره 3.86 عبر معايير مختلفة. ومع ذلك، لا تزال هناك تحديات في ضمان دقة البرنامج العالية، حيث واجه روبوت المحادثة صعوبة في تقديم معلومات محددة حول برنامج PGS نفسه. يؤكد المؤلفون على أهمية التحسين المستمر، بما في ذلك دمج ملاحظات المرضى في التقييمات المستقبلية، لتعزيز فعالية روبوت المحادثة في تقديم المعلومات الجينية. يقترحون نهجًا هجينًا يجمع بين المحتوى المكتوب وقدرات نموذج اللغة الكبير لتحسين كل من دقة المعلومات وجودة المحادثات، بهدف تسهيل مشاركة أفضل للمرضى وفهم نتائج الاختبارات الجينية.
DOI: https://doi.org/10.2196/65848
PMID: https://pubmed.ncbi.nlm.nih.gov/40493514
Publication Date: 2025-03-21
Author(s): Emma Coen et al.
Primary Topic: Digital Mental Health Interventions
Overview
The research highlights the urgent need for innovative service delivery models in genomic screening due to the rising demand for population-wide genomic testing and limited genetic counseling resources. The study focuses on the development and evaluation of a chatbot powered by the GPT-4 large language model (LLM) to assist in returning positive genomic screening results, specifically within South Carolina’s In Our DNA SC program, which offers free genetic screening to a large population. The chatbot was designed using a three-step prompt engineering process that incorporated retrieval-augmented generation and few-shot techniques, utilizing training materials such as patient FAQs and genetic counseling scripts.
Evaluation of the chatbot’s performance involved expert ratings based on two hypothetical patient scenarios, assessing criteria such as tone, clarity, and accuracy. The chatbot received an average score of 3.86 (SD 0.89), with particularly high ratings for tone and usability (mean 4.25, SD 0.71). While the chatbot effectively managed open-ended queries and maintained conversational boundaries, it showed a need for improvement in program-specific accuracy (mean 3.25, SD 1.39). The findings suggest that LLM-powered chatbots can feasibly support the return of positive genomic results, with potential benefits for healthcare efficiency and patient education. Future work will focus on developing a hybrid chatbot model that combines LLM and rule-based components to enhance the delivery of genomic information and address the limitations of current genetic counseling resources.
Introduction
The introduction of this research paper highlights the growing demand for genomic testing and the challenges posed by limited access to genomic expertise. To address these issues, innovative genetic service delivery models, particularly the integration of chatbots, have emerged as effective tools for enhancing communication throughout the genomic service process. Chatbots provide a scalable and accessible platform for delivering personalized health education, supporting tasks such as pretest counseling, informed consent, and the communication of genetic test results. Previous studies, including the BRIDGE trial, have demonstrated that chatbot-based approaches can be as effective as standard care in completing pretest genetic education and testing.
Despite the advantages of chatbots, their application in delivering positive genetic test results remains underexplored. Current practices primarily involve direct communication due to the sensitive nature of the information. However, qualitative data suggest that patients are open to receiving results via chatbots, which offer convenience and the opportunity for reflection and inquiry. The paper proposes the development of a hybrid chatbot that combines rule-based and large language model (LLM) components to facilitate nuanced interactions regarding positive population-wide genomic screening (PGS) results. The objectives of the project are to engineer an LLM-based chatbot for this purpose and to evaluate its effectiveness through hypothetical case studies and expert assessments. This work aims to enhance patient engagement and understanding in the hereditary cancer diagnostic process.
Methods
The “Methods” section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, employing statistical analyses to evaluate the data collected from various experiments. Specific methodologies included controlled laboratory experiments, where variables were systematically manipulated to observe their effects on the outcomes of interest.
Data collection involved the use of standardized instruments to ensure reliability and validity. The analysis was conducted using software tools that facilitated complex statistical modeling, allowing for the assessment of relationships between variables. Key findings were derived from regression analyses, which provided insights into the predictive power of the independent variables on the dependent outcomes. Overall, the methods employed were rigorous and aimed at ensuring the robustness of the findings.
Results
In this section, the authors present the results of their evaluation of a chatbot designed to assist in delivering positive results from genomic screening. Utilizing the Retrieval-Augmented Generation (RAG) technique, they engineered a specific prompt for the chatbot, which was subsequently assessed by eight experts. The intrinsic evaluation involved the chatbot’s responses to 14 questions across eight domains within two hypothetical case scenarios. The chatbot received an overall average score of 3.88, with the highest ratings in the tone domain and the lowest in program accuracy, indicating areas for improvement.
The findings highlight the potential for further refinement of the prompt and the integration of the chatbot’s large language model (LLM) with an existing rule-based system. This integration aims to create a hybrid chatbot capable of effectively supporting the return of genomic screening results. Notably, the variability in scores among raters suggests differing opinions on the chatbot’s performance, which may guide future enhancements.
Discussion
In this section, the authors discuss the development and evaluation of a chatbot designed to assist individuals receiving results from a population genomic screening (PGS) program at the Medical University of South Carolina (MUSC). The PGS program, initiated in November 2021, aims to provide genetic screening to 100,000 individuals, with significant progress made by recruiting over 59,000 participants and identifying various genetic conditions. The chatbot was trained using a prompt engineering approach that incorporated retrieval-augmented generation (RAG) techniques, allowing it to generate contextually informed responses by accessing supplementary materials in real-time. The evaluation of the chatbot’s performance involved domain experts assessing its responses to hypothetical patient scenarios, focusing on criteria such as accuracy, clarity, and usability.
The findings indicate that the chatbot performed well in terms of boundaries, domain accuracy, and robustness, with an average expert rating of 3.86 across various criteria. However, challenges remain in ensuring high program accuracy, as the chatbot struggled to provide specific information about the PGS program itself. The authors emphasize the importance of ongoing refinement, including incorporating patient feedback in future evaluations, to enhance the chatbot’s effectiveness in delivering genetic information. They propose a hybrid approach that combines scripted content with the LLM capabilities to improve both the accuracy of information and the conversational quality of interactions, ultimately aiming to facilitate better patient engagement and understanding of genetic testing results.
