DOI: https://doi.org/10.7759/cureus.102219
PMID: https://pubmed.ncbi.nlm.nih.gov/41737097
تاريخ النشر: 2026-01-24
المؤلف: Amol Kodan
الموضوع الرئيسي: أنظمة السجلات الصحية الإلكترونية
نظرة عامة
تقيّم هذه الدراسة فعالية نموذج محادثة دليل مقدمي الخدمات المدعوم بنموذج لغة كبير (LLM) يهدف إلى تحسين التنقل في الرعاية المعتمدة على الحلقات داخل نظام الرعاية الصحية في الولايات المتحدة. يتم انتقاد أدلة مقدمي الخدمات التقليدية بسبب نقص الشفافية، والاعتماد على نماذج بحث صارمة، ووجود معلومات غير دقيقة، مما يعيق اختيار مقدمي الخدمات بشكل مستنير وفعالية الرعاية المعتمدة على القيمة (VBC). قامت الدراسة بتقييم أربعة نماذج LLM – GPT-3.5-turbo وGPT-4o-mini وGPT-4o وGPT-5.1 – باستخدام 87 سيناريو اختبار للغة الطبيعية لقياس مقاييس الأداء المختلفة، بما في ذلك دقة تحديد الحلقات، وموثوقية تصنيف مقدمي الخدمات، وموثوقية الأرقام. ومن الجدير بالذكر أن جميع النماذج حققت دقة عالية في تحديد الحلقات، حيث اقتربت من 91%، على الرغم من وجود تباين واضح في تصنيف مقدمي الخدمات ودقة الأرقام.
تشير النتائج إلى أن أدلة مقدمي الخدمات المدعومة بنموذج LLM يمكن أن تحسن بشكل كبير الشفافية وتجربة المستخدم في إعدادات VBC. من خلال إعطاء الأولوية لتحديد الحلقات بدقة في صيغ تصنيفها، تسلط الدراسة الضوء على قوة حاسمة للنماذج الحديثة التي تم تجاهلها في التقييمات السابقة. مع موثوقية رقمية قوية، ومعدلات هلوسة منخفضة، وسلوك امتناع مناسب، يمكن لهذه الأنظمة تحويل عملية البحث عن مقدمي الخدمات إلى تجربة دعم قرار أكثر توجيهًا وتوافقًا مع القيمة. تشير الإمكانية لمزيد من التقدم في النماذج ودمج بيانات إضافية، مثل توفر مقدمي الخدمات، إلى أن الأدلة المدعومة بنموذج LLM يمكن أن تصبح مكونات أساسية في بنية الرعاية الصحية الشفافة والمركزة على الأعضاء والمعتمدة على القيمة.
مقدمة
تناقش مقدمة هذه الورقة البحثية التحديات المرتبطة بنماذج الدفع المعتمدة على الحلقات في الرعاية الصحية، والتي تتطلب توافقًا واضحًا بين نية المريض، والحلقات المحددة للرعاية، وبيانات أداء مقدمي الخدمات. حاليًا، غالبًا ما يكافح المرضى لتمييز الصلة والقيمة لحلقات الرعاية بسبب المعلومات المجزأة في أدلة مقدمي الخدمات، مما يعيق قدرتهم على التفاعل مع حوافز الرعاية المعتمدة على القيمة (VBC). تعرف الدراسة الرعاية المعتمدة على الحلقات بأنها تجميع لجميع الخدمات المتعلقة بحالة معينة على مدى فترة محددة، وتبرز دور نطاقات المخاطر في تصنيف المرضى بناءً على التعقيد السريري لضمان تفسير عادل للتكاليف.
يقترح المؤلفون أن التقدم الأخير في نماذج اللغة الكبيرة (LLMs) يمكن أن يعزز شفافية أدلة مقدمي الخدمات من خلال تمكين واجهات محادثة تفسر استفسارات المرضى وتحلل البيانات السريرية والمالية المنظمة. يمكن أن يسهل ذلك توافقًا أفضل بين نية المريض والحلقات المناسبة للرعاية، ويحسن اختيار مقدمي الخدمات بناءً على مقاييس مستوى الحلقة. ومع ذلك، تؤكد الورقة على أن تطبيقات الرعاية الصحية تتطلب معايير دقة وسلامة أعلى من الأنظمة المحادثة التقليدية، حيث يمكن أن تؤدي الأخطاء إلى قرارات رعاية غير مناسبة وعواقب مالية. لمعالجة هذه المخاوف، تقيم الدراسة قدرات LLMs باستخدام مجموعات بيانات تركيبية منظمة، مع التركيز على تحديد الحلقات، وتصنيف مقدمي الخدمات، ودقة الأرقام، مما يمهد الطريق لتطبيقات العالم الحقيقي المستقبلية.
الطرق
في هذه الدراسة، تم تقييم أربعة نماذج لغة كبيرة (LLMs) – GPT-3.5-turbo وGPT-4o-mini وGPT-4o وGPT-5.1 – بشكل منهجي تحت ظروف تجريبية محكومة لتقييم أدائها في معالجة حالات اختبار اللغة الطبيعية. تم الوصول إلى النماذج عبر واجهة برمجة تطبيقات موحدة وعملت بإعدادات استدلال حتمية عند درجة حرارة صفر لضمان إمكانية إعادة الإنتاج. عالج كل نموذج نفس مجموعة من 87 حالة اختبار بشكل مستقل، دون الاحتفاظ بذاكرة محادثة، للتركيز على سلوك كل استفسار. تم استخدام موجه نظام ثابت لتقييد النماذج من إنشاء مقدمي خدمات خياليين، أو تكاليف، أو مقاييس أداء، مما يفرض الاعتماد فقط على جداول CSV المقدمة. بالإضافة إلى ذلك، كانت الطلبات للرعاية داخل الشبكة محدودة بمقدمي الخدمات من المستوى 1، وكانت المخرجات مقيدة بتلك التي تحتوي على بيانات أداء على مستوى الحلقة متاحة.
كانت المخرجات مطلوبة للامتثال لمخطط JSON صارم، مدعومًا بشرحات بلغة بسيطة، مما سهل التحقق التلقائي والتقييم بما يتماشى مع معايير نشر الرعاية الصحية. أظهرت النتائج موثوقية عالية في المخرجات المنظمة عبر جميع النماذج، مع معدلات نجاح تحليل JSON تتراوح بين 96.6% و100.0%. ومع ذلك، سلطت الإخفاقات المتبقية الملحوظة في GPT-4o وGPT-5.1 الضوء على ضرورة وجود آليات قوية للاستخراج والتحقق والاحتياطي لضمان السلامة في التطبيقات الحرجة، حتى عند العمل تحت ظروف حتمية.
النتائج
قيمت نتائج الدراسة 87 حالة اختبار للغة الطبيعية عبر نماذج مختلفة، مع التركيز على الشفافية المعتمدة على الحلقة. تم حساب مقاييس الملخص مع فترات ثقة بنسبة 95% للنسب وفترات ثقة bootstrap للمتوسطات، مع بيانات مصنفة حسب اللغة (الإنجليزية: ن = 71؛ غير الإنجليزية: ن = 16). أظهرت جميع النماذج صلاحية عالية للمخرجات المنظمة، مع معدلات نجاح تحليل JSON تتراوح بين 96.6% و100.0%. ومع ذلك، سلطت الإخفاقات المتبقية في النماذج GPT-4o وGPT-5.1 الضوء على ضرورة وجود تحقق قوي من المخطط وآليات احتياطي آمنة في بيئات الإنتاج.
كانت دقة تحديد الحلقات ملحوظة للغاية، مما يشير إلى أن النماذج الكبيرة الحديثة (LLMs) يمكن أن تترجم بفعالية نية الأعضاء النصية الحرة إلى هياكل الحلقات، وهو أمر ضروري لمقارنات التكاليف المعتمدة على الحلقات في إعدادات الرعاية المعتمدة على القيمة (VBC). من ناحية أخرى، أظهر تحديد نطاق المخاطر تباينًا أكبر، مما يشير إلى أنه بينما يؤثر تصنيف المخاطر على تفسير التكاليف، يمكن التخفيف من الأخطاء من خلال مطالبات التوضيح أو المنطق الحتمي. تؤكد الدراسة على أن دقة الحلقة أمر حاسم لضمان شفافية ذات مغزى في تفاعلات الدليل، حيث تحدد ما إذا كان النموذج يحدد بشكل صحيح الحلقة المناسبة للرعاية. تم تقديم مقياس دقة تصنيف الحلقة المعتمد على الوزن، حيث تعمل صحة الحلقة كشرط بوابة ثنائي، ومعامل قابل للتعديل $\alpha$ يضبط التركيز على صحة التصنيف الأعلى مقابل جودة التصنيف العامة، مما يضمن عدم منح أي ائتمان للتصنيفات عندما تكون الحلقة غير صحيحة. تتماشى هذه الطريقة مع أهداف الشفافية وحماية الأعضاء.
المناقشة
تؤكد قسم المناقشة في الورقة البحثية على الدور الحاسم لدقة الحلقة في تعزيز الشفافية وتجربة الأعضاء ضمن أطر الرعاية المعتمدة على القيمة المعتمدة على الحلقة (VBC). أظهرت النماذج التي تم تقييمها دقة عالية في تحديد الحلقات، مع معدلات تقترب من 91%، وهو أمر ضروري لضمان حصول الأعضاء على معلومات سريرية ومالية ذات صلة قبل إجراء مقارنات التكاليف أو اختيار مقدمي الخدمات. تجادل الدراسة من أجل إعادة صياغة مقاييس تقييم التصنيف لإعطاء الأولوية لدقة الحلقة، مؤكدة أن التصنيفات غير المثالية المستندة إلى حلقات صحيحة يمكن أن تسهل اتخاذ قرارات مستنيرة، بينما يمكن أن تؤدي التصنيفات المستندة إلى حلقات غير صحيحة إلى تضليل الأعضاء.
تسلط النتائج الضوء أيضًا على أهمية تحقيق التوازن بين مقاييس أداء النموذج المختلفة، مثل الموثوقية الرقمية ومعدلات الهلوسة، لضمان ملاءمتها للتطبيقات الموجهة للأعضاء. أظهر النموذج الأقوى دقة عالية في الحلقة إلى جانب معدلات هلوسة منخفضة وسلوك امتناع قوي، مما يبرز أن الشفافية هي خاصية ناشئة من عوامل متعددة بدلاً من مقياس واحد. علاوة على ذلك، يتم تقديم الانتقال من أدلة مقدمي الخدمات التقليدية إلى واجهات المحادثة المدعومة بنموذج LLM كتحسين كبير، مما يقلل من العبء المعرفي ويعزز تفاعل الأعضاء مع إشارات التكلفة والجودة المتأصلة في نماذج VBC.
تختتم الورقة بالدعوة إلى نهج منظم للحوكمة والسلامة في نشر LLMs، مشددة على أن هذه النماذج يجب ألا تعمل كصانعي قرارات مستقلين. بدلاً من ذلك، يجب أن تتولى الأنظمة الحتمية الوظائف الحرجة مثل تصفية الشبكة والتحقق الرقمي. يقترح المؤلفون أن دمج بيانات إضافية، مثل توفر مقدمي الخدمات، واستمرار التقدم في تدريب النماذج يمكن أن يعزز المزيد من الشفافية والفعالية لأدلة مقدمي الخدمات المدعومة بنموذج LLM، مما يدعم في النهاية نهجًا أكثر تركيزًا على الأعضاء في الرعاية الصحية.
القيود
تسلط قيود هذه الدراسة الضوء على عدة عوامل حاسمة قد تؤثر على قابلية تعميم نتائجها. أولاً، قد لا تعكس استخدام البيانات التركيبية، على الرغم من فائدتها للتقييم المنضبط، التعقيدات والتباينات الموجودة في بيانات المطالبات الواقعية بشكل كامل، خاصة من حيث التباين، والضوضاء، والافتقار. بالإضافة إلى ذلك، كانت التحليل محدودًا بالتفاعلات ذات الدور الواحد، مما قد يغفل مزايا تدفقات التوضيح متعددة الأدوار التي يمكن أن تعزز أداء النموذج. تم إجراء التقييم تحت إعدادات استدلال حتمية، والتي، على الرغم من تحسين إمكانية إعادة الإنتاج، قد لا تعكس بدقة التباين الذي يتم مواجهته في التطبيقات العملية.
علاوة على ذلك، قد لا يتماشى تقييم تصنيف نطاق المخاطر كمهام مطابقة دقيقة مع التفسيرات الاحتمالية أو الترتيبية التي تتطلبها غالبًا تصنيفات المخاطر السريرية. كما لاحظت الدراسة أن مقاييس الأداء، بما في ذلك زمن الاستجابة واستخدام الرموز، كانت تستند إلى مطالبات مع سياق جدولي كامل، مما قد يختلف في سيناريوهات الإنتاج التي تستخدم الاسترجاع المستهدف. الأخطاء التي تم تحديدها في الدراسة، مثل قيم تكاليف الحلقات المحتملة ولكن غير الصحيحة، تسلط الضوء على الفشل الهيكلي بدلاً من القضايا المحادثة. أخيرًا، يحد الحجم الصغير لمجموعة غير الإنجليزية من الاستنتاجات حول قوة تعدد اللغات. بشكل عام، تؤكد النتائج على الحاجة إلى التحقق من البيانات الواقعية للمطالبات وتقترح أنه بينما أظهر النموذج الأقوى أداءً عاليًا في الموثوقية الرقمية ومعدلات هلوسة منخفضة، فإن المزيد من البحث ضروري لتقييم النتائج السريرية اللاحقة وتأثيرها على تفاعل الأعضاء واختيار الرعاية الفعالة من حيث التكلفة.
DOI: https://doi.org/10.7759/cureus.102219
PMID: https://pubmed.ncbi.nlm.nih.gov/41737097
Publication Date: 2026-01-24
Author(s): Amol Kodan
Primary Topic: Electronic Health Records Systems
Overview
The research evaluates the effectiveness of a large language model (LLM)-driven provider directory chatbot aimed at enhancing episode-based care navigation within the U.S. healthcare system. Conventional provider directories are criticized for their lack of transparency, reliance on rigid search paradigms, and the presence of inaccurate information, which impede informed provider selection and the efficacy of value-based care (VBC). The study assessed four LLMs—GPT-3.5-turbo, GPT-4o-mini, GPT-4o, and GPT-5.1—using 87 natural language test scenarios to measure various performance metrics, including episode identification accuracy, provider ranking reliability, and numeric fidelity. Notably, all models achieved high episode identification accuracy, nearing 91%, although variability was evident in provider ranking and numeric precision.
The findings suggest that LLM-driven provider directories can significantly improve transparency and user experience in VBC settings. By prioritizing accurate episode identification in their ranking formulations, the study underscores a critical strength of contemporary LLMs that has been overlooked in previous assessments. With strong numeric fidelity, low hallucination rates, and appropriate abstention behavior, these systems can transform the provider search process into a more guided and value-aligned decision support experience. The potential for further model advancements and the integration of additional data, such as provider availability, indicates that LLM-driven directories could become essential components of a transparent, member-centered, and value-based healthcare infrastructure.
Introduction
The introduction of this research paper discusses the challenges associated with episode-based payment models in healthcare, which require a clear alignment between patient intent, defined episodes of care, and provider performance data. Currently, patients often struggle to discern the relevance and value of care episodes due to fragmented information in provider directories, which hinders their ability to engage with value-based care (VBC) incentives. The study defines episode-based care as the aggregation of all services related to a specific condition over a designated period, and it highlights the role of risk bands in categorizing patients based on clinical complexity to ensure fair cost interpretation.
The authors propose that recent advancements in large language models (LLMs) could enhance the transparency of provider directories by enabling conversational interfaces that interpret patient queries and analyze structured clinical and financial data. This could facilitate better alignment of patient intent with appropriate episodes of care and improve provider selection based on episode-level metrics. However, the paper emphasizes that healthcare applications demand higher accuracy and safety standards than typical conversational systems, as errors could lead to inappropriate care decisions and financial repercussions. To address these concerns, the study evaluates the capabilities of LLMs using structured synthetic datasets, focusing on episode identification, provider ranking, and numeric accuracy, thereby laying the groundwork for future real-world applications.
Methods
In this study, four large language models (LLMs)—GPT-3.5-turbo, GPT-4o-mini, GPT-4o, and GPT-5.1—were systematically evaluated under controlled experimental conditions to assess their performance in processing natural language test cases. The models were accessed via a unified API and operated with deterministic inference settings at zero temperature to ensure reproducibility. Each model processed the same set of 87 test cases independently, without retaining conversational memory, to focus on per-query behavior. A fixed system prompt was employed to restrict the models from generating fictitious providers, costs, or performance metrics, mandating reliance solely on provided CSV tables. Additionally, requests for in-network care were limited to Tier 1 providers, and outputs were constrained to those with available episode-level performance data.
The outputs were required to adhere to a strict JSON schema, supplemented by plain language explanations, which facilitated automated validation and scoring in line with healthcare deployment standards. The results indicated high reliability in structured output across all models, with JSON parse success rates ranging from 96.6% to 100.0%. However, residual failures noted in GPT-4o and GPT-5.1 highlighted the necessity for robust extraction, validation, and fallback mechanisms to ensure safety in critical applications, even when operating under deterministic conditions.
Results
The results of the study evaluated 87 natural-language test cases across various models, focusing on episode-weighted transparency. Summary metrics were calculated with 95% confidence intervals for proportions and bootstrap confidence intervals for means, with data stratified by language (English: n = 71; non-English: n = 16). All models exhibited high structured output validity, with JSON parse success rates between 96.6% and 100.0%. However, residual failures in models GPT-4o and GPT-5.1 highlighted the necessity for robust schema validation and safe fallback mechanisms in production environments.
The accuracy of episode identification was notably high, indicating that contemporary large language models (LLMs) can effectively translate free-text member intent into episode constructs, which is essential for episode-based cost comparisons in value-based care (VBC) settings. Conversely, risk-band identification displayed greater variability, suggesting that while risk stratification affects cost interpretation, errors can be mitigated through clarification prompts or deterministic logic. The study emphasizes that episode accuracy is crucial for ensuring meaningful transparency in directory interactions, as it determines whether the model correctly identifies the appropriate episode of care. The Episode Weighted Ranking Accuracy metric was introduced, where episode correctness serves as a binary gating condition, and a tunable parameter $\alpha$ adjusts the emphasis on exact top-ranked correctness versus overall ranking quality, ensuring that no credit is given for rankings when the episode is incorrect. This approach aligns the metric with transparency and member protection objectives.
Discussion
The discussion section of the research paper emphasizes the critical role of episode accuracy in enhancing transparency and member experience within episode-based value-based care (VBC) frameworks. The evaluated models demonstrated high episode identification accuracy, with rates nearing 91%, which is essential for ensuring that members receive relevant clinical and financial information before making cost comparisons or provider selections. The study argues for a reframing of ranking evaluation metrics to prioritize episode accuracy, asserting that even imperfect rankings grounded in correct episodes can facilitate informed decision-making, while rankings based on incorrect episodes can mislead members.
The findings also highlight the importance of balancing various model performance metrics, such as numeric fidelity and hallucination rates, to ensure suitability for member-facing applications. The strongest model exhibited high episode accuracy alongside low hallucination rates and strong abstention behavior, underscoring that transparency is an emergent property of multiple factors rather than a single metric. Furthermore, the transition from traditional provider directories to LLM-driven conversational interfaces is presented as a significant improvement, reducing cognitive burden and enhancing member engagement with cost and quality signals inherent in VBC models.
The paper concludes by advocating for a structured approach to governance and safety in deploying LLMs, emphasizing that these models should not operate as standalone decision-makers. Instead, deterministic systems should handle critical functions like network filtering and numeric validation. The authors suggest that integrating additional data, such as provider availability, and continuing advancements in model training could further enhance the transparency and effectiveness of LLM-driven provider directories, ultimately supporting a more member-centered approach to healthcare.
Limitations
The limitations of this study highlight several critical factors that may affect the generalizability of its findings. Firstly, the use of synthetic data, while beneficial for controlled evaluation, may not fully encapsulate the complexities and variabilities present in real-world claims data, particularly in terms of heterogeneity, noise, and missingness. Additionally, the analysis was limited to single-turn interactions, potentially overlooking the advantages of multi-turn clarification flows that could enhance model performance. The evaluation was conducted under deterministic inference settings, which, while improving reproducibility, may not accurately reflect the variability encountered in practical applications.
Moreover, the assessment of risk band classification as an exact match task may not align with the probabilistic or ordinal interpretations that clinical risk stratification often requires. The study also noted that the performance metrics, including latency and token usage, were based on prompts with full tabular context, which may differ in production scenarios utilizing targeted retrieval. Errors identified in the study, such as plausible but incorrect episode cost values, underscore structural failures rather than conversational issues. Lastly, the small size of the non-English subset limits conclusions about multilingual robustness. Overall, the findings emphasize the need for validation with real-world claims data and suggest that while the strongest performing model demonstrated high numeric fidelity and low hallucination rates, further research is necessary to assess downstream clinical outcomes and the impact on member engagement and cost-effective care selection.
