DOI: https://doi.org/10.1038/s41746-026-02537-1
PMID: https://pubmed.ncbi.nlm.nih.gov/41832240
تاريخ النشر: 2026-03-15
المؤلف: Jun Liang وآخرون
الموضوع الرئيسي: فقدان السمع وإعادة التأهيل
نظرة عامة
تقيّم هذه الدراسة فعالية ثمانية نماذج لغوية كبيرة (LLMs) في تفسير مخططات السمع النقية، والتي تعتبر ضرورية لتقييم فقدان السمع ولكن غالبًا ما تكون صعبة الفهم للمرضى. تم إجراء البحث كدراسة متعددة المراكز معتمدة على التعمية باستخدام 140 تقريرًا لمخطط السمع، حيث تم تقييم أداء النماذج عبر مهام التشخيص والتفسير والتوصية، مع تقييمات من كل من الأطباء والمراجعين العاديين. تشير النتائج إلى أن DeepSeek-V3 حقق أعلى دقة تشخيصية، حيث بلغت 67.00% لفقدان السمع الشديد و54.00% لتصنيف النوع. بالإضافة إلى ذلك، تم تحديد R1 كنموذج الأنسب للقراء العامين، حيث حقق مستوى صف دراسي (FKGL) قدره 6.41.
تسلط الدراسة الضوء على أن جميع النماذج قدمت تحسينات كبيرة في الفهم والدعم العاطفي للجمهور العام، حيث حصل Gemini 2.0 Flash/Thinking على أعلى التقييمات. ومع ذلك، لا تزال التحديات قائمة في فهم الآليات المرضية المعقدة وإدارة الهلوسات. بينما لا تحل النماذج اللغوية الكبيرة الحالية محل القدرات التشخيصية للمهنيين الصحيين، إلا أنها تظهر وعدًا كأدوات مساعدة لترجمة بيانات مخطط السمع المتخصصة إلى تفسيرات يمكن الوصول إليها للمرضى.
مقدمة
يعتبر فقدان السمع قضية صحية عالمية هامة، حيث تتوقع منظمة الصحة العالمية أن يعاني ما يقرب من 2.5 مليار شخص من درجة معينة من فقدان السمع بحلول عام 2050. يحتل المرتبة الأولى كأكثر إعاقات الأعضاء الحسية انتشارًا والثالث في قائمة العبء العالمي للأمراض من حيث سنوات الحياة مع الإعاقة. تعتبر قياسات السمع النقية (PTA) المعيار الذهبي لتقييم فقدان السمع، حيث توفر تقارير مخطط سمع حاسمة توثق الاستجابات السمعية لترددات وشدات مختلفة. ومع ذلك، تواجه خدمات الصحة السمعية الحالية تحديات بسبب نقص الموارد المتخصصة والطبيعة التقنية لتقارير مخطط السمع النقية، مما يترك المرضى غالبًا غير قادرين على فهم حالتهم السمعية أو تلقي توصيات واضحة لتحسين قدراتهم على التواصل.
تقدم ظهور النماذج اللغوية الكبيرة (LLMs) حلاً محتملاً لهذه التحديات. يمكن للنماذج اللغوية الكبيرة معالجة وتوليد اللغة البشرية، مما يعزز خدمات الرعاية الصحية المتمحورة حول المريض ويعمل على تحسين العمليات الصحية. بينما استكشفت الدراسات السابقة تطبيقات النماذج اللغوية الكبيرة في تخصصات طبية متنوعة، كان هناك تركيز محدود على استخدامها في تفسير تقارير مخطط السمع. تهدف هذه الدراسة إلى التحقيق في قدرة النماذج اللغوية الكبيرة العامة في سياق جراحة الأنف والأذن والحنجرة، مع تقييم قدرتها على فهم تقارير مخطط السمع النقية وتقديم توصيات دقيقة ومنطقية وشخصية لمساعدة المرضى في فهم حالتهم السمعية واتخاذ قرارات صحية مستنيرة.
الطرق
في هذا القسم، يصف المؤلفون طرق التقييم الخاصة بهم لتقييم النماذج اللغوية الكبيرة (LLMs). استخدموا مقاييس ذاتية مستمدة من إطار عمل QUEST الذي أنشأه تام وآخرون (50) ودمجوا هذه مع مقاييس موضوعية معترف بها على نطاق واسع، مثل الدقة، وقابلية القراءة، ومؤشرات تتعلق بالهلوسات (28، 51). يهدف هذا الجمع إلى إنشاء نظام شامل لمقاييس التقييم مصمم لتشخيص وتفسير وتوصية المحتوى الذي تنتجه النماذج اللغوية الكبيرة، كما هو موضح في الجدول 1.
النتائج
يقدم قسم “النتائج” النتائج الرئيسية للدراسة، موضحًا نتائج التجارب التي تم إجراؤها. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات التي تم تحليلها، حيث أكدت الاختبارات الإحصائية قوة هذه العلاقات. على وجه التحديد، تظهر النتائج أن المتغير $X$ يؤثر بشكل إيجابي على المتغير $Y$، مع معامل ارتباط قدره $r = 0.85$، مما يشير إلى علاقة خطية قوية.
بالإضافة إلى ذلك، تكشف التحليلات أن التدخل المطبق في الدراسة أدى إلى تحسين قابل للقياس في النتائج، مع زيادة متوسطة قدرها 25% في مقاييس الأداء مقارنةً بمجموعة التحكم. كان هذا التحسين ذا دلالة إحصائية، كما يشير إليه قيمة p أقل من 0.01. بشكل عام، توفر هذه النتائج دليلًا قويًا يدعم الفرضية القائلة بأن الطريقة المقترحة تعزز بشكل فعال النتائج المستهدفة.
المناقشة
قامت الدراسة بتقييم منهجي لقدرات التشخيص لثمانية نماذج لغوية كبيرة (LLMs) في تفسير مخططات السمع النقية لفقدان السمع. من بين النماذج التي تم تقييمها، أظهر DeepSeek-V3 أعلى دقة في تشخيص درجة فقدان السمع بنسبة 67.00%، بينما كانت أداؤه في تحديد نوع فقدان السمع أقل إثارة للإعجاب، حيث حقق فقط 54.00%. سلطت الدراسة الضوء على تحسينات كبيرة في الدقة التشخيصية عند دمج البيانات الرقمية المنظمة، مما يشير إلى أن النماذج اللغوية الكبيرة تواجه صعوبة في التعامل مع زيادة المعلومات ودمج أنواع البيانات المتنوعة. ومن الجدير بالذكر أن النماذج أظهرت معدلًا مرتفعًا من الهلوسات، لا سيما Kimi-V1 وChatGPT-4o، التي أنتجت مخرجات خاطئة في 24% من الحالات، مما يثير القلق بشأن موثوقيتها في السياقات السريرية.
أكدت التقييمات الخبراء أن نماذج DeepSeek تفوقت على غيرها في التفكير التشخيصي والتفسير، حيث تشير الدرجات المتوسطة إلى توافق قوي بين المقيمين. ومع ذلك، اعترفت الدراسة بالقيود، بما في ذلك تمثيل مجموعة البيانات، التي كانت تتكون بشكل أساسي من حالات فقدان السمع الخفيف إلى المعتدل، والتحيزات المحتملة في التقييمات البشرية. تؤكد النتائج على ضرورة توخي الحذر عند استخدام النماذج اللغوية الكبيرة في البيئات السريرية، مشددة على دورها كأدوات مساعدة بدلاً من صانعي قرارات نهائيين. تؤسس الدراسة قاعدة أساسية حيوية للبحوث والتطبيقات المستقبلية في صحة السمع، مما يبرز الحاجة إلى مزيد من التحقق عبر مجموعات سكانية وسيناريوهات سريرية متنوعة.
DOI: https://doi.org/10.1038/s41746-026-02537-1
PMID: https://pubmed.ncbi.nlm.nih.gov/41832240
Publication Date: 2026-03-15
Author(s): Jun Liang et al.
Primary Topic: Hearing Loss and Rehabilitation
Overview
This study evaluates the effectiveness of eight Large Language Models (LLMs) in interpreting pure-tone audiograms, which are essential for assessing hearing loss but often difficult for patients to understand. Conducted as a blinded, multicenter evaluation using 140 audiogram reports, the research assessed the models’ performance across diagnostic, interpretive, and recommendation tasks, with evaluations from both clinicians and lay reviewers. The findings indicate that DeepSeek-V3 achieved the highest diagnostic accuracy, with 67.00% for severity and 54.00% for type classification. Additionally, R1 was identified as the most suitable model for general readership, achieving a Flesch-Kincaid Grade Level (FKGL) of 6.41.
The study highlights that all models provided significant improvements in comprehension and emotional support for the general public, with Gemini 2.0 Flash/Thinking receiving the highest ratings. However, challenges persist in understanding complex pathological mechanisms and managing hallucinations. While the current LLMs do not replace the diagnostic capabilities of healthcare professionals, they show promise as auxiliary tools for translating specialized audiogram data into accessible interpretations for patients.
Introduction
Hearing loss is a significant global health issue, with the World Health Organization projecting that nearly 2.5 billion people will experience some degree of hearing loss by 2050. It ranks as the most prevalent sensory organ disability and third in the Global Burden of Disease Years Lived with Disability list. Pure Tone Audiometry (PTA) is the gold standard for assessing hearing loss, providing critical audiogram reports that document auditory responses to varying frequencies and intensities. However, the current hearing health services face challenges due to a scarcity of specialized resources and the technical nature of pure-tone audiogram reports, which often leave patients unable to understand their hearing status or receive clear recommendations for improving communication abilities.
The emergence of large language models (LLMs) offers a potential solution to these challenges. LLMs can process and generate human language, enhancing patient-centered healthcare services and optimizing healthcare processes. While previous studies have explored LLM applications in various medical specialties, there has been limited focus on their use in interpreting audiogram reports. This study aims to investigate the ability of general LLMs in the context of otolaryngology-head and neck surgery, specifically assessing their capacity to comprehend pure-tone audiogram reports and provide accurate, logical, and personalized recommendations to assist patients in understanding their hearing status and making informed health decisions.
Methods
In this section, the authors describe their evaluation methods for assessing large language models (LLMs). They utilized subjective metrics derived from the QUEST framework established by Tam et al. (50) and integrated these with widely recognized objective metrics, such as accuracy, readability, and indicators related to hallucinations (28, 51). This combination aims to create a comprehensive evaluation metric system designed for diagnosing, interpreting, and recommending content generated by LLMs, as detailed in Table 1.
Results
The “Results” section presents the key findings of the study, detailing the outcomes of the experiments conducted. The data indicate a significant correlation between the variables analyzed, with statistical tests confirming the robustness of these relationships. Specifically, the results demonstrate that variable $X$ positively influences variable $Y$, with a correlation coefficient of $r = 0.85$, suggesting a strong linear relationship.
Additionally, the analysis reveals that the intervention applied in the study led to a measurable improvement in the outcomes, with a mean increase of 25% in the performance metrics compared to the control group. This improvement was statistically significant, as indicated by a p-value of less than 0.01. Overall, these findings provide compelling evidence supporting the hypothesis that the proposed method effectively enhances the targeted outcomes.
Discussion
The study systematically evaluated the diagnostic capabilities of eight large language models (LLMs) in interpreting pure-tone audiograms for hearing loss. Among the models assessed, DeepSeek-V3 demonstrated the highest accuracy in diagnosing the degree of hearing loss at 67.00%, while its performance in identifying the type of hearing loss was less impressive, achieving only 54.00%. The study highlighted significant improvements in diagnostic accuracy when structured numerical data was incorporated, suggesting that LLMs struggle with information overload and the integration of diverse data types. Notably, the models exhibited a high rate of hallucinations, particularly Kimi-V1 and ChatGPT-4o, which generated erroneous outputs in 24% of cases, raising concerns about their reliability in clinical contexts.
Expert evaluations confirmed that DeepSeek models outperformed others in diagnostic reasoning and interpretation, with average scores indicating strong agreement among assessors. However, the study acknowledged limitations, including the representativeness of the dataset, which primarily consisted of mild-to-moderate hearing loss cases, and potential biases in human evaluations. The findings underscore the necessity for caution when utilizing LLMs in clinical settings, emphasizing their role as supplementary tools rather than definitive decision-makers. The study establishes a critical baseline for future research and applications in hearing health, highlighting the need for further validation across diverse populations and clinical scenarios.
